mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do? Type of change: Bug fix `scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has no effect on the `lm_eval` task: the value is parsed by `parser.sh`, printed, and then dropped. lm-eval's built-in `trtllm` backend (`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example switched to in #2066) accepts `**kwargs`, but builds `KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed set of keys — `kwargs` is never merged in. So an extra `--model_args` entry is accepted by the CLI and silently discarded, and the KV cache is sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There is no way to fix this from the caller: `--model_args` only yields scalars, so a `KvCacheConfig` object cannot be passed in either. On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization then OOMs asking for 2.82 GiB. `examples/llm_eval/lm_eval_trtllm.py` already exists to patch this backend (its `_parse_logprobs` misaligns TensorRT-LLM's `prompt_logprobs` by one). It now also injects the fraction into the `KvCacheConfig` the backend builds, defaulting to 0.8 — the same default `parser.sh` declares, and below TensorRT-LLM's 0.9. `huggingface_example.sh` passes the parsed value through in `--model_args`. Scoped deliberately to the `lm_eval` path: the `quant` smoke test and `mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and `simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are left as they are. ### Usage ```bash # Via the example script (parser.sh default 0.8) scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \ --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \ --kv_cache_free_gpu_memory_fraction 0.5 ``` ```bash # Standalone, via lm-eval's --model_args python lm_eval_trtllm.py --model trtllm \ --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \ --tasks mmlu --batch_size 8 ``` ### Testing - `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed (lm-eval 0.4.12, no GPU). - The new tests instantiate the **real** upstream `TRTLLM.__init__` through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer stubbed, and assert the engine receives `KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`; that an unset key still yields 0.8 rather than 0.9; and that the patch does not outlive the constructor. Reverting the fix fails 3 of them. - Tripwire test asserts upstream still neither declares nor forwards the argument, so this shim gets deleted rather than silently kept once lm-eval fixes it. - `pre-commit run --files <changed>` clean (ruff, mypy, bandit, markdownlint); `bash -n` on the modified script. - Not run: the GPU end-to-end `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which exercises `lm_eval` through the modified script — no GPU in this environment. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative; `parser.sh`'s declared default is unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information NVBug 6701763. The 0.9 default on this path arrived with #2066 and was documented as a known limitation in `examples/llm_eval/README.md` ("the KV cache uses 90% of free GPU memory rather than 70%"); that note is replaced by the working knob. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed the TensorRT-LLM evaluation workflow so `kv_cache_free_gpu_memory_fraction` is correctly passed to the backend. - The setting now defaults to `0.8`, providing more predictable GPU memory allocation for KV-cache usage. - **Documentation** - Updated the TensorRT-LLM evaluation example and usage guidance to describe the KV-cache memory setting and its default behavior. - Updated the Hugging Face example to pass the configured KV-cache memory fraction. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
2aefe08f20
commit
19de0075cb
@@ -296,7 +296,7 @@ if [[ $TASKS =~ "lm_eval" ]]; then
|
||||
# explicitly; the engine's max_seq_len is max_input_len + max_output_len.
|
||||
python lm_eval_trtllm.py \
|
||||
--model trtllm \
|
||||
--model_args "model=$SAVE_PATH,tokenizer=$MODEL_ABS_PATH,tensor_parallel_size=$LM_EVAL_TP,max_batch_size=$BUILD_MAX_BATCH_SIZE,max_gen_toks=$BUILD_MAX_OUTPUT_LEN,max_input_len=$BUILD_MAX_INPUT_LEN,max_output_len=$BUILD_MAX_OUTPUT_LEN" \
|
||||
--model_args "model=$SAVE_PATH,tokenizer=$MODEL_ABS_PATH,tensor_parallel_size=$LM_EVAL_TP,max_batch_size=$BUILD_MAX_BATCH_SIZE,max_gen_toks=$BUILD_MAX_OUTPUT_LEN,max_input_len=$BUILD_MAX_INPUT_LEN,max_output_len=$BUILD_MAX_OUTPUT_LEN,kv_cache_free_gpu_memory_fraction=$KV_CACHE_FREE_GPU_MEMORY_FRACTION" \
|
||||
--tasks $LM_EVAL_TASKS \
|
||||
--batch_size $BUILD_MAX_BATCH_SIZE $lm_eval_flags | tee $LM_EVAL_RESULT
|
||||
|
||||
|
||||
@@ -117,7 +117,7 @@ checkpoint directly with the TensorRT-LLM LLM API.
|
||||
|
||||
```sh
|
||||
python lm_eval_trtllm.py --model trtllm \
|
||||
--model_args model=<Quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096,max_output_len=512 \
|
||||
--model_args model=<Quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096,max_output_len=512,kv_cache_free_gpu_memory_fraction=0.8 \
|
||||
--tasks <comma separated tasks> \
|
||||
--batch_size <max batch size>
|
||||
```
|
||||
@@ -140,11 +140,18 @@ python lm_eval_trtllm.py --model trtllm \
|
||||
> loglikelihood task (hellaswag, mmlu, arc, ...) fails with a `KeyError`;
|
||||
> `lm_eval_trtllm.py` overrides the alignment. It goes away once the fix lands upstream.
|
||||
|
||||
> **_NOTE:_** The backend forwards only a fixed set of arguments to TensorRT-LLM, so the
|
||||
> tuning the old `lm_eval_tensorrt_llm.py` applied is not reachable: expert parallelism is
|
||||
> left at the TensorRT-LLM default (MoE checkpoints can fail in DeepEP kernels on some
|
||||
> GPUs, e.g. SM 12.0) and the KV cache uses 90% of free GPU memory rather than 70%. Lower
|
||||
> `tensor_parallel_size` if you hit either.
|
||||
> **_NOTE:_** `kv_cache_free_gpu_memory_fraction` is the share of the GPU memory left after
|
||||
> loading the weights that the KV cache may take. TensorRT-LLM's own default of 0.9 can leave
|
||||
> too little room for the `prompt_logprobs` buffers and OOM on a large-memory GPU, so
|
||||
> `lm_eval_trtllm.py` defaults it to 0.8; lm-eval's backend drops the key, which is why this
|
||||
> entry point forwards it. `huggingface_example.sh` passes its
|
||||
> `--kv_cache_free_gpu_memory_fraction` (default 0.8) through.
|
||||
|
||||
> **_NOTE:_** Other than the KV cache fraction, the backend forwards only a fixed set of
|
||||
> arguments to TensorRT-LLM, so the remaining tuning the old `lm_eval_tensorrt_llm.py`
|
||||
> applied is not reachable: expert parallelism is left at the TensorRT-LLM default (MoE
|
||||
> checkpoints can fail in DeepEP kernels on some GPUs, e.g. SM 12.0). Lower
|
||||
> `tensor_parallel_size` if you hit that.
|
||||
|
||||
`lm_eval_tensorrt_llm.py` (`--model trt-llm`) has been removed; use the command above.
|
||||
|
||||
|
||||
@@ -16,18 +16,22 @@
|
||||
"""Run lm-evaluation-harness against a TensorRT-LLM checkpoint.
|
||||
|
||||
Entry point around lm-eval's built-in ``trtllm`` backend
|
||||
(``lm_eval.models.trtllm_causallms``, new in 0.4.12). It exists only to correct that
|
||||
backend's ``prompt_logprobs`` handling -- everything else is upstream. Drop this file and
|
||||
call ``lm_eval`` directly once the fix lands upstream.
|
||||
(``lm_eval.models.trtllm_causallms``, new in 0.4.12). It exists to correct that backend's
|
||||
``prompt_logprobs`` handling and to forward ``kv_cache_free_gpu_memory_fraction`` to
|
||||
TensorRT-LLM -- everything else is upstream. Drop this file and call ``lm_eval`` directly
|
||||
once both land upstream.
|
||||
|
||||
python lm_eval_trtllm.py --model trtllm \
|
||||
--model_args model=<quantized checkpoint dir>,tokenizer=<HF model folder>,\
|
||||
tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096 \
|
||||
tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096,\
|
||||
kv_cache_free_gpu_memory_fraction=0.8 \
|
||||
--tasks <comma separated tasks> --batch_size <max batch size>
|
||||
"""
|
||||
|
||||
import sys
|
||||
from functools import partial
|
||||
from importlib.metadata import version
|
||||
from importlib.util import find_spec
|
||||
|
||||
from lm_eval.__main__ import cli_evaluate
|
||||
from packaging.version import Version
|
||||
@@ -36,6 +40,7 @@ if Version(version("lm_eval")) < Version("0.4.12"):
|
||||
# 0.4.12 is the first release shipping lm_eval.models.trtllm_causallms.
|
||||
raise ImportError(f"lm_eval_trtllm.py requires lm-eval >= 0.4.12; found {version('lm_eval')}.")
|
||||
|
||||
from lm_eval.models import trtllm_causallms
|
||||
from lm_eval.models.trtllm_causallms import TRTLLM
|
||||
|
||||
# TensorRT-LLM only started passing the prompt token ids into `compute_logprobs` in
|
||||
@@ -125,6 +130,42 @@ _UPSTREAM_PARSE_LOGPROBS = TRTLLM._parse_logprobs
|
||||
TRTLLM._parse_logprobs = staticmethod(_parse_logprobs)
|
||||
|
||||
|
||||
# lm-eval's backend builds `KvCacheConfig(enable_block_reuse=False)` itself and passes
|
||||
# `LLM(...)` a fixed set of keys, dropping every other `--model_args` entry, so patching the
|
||||
# class it calls is the only way to size the KV cache. Left at TensorRT-LLM's default of
|
||||
# 0.9, the cache leaves too little room for the `prompt_logprobs` buffers and a
|
||||
# loglikelihood run dies with a CUDA OOM on a large-memory GPU.
|
||||
#
|
||||
# Rebinding a module global is process-wide, but `evaluator.py` builds exactly one model per
|
||||
# run, and nothing introspects the constructor it calls -- `create_from_arg_obj` just does
|
||||
# `cls(**model_args)` -- so nothing else can observe the swap.
|
||||
if find_spec("tensorrt_llm") and not hasattr(trtllm_causallms, "KvCacheConfig"):
|
||||
# Guarded on find_spec because without tensorrt_llm the name is absent by design, and
|
||||
# the backend's own "package is not installed" error is the useful one.
|
||||
raise RuntimeError(
|
||||
"lm_eval.models.trtllm_causallms no longer imports KvCacheConfig at module scope, so "
|
||||
f"the KV cache size cannot be set; the backend changed shape in lm-eval "
|
||||
f"{version('lm_eval')}. Recheck whether this file is still needed."
|
||||
)
|
||||
|
||||
_UPSTREAM_INIT = TRTLLM.__init__
|
||||
|
||||
|
||||
def _init(self, *args, kv_cache_free_gpu_memory_fraction: float = 0.8, **kwargs) -> None:
|
||||
"""``TRTLLM.__init__``, with the KV cache share of free GPU memory made settable."""
|
||||
kv_cache_config = trtllm_causallms.KvCacheConfig
|
||||
trtllm_causallms.KvCacheConfig = partial(
|
||||
kv_cache_config, free_gpu_memory_fraction=float(kv_cache_free_gpu_memory_fraction)
|
||||
)
|
||||
try:
|
||||
_UPSTREAM_INIT(self, *args, **kwargs)
|
||||
finally:
|
||||
trtllm_causallms.KvCacheConfig = kv_cache_config
|
||||
|
||||
|
||||
TRTLLM.__init__ = _init
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
# Warn up front so an unusable container is obvious before the model loads, but do not
|
||||
# abort: generative tasks are unaffected by the old prompt_logprobs layout.
|
||||
|
||||
Reference in New Issue
Block a user