Files
Model-Optimizer/examples
Chenjie LuoandClaude Opus 5 19de0075cb Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do?

Type of change: Bug fix

`scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has
no effect on the `lm_eval` task: the value is parsed by `parser.sh`,
printed, and then dropped.

lm-eval's built-in `trtllm` backend
(`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example
switched to in #2066) accepts `**kwargs`, but builds
`KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed
set of keys — `kwargs` is never merged in. So an extra `--model_args`
entry is accepted by the CLI and silently discarded, and the KV cache is
sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There
is no way to fix this from the caller: `--model_args` only yields
scalars, so a `KvCacheConfig` object cannot be passed in either.

On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB
free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization
then OOMs asking for 2.82 GiB.

`examples/llm_eval/lm_eval_trtllm.py` already exists to patch this
backend (its `_parse_logprobs` misaligns TensorRT-LLM's
`prompt_logprobs` by one). It now also injects the fraction into the
`KvCacheConfig` the backend builds, defaulting to 0.8 — the same default
`parser.sh` declares, and below TensorRT-LLM's 0.9.
`huggingface_example.sh` passes the parsed value through in
`--model_args`.

Scoped deliberately to the `lm_eval` path: the `quant` smoke test and
`mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and
`simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are
left as they are.

### Usage

```bash
# Via the example script (parser.sh default 0.8)
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \
    --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \
    --kv_cache_free_gpu_memory_fraction 0.5
```

```bash
# Standalone, via lm-eval's --model_args
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \
    --tasks mmlu --batch_size 8
```

### Testing

- `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed
(lm-eval 0.4.12, no GPU).
- The new tests instantiate the **real** upstream `TRTLLM.__init__`
through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer
stubbed, and assert the engine receives
`KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`;
that an unset key still yields 0.8 rather than 0.9; and that the patch
does not outlive the constructor. Reverting the fix fails 3 of them.
- Tripwire test asserts upstream still neither declares nor forwards the
argument, so this shim gets deleted rather than silently kept once
lm-eval fixes it.
- `pre-commit run --files <changed>` clean (ruff, mypy, bandit,
markdownlint); `bash -n` on the modified script.
- Not run: the GPU end-to-end
`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which
exercises `lm_eval` through the modified script — no GPU in this
environment.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes
from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative;
`parser.sh`'s declared default is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

NVBug 6701763. The 0.9 default on this path arrived with #2066 and was
documented as a known limitation in `examples/llm_eval/README.md` ("the
KV cache uses 90% of free GPU memory rather than 70%"); that note is
replaced by the working knob.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed the TensorRT-LLM evaluation workflow so
`kv_cache_free_gpu_memory_fraction` is correctly passed to the backend.
- The setting now defaults to `0.8`, providing more predictable GPU
memory allocation for KV-cache usage.

- **Documentation**
- Updated the TensorRT-LLM evaluation example and usage guidance to
describe the KV-cache memory setting and its default behavior.
- Updated the Hugging Face example to pass the configured KV-cache
memory fraction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 13:06:51 -07:00
..