Files
Chenjie LuoandClaude Opus 5 19de0075cb Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do?

Type of change: Bug fix

`scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has
no effect on the `lm_eval` task: the value is parsed by `parser.sh`,
printed, and then dropped.

lm-eval's built-in `trtllm` backend
(`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example
switched to in #2066) accepts `**kwargs`, but builds
`KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed
set of keys — `kwargs` is never merged in. So an extra `--model_args`
entry is accepted by the CLI and silently discarded, and the KV cache is
sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There
is no way to fix this from the caller: `--model_args` only yields
scalars, so a `KvCacheConfig` object cannot be passed in either.

On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB
free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization
then OOMs asking for 2.82 GiB.

`examples/llm_eval/lm_eval_trtllm.py` already exists to patch this
backend (its `_parse_logprobs` misaligns TensorRT-LLM's
`prompt_logprobs` by one). It now also injects the fraction into the
`KvCacheConfig` the backend builds, defaulting to 0.8 — the same default
`parser.sh` declares, and below TensorRT-LLM's 0.9.
`huggingface_example.sh` passes the parsed value through in
`--model_args`.

Scoped deliberately to the `lm_eval` path: the `quant` smoke test and
`mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and
`simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are
left as they are.

### Usage

```bash
# Via the example script (parser.sh default 0.8)
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \
    --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \
    --kv_cache_free_gpu_memory_fraction 0.5
```

```bash
# Standalone, via lm-eval's --model_args
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \
    --tasks mmlu --batch_size 8
```

### Testing

- `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed
(lm-eval 0.4.12, no GPU).
- The new tests instantiate the **real** upstream `TRTLLM.__init__`
through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer
stubbed, and assert the engine receives
`KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`;
that an unset key still yields 0.8 rather than 0.9; and that the patch
does not outlive the constructor. Reverting the fix fails 3 of them.
- Tripwire test asserts upstream still neither declares nor forwards the
argument, so this shim gets deleted rather than silently kept once
lm-eval fixes it.
- `pre-commit run --files <changed>` clean (ruff, mypy, bandit,
markdownlint); `bash -n` on the modified script.
- Not run: the GPU end-to-end
`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which
exercises `lm_eval` through the modified script — no GPU in this
environment.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes
from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative;
`parser.sh`'s declared default is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

NVBug 6701763. The 0.9 default on this path arrived with #2066 and was
documented as a known limitation in `examples/llm_eval/README.md` ("the
KV cache uses 90% of free GPU memory rather than 70%"); that note is
replaced by the working knob.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed the TensorRT-LLM evaluation workflow so
`kv_cache_free_gpu_memory_fraction` is correctly passed to the backend.
- The setting now defaults to `0.8`, providing more predictable GPU
memory allocation for KV-cache usage.

- **Documentation**
- Updated the TensorRT-LLM evaluation example and usage guidance to
describe the KV-cache memory setting and its default behavior.
- Updated the Hugging Face example to pass the configured KV-cache
memory fraction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 13:06:51 -07:00
..
2025-01-29 01:45:47 +05:30
2025-03-03 22:54:22 +05:30
2025-03-03 22:54:22 +05:30
2025-03-03 22:54:22 +05:30
2025-05-08 23:43:51 +05:30

Evaluation scripts for LLM tasks

This folder includes popular 3rd-party LLM benchmarks for LLM accuracy evaluation.

The following instructions show how to evaluate the Model Optimizer quantized LLM with the benchmarks, including the TensorRT-LLM deployment.

NeMo Evaluator

NeMo Evaluator is the recommended way to evaluate a large choice of benchmarks on quantized checkpoints generated from hf_ptq. Quantized checkpoints can be served with TensorRT-LLM, vLLM, or SGLang and then evaluated using NeMo Evaluator.

LM-Eval-Harness

LM-Eval-Harness provides a unified framework to test generative language models on a large number of different evaluation tasks.

The supported eval tasks are here.

For guidance on shortening research iteration cycles while preserving meaningful model comparisons, see ModelOpt for Researchers: Fast Experimentation Workflows.

Baseline

Both standard HuggingFace models and heterogeneous pruned checkpoints produced by Puzzletron are supported.

  • For models which fit on a single GPU:
python lm_eval_hf.py --model hf --model_args pretrained=<HF model folder or model card> --tasks <comma separated tasks> --batch_size 4

For a quick smoke test, add --limit 10 to any of the above commands to evaluate on only 10 samples per task.

  • To fit one model across multiple GPUs (model sharding) and enable larger batches that may speed up evaluation:
python lm_eval_hf.py --model hf --model_args pretrained=<HF model folder or model card>,parallelize=True --tasks <comma separated tasks> --batch_size 4

Note (Slurm interactive nodes): On Slurm interactive nodes, WORLD_SIZE is set to the number of available GPUs in the shell environment. Running python directly causes lm_eval to hang waiting for peer ranks that were never spawned. Prepend WORLD_SIZE=1 to the python commands above to fix this. This does not limit GPU usage — parallelize=True independently enables model parallelism across all available GPUs within the single process. The accelerate launch command manages WORLD_SIZE itself and does not require this workaround.

  • For data-parallel evaluation with model-sharding:

--num_processes controls how many model copies evaluate samples concurrently. More copies usually make evaluation faster but leave fewer GPUs for each copy. With N GPUs, each copy uses approximately N / num_processes GPUs. For example, on 8 GPUs, 8 processes run eight single-GPU copies. Choose the largest number of processes for which each model copy fits.

accelerate launch --multi_gpu --num_processes <num_copies_of_your_model> \
    lm_eval_hf.py --model hf \
    --tasks <comma separated tasks> \
    --model_args pretrained=<HF model folder or model card>,parallelize=True \
    --batch_size 4

Quantized (simulated)

  • For simulated quantization with any of the default quantization formats:

Multi-GPU evaluation without data-parallelism:

# MODELOPT_QUANT_CFG: Choose from [INT8_SMOOTHQUANT_CFG|FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG|INT4_AWQ_CFG|W4A8_AWQ_BETA_CFG|MXFP8_DEFAULT_CFG]
python lm_eval_hf.py --model hf \
    --tasks <comma separated tasks> \
    --model_args pretrained=<HF model folder or model card>,parallelize=True \
    --quant_cfg <MODELOPT_QUANT_CFG> \
    --batch_size 4

NOTE: MXFP8_DEFAULT_CFG is one the OCP Microscaling Formats (MX Formats) family which defines a set of block-wise dynamic quantization formats. The specifications can be found in the official documentation. Currently we support all MX formats for simulated quantization, including MXFP8 (E5M2, E4M3), MXFP6 (E3M2, E2M3), MXFP4, MXINT8. However, only MXFP8 (E4M3) is in our example configurations, users can create their own configurations for other MX formats by simply modifying the num_bits field in the MXFP8_DEFAULT_CFG.

NOTE: ModelOpt's triton kernels give faster NVFP4 simulated quantization. For details, please see the installation guide.

For data-parallel evaluation, launch with accelerate launch --multi_gpu --num_processes <num_copies_of_your_model> (as shown earlier).

  • For simulated optimal per-layer quantization with auto_quantize:

Multi-GPU evaluation without data-parallelism:

# MODELOPT_QUANT_CFG_TO_SEARCH: Comma-separated list of the formats auto_quantize searches over.
# Pick each one from [INT8_SMOOTHQUANT_CFG|FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG|INT4_AWQ_CFG|W4A8_AWQ_BETA_CFG|MXFP8_DEFAULT_CFG|NONE],
# where NONE lets auto_quantize leave a layer unquantized.
# EFFECTIVE_BITS: Effective bits constraint for auto_quantize

# Example settings for an optimally quantized model with W4A8 & FP8 with effective bits of 4.8:
# MODELOPT_QUANT_CFG_TO_SEARCH=W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG,NONE
# EFFECTIVE_BITS=4.8

python lm_eval_hf.py --model hf \
    --tasks <comma separated tasks> \
    --model_args pretrained=<HF model folder or model card>,parallelize=True \
    --quant_cfg <MODELOPT_QUANT_CFG_TO_SEARCH> \
    --auto_quantize_bits <EFFECTIVE_BITS> \
    --batch_size 4

For data-parallel evaluation, launch with accelerate launch --multi_gpu --num_processes <num_copies_of_your_model> (as shown earlier).

  • If evaluating encoder-decoder models such as T5, keep --model hf: lm-eval detects the encoder-decoder architecture from config.json. There is no hf-seq2seq backend in the supported lm-eval versions (>= 0.4.12); add backend=seq2seq to --model_args only for checkpoints lm-eval cannot classify on its own.
# MODELOPT_QUANT_CFG: Choose from [INT8_SMOOTHQUANT_CFG|FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG|INT4_AWQ_CFG|W4A8_AWQ_BETA_CFG|MXFP8_DEFAULT_CFG]
python lm_eval_hf.py --model hf --model_args pretrained=t5-small --quant_cfg <MODELOPT_QUANT_CFG> --tasks <comma separated tasks> --batch_size 4

If trust_remote_code needs to be true, please append the command with the --trust_remote_code flag.

TensorRT-LLM

Uses the trtllm backend built into lm-eval (>= 0.4.12), which loads the quantized checkpoint directly with the TensorRT-LLM LLM API.

python lm_eval_trtllm.py --model trtllm \
    --model_args model=<Quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096,max_output_len=512,kv_cache_free_gpu_memory_fraction=0.8 \
    --tasks <comma separated tasks> \
    --batch_size <max batch size>

NOTE: Loglikelihood tasks (mmlu, hellaswag, arc, ...) need TensorRT-LLM >= 1.3.0rc11, which is when the engine started returning the requested token in every prompt_logprobs entry. Earlier releases return only the top-1 token per position, so a continuation token's logprob cannot be recovered and the run aborts with a clear error. Generative tasks (gsm8k, ifeval) are unaffected.

NOTE: Set max_input_len and max_output_len explicitly. They default to 2048 and 512, and prompts longer than max_input_len are silently truncated — 5-shot MMLU or gsm8k prompts exceed 2048 tokens. max_seq_len of the engine is their sum.

NOTE: tensor_parallel_size defaults to 1; set it to the number of GPUs the checkpoint needs. pipeline_parallel_size is also supported.

NOTE: Use lm_eval_trtllm.py rather than the plain lm_eval CLI. lm-eval 0.4.12's trtllm backend misaligns TensorRT-LLM's prompt_logprobs by one position, so every loglikelihood task (hellaswag, mmlu, arc, ...) fails with a KeyError; lm_eval_trtllm.py overrides the alignment. It goes away once the fix lands upstream.

NOTE: kv_cache_free_gpu_memory_fraction is the share of the GPU memory left after loading the weights that the KV cache may take. TensorRT-LLM's own default of 0.9 can leave too little room for the prompt_logprobs buffers and OOM on a large-memory GPU, so lm_eval_trtllm.py defaults it to 0.8; lm-eval's backend drops the key, which is why this entry point forwards it. huggingface_example.sh passes its --kv_cache_free_gpu_memory_fraction (default 0.8) through.

NOTE: Other than the KV cache fraction, the backend forwards only a fixed set of arguments to TensorRT-LLM, so the remaining tuning the old lm_eval_tensorrt_llm.py applied is not reachable: expert parallelism is left at the TensorRT-LLM default (MoE checkpoints can fail in DeepEP kernels on some GPUs, e.g. SM 12.0). Lower tensor_parallel_size if you hit that.

lm_eval_tensorrt_llm.py (--model trt-llm) has been removed; use the command above.

MMLU

Massive Multitask Language Understanding. A score (0-1, higher is better) will be printed at the end of the benchmark.

Setup

Download data

mkdir -p data
wget --connect-timeout=20 --read-timeout=60 --tries=3 -c \
    https://huggingface.co/datasets/cais/mmlu/resolve/c30699e8356da336a370243923dbaf21066bb9fe/data.tar -O data/mmlu.tar
tar -xf data/mmlu.tar -C data && mv data/data data/mmlu

Run the commands below from examples/llm_eval; mmlu.py resolves its default --data_dir data/mmlu relative to the current directory.

Baseline

python mmlu.py --model_name causal --model_path <HF model folder or model card>

Quantized (simulated)

# MODELOPT_QUANT_CFG: Choose from [INT8_SMOOTHQUANT_CFG|FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG|INT4_AWQ_CFG|W4A8_AWQ_BETA_CFG|MXFP8_DEFAULT_CFG]
python mmlu.py --model_name causal --model_path <HF model folder or model card> --quant_cfg <MODELOPT_QUANT_CFG>

auto_quantize (simulated)

# MODELOPT_QUANT_CFG_TO_SEARCH: Comma-separated list of the formats auto_quantize searches over.
# Pick each one from [INT8_SMOOTHQUANT_CFG|FP8_DEFAULT_CFG|NVFP4_DEFAULT_CFG|INT4_AWQ_CFG|W4A8_AWQ_BETA_CFG|MXFP8_DEFAULT_CFG|NONE],
# where NONE lets auto_quantize leave a layer unquantized.
# EFFECTIVE_BITS: Effective bits constraint for auto_quantize

# Example settings for an optimally quantized model with W4A8 & FP8 with effective bits of 4.8:
# MODELOPT_QUANT_CFG_TO_SEARCH=W4A8_AWQ_BETA_CFG,FP8_DEFAULT_CFG,NONE
# EFFECTIVE_BITS=4.8

python mmlu.py --model_name causal --model_path <HF model folder or model card> --quant_cfg $MODELOPT_QUANT_CFG_TO_SEARCH --auto_quantize_bits $EFFECTIVE_BITS --batch_size 4

Evaluate with TensorRT-LLM

python mmlu.py --model_name causal --model_path <HF model folder or model card> --checkpoint_dir <Quantized checkpoint dir>

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark of LLMs for code that continuously collects new problems over time.

We support running LiveCodeBench against a local running OpenAI API compatible server. For example, quantized TensorRT-LLM checkpoint or engine can be loaded using trtllm-serve command. Once the local server is up, the following command can be used to run the LiveCodeBench:

bash run_livecodebench.sh <custom defined model name> <prompt batch size in parallel> <max output tokens> <local model server port>

Simple Evals

Simple Evals is a lightweight library for evaluating language models published from OpenAI. This eval includes "simpleqa", "mmlu", "math", "gpqa", "mgsm", "drop" and "humaneval" benchmarks.

Similarly, we support running simple evals against a local running OpenAI API compatible server. Once the local server is up, the following command can be used to run the Simple Evals:

bash run_simple_eval.sh <custom defined model name> <comma separated eval names> <max output tokens> <local model server port> [num examples per eval]

The optional fifth argument caps the number of examples per eval (--examples); omit it to run the full eval.

Customize quantization method for evaluation

An example of customized quantization config is shown in quantization_utils.py. It allows users to test accuracy of a custom method without the need of modifying the whole deployment framework, e.g., TensorRT-LLM, vLLM, SGLang, etc. Users can disable quantization of specific layers to debug the cause of accuracy drop, or explore a promising new quantization method.

python lm_eval_hf.py --model hf \
    --tasks <comma separated tasks> \
    --model_args pretrained=<HF model folder or model card>,parallelize=True \
    --quant_cfg MY_QUANT_CONFIG \
    --batch_size 4

Evaluating with LM-Eval-Harness via vLLM

The run_lm_eval_vllm.sh script provides a convenient way to run evaluations using the lm-evaluation-harness library against a model served with vLLM's OpenAI-compatible API endpoint.

This is useful for evaluating quantized models deployed with vLLM or any model served via its OpenAI API interface. More importantly, for new models that are neither supported natively by vLLM or Transformers, but Transformers compatible, they can still be evaluated with vLLM's endpoint! By Transformers compatible, it needs to satisfy the following:

  • The model directory must have the correct structure (e.g. config.json is present)
  • config.json must contain auto_map.AutoModel.
  • Customization should be done in the base model (e.g. in MyModel, not MyModelForCausalLM).

Prerequisites

  1. Install vLLM: Follow the installation instructions at https://docs.vllm.ai/en/latest/getting_started/installation.html.

Usage

  1. Start the vLLM OpenAI-compatible Server: In a separate terminal, launch the vLLM server with your desired model. For example:

    # Example using vLLM's built-in server
    vllm serve <your_model_name_or_path> \
        --port 8000 \
        --tensor-parallel-size <tp_size> # Adjust as needed
    

    Replace <your_model_name_or_path> with the actual model identifier (e.g., Qwen/Qwen3-30B-A3B) and adjust the --port and --tensor-parallel-size if necessary. You may also need to disable vllm v1 by export VLLM_USE_V1=0 if you encounter issues.

    To serve a modelopt quantized model, add --quantization modelopt, for example:

    # Example using vLLM's built-in server
    vllm serve nvidia/Llama-3.1-8B-Instruct-FP8 \
        --quantization modelopt \
        --port 8000 \
        --tensor-parallel-size <tp_size> # Adjust as needed
    

    To generate the quantized model such as nvidia/Llama-3.1-8B-Instruct-FP8, please refer to instructions here. Note currently modelopt quantized model support in vLLM is limited, we are working on expanding the model and quant formats support.

  2. Make the script executable (if not already):

    chmod +x run_lm_eval_vllm.sh
    
  3. Run the evaluation script from the examples/llm_eval directory:

    ./run_lm_eval_vllm.sh <model_name> [port] [task]
    
    • <model_name>: The name of the model being served (this is passed to lm_eval, e.g., Qwen/Qwen3-30B-A3B).
    • [port]: (Optional) The port the vLLM server is listening on. Defaults to 8000. Note, it must match the number when launch the server.
    • [task]: (Optional) The lm_eval task(s) to run. Defaults to mmlu. Can be a single task or a comma-separated list (e.g., "mmlu,hellaswag").

Examples

  • Evaluate Qwen3-30B-A3B on MMLU (default task and port):

    # Assumes vLLM server running with Qwen/Qwen3-30B-A3B on port 8000
    ./run_lm_eval_vllm.sh Qwen/Qwen3-30B-A3B
    
  • Evaluate a model on Hellaswag using port 8001:

    # Assumes vLLM server running with <model_name> on port 8001
    ./run_lm_eval_vllm.sh <model_name> 8001 hellaswag
    
  • Evaluate on multiple tasks:

    # Assumes vLLM server running with <model_name> on port 8000
    ./run_lm_eval_vllm.sh <model_name> 8000 "arc_easy,winogrande"