Files
Keval MorabiaandClaude Opus 4.8 01415c2788 fix(llm_eval): repair test_qwen3_eval_fp8 end-to-end (#1650)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` was
silently passing while its evals crashed, then began failing as a
timeout. This repairs the whole pipeline:

- **lm_eval `IndexError` (root cause):** TRT-LLM KV-cache prefix reuse
returns truncated `context_logits` for shared-prefix requests (e.g.
hellaswag's one-context / many-endings), which breaks `parse_logprobs`.
Add an `enable_kv_cache_reuse` flag to `modelopt.deploy.llm.LLM`
(default `True`, unchanged) and disable it for the eval deployment so
full-length context logits are returned.
- **Silent CI green:** `python eval.py | tee result.txt` returns `tee`'s
exit code, so a crashing eval was masked. Add `set -o pipefail` to
`huggingface_example.sh` so failures fail the test.
- **Long-prompt overflows:** with the tiny test model's toy tokenizer,
gsm8k/MMLU prompts exceed `max_seq_len`. Bump test
`max_position_embeddings` to 8192, skip MMLU prompts that don't fit even
at zero-shot, and add an MMLU sample limit (`--mmlu_limit`).
- **human-eval build failures:** install with `--no-build-isolation`
(`pkg_resources` is absent in pip's isolated build env), patch its
malformed `console_scripts` entry point, and pin the clone.
- **Cleanups:** gate the post-quant `run_tensorrt_llm.py` smoke test
behind the `quant` task (eval tasks deploy on their own; ~45s saved for
eval-only runs); replace the SIGPIPE-prone serve-readiness `tail -f |
while` with a poll loop (required under `pipefail`).

### Usage

N/A — example/test fix.

### Testing

All four eval tasks verified end-to-end in the CI container (TRT-LLM
1.3.0rc17, RTX 6000 Ada): lm_eval (hellaswag + gsm8k), MMLU, and
simple_eval (humaneval) all complete with exit 0 and no
`IndexError`/overflow. Cold full run ≈ 340s on this GPU.

CI test on 2-gpu:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/27154417497/job/80153551154

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new `enable_kv_cache_reuse`
defaults to current behavior; new script flags are optional)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: N/A (fixes and strengthens an
existing test)
- Did you update Changelog?: N/A (bug fix to examples/tests)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

The full test runs ~340s on an RTX 6000 Ada; CI runners are historically
slower, while `@pytest.mark.timeout` is set to 600 — worth watching the
first CI run and bumping if it's close.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added an option to limit MMLU evaluation length.

* **Bug Fixes**
* Disabled KV-cache prefix reuse for evaluations needing per-token
context logits to prevent truncated/incorrect logprobs.
* Skip examples whose prompts remain too long; warn and report accuracy
as NaN if all examples are skipped.

* **Chores / Scripts**
* Improved example scripts for reproducible installs, patched entry
point handling, pipeline failure detection, conditional test invocation,
polling-based log wait, and a new CLI flag for MMLU limits.

* **Tests**
* Increased timeout and prompt headroom; capped MMLU smoke tests for
speed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 18:11:51 +00:00
..
2025-01-29 01:45:47 +05:30