Files
Keval MorabiaandClaude Opus 5 6a2ae5a25b Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been
failing with `Failed: Timeout (>900.0s) from pytest-timeout` on
unrelated branches (runs 33027056246 and the one for `ad83a428`, while
the 2026-08-24 nightly passed). It is not the test being slow — it is
the harness deadlocking, and the deadlock also destroys the diagnostics
that would explain the underlying kill.

**Mechanism.** The traceback shows `self = <Popen: returncode: -9 args:
['scripts/huggingface_example.sh', ...]>` while still blocked in
`stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends
SIGKILL — pytest-timeout raises in the main thread, and the test's
`finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM
kill is the likely source). But `subprocess.run(..., stdout=PIPE,
stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and
a surviving grandchild (the TRT-LLM serve/build worker) still holds the
write end. EOF never arrives, so the test blocks until the 900 s alarm.
Because the pipe is never drained, **every line of child output is
discarded**, which is why the CI log says nothing about what the script
was doing when it died.

Reduced to a self-contained reproducer:

```python
script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$"
subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20)
# -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost
```

**Fix.** `_run_capturing` now starts the command in its own session,
drains its output on a reader thread (so logs stream as they arrive
instead of being buffered until the end), waits on the *process*, and
kills the process group if descendants still hold the pipe after a 30 s
grace period. A killed launcher now fails in seconds with its logs
intact instead of silently burning the test's whole timeout.

This does not fix whatever kills the script; it makes it diagnosable.
Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against
its 900 s mark** on the last green nightly, so it is fragile regardless
and may want its work trimmed or its budget raised once the logs show
where the time goes.

### Usage

```python
# unchanged public API
run_example_command(cmd_parts, example_path="llm_eval")
```

### Testing

Verified against the reproducer above and on the normal paths:

| scenario | before | after |
| --- | --- | --- |
| launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900
s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured |
| the surviving descendant | keeps running | killed with the process
group (stopped ticking, 20 -> 20 bytes) |
| normal exit | ok | `rc=0`, stdout and stderr interleaved in order |
| non-zero exit | ok | `rc=3`, output captured |

The example-test suites that use this helper run through the same code
path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised
it on nemo:26.08 in the branch this was extracted from.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `_run_capturing` keeps its
`(returncode, output)` contract; only the buffering strategy changed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — this is test
infrastructure; the scenario needs a process that outlives a SIGKILLed
parent, which is awkward to assert in CI. Verified manually with the
reproducer above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test-infrastructure fix.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved command execution reliability with real-time output capture.
* Ensured lingering child processes are cleaned up after commands exit
or are interrupted.
  * Added warnings when forced cleanup may truncate output.
  * Prevented hangs when descendant processes keep output streams open.

* **Documentation**
* Updated MMLU setup instructions to use the Hugging Face dataset
repository.
  * Improved Windows instructions by explicitly using `curl.exe`.

* **Examples**
* Improved MMLU downloads with retries, separate timeouts, resume
support, and automatic temporary-file cleanup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

### Update: the second half of the failure

With the streaming fix in place, the next CI run showed the *actual*
cause, which the old code had been hiding. That run blocked in
`process.wait()` with `<Popen: returncode: None ...>` — child alive and
working, not the previous dead-child pipe deadlock — and the now-visible
output was:

```
--2026-08-27 18:09:20--  (try: 4)  https://people.eecs.berkeley.edu/~hendrycks/data.tar
Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out.
Retrying.
--2026-08-27 18:11:40--  (try: 5)  ...
```

`huggingface_example.sh` downloads the MMLU tarball from
`people.eecs.berkeley.edu`, that host stopped answering around
2026-08-25, and wget's default retry policy (20 tries, ~2 min per
connect timeout) consumed the whole 900 s budget. Not runner-specific:
the URL also times out from a developer workstation, and the nightlies
flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage.

So this PR now carries both halves of the same failure:

1. the harness no longer deadlocks and no longer swallows the logs
(`985809cc2d`), and
2. the MMLU data comes from HuggingFace's copy of the same tarball, with
bounded retries (`40f1d89154`).

The mirror is byte-for-byte the same dataset in the same layout the
script already expects — verified by running the exact download/extract
commands:

```
https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar  ->  HTTP 200, 166 MB
data/mmlu/{dev,test,val}/  ->  57 subject CSVs each, plus auxiliary_train/
```

`wget --timeout=20 --tries=3` plus an explicit error means the next
dataset-host outage fails in about a minute with "Could not download the
MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently
eating a test's timeout. The same URL is updated in
`examples/llm_eval/README.md` so a manual run does not hit the dead host
either.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:38:25 +05:30

8.3 KiB

Table of Contents

Overview

This repository provides scripts, popular third-party benchmarks, and instructions for evaluating the accuracy of Large Language Models (LLMs). It demonstrates how to use a ModelOpt quantized LLM with various established benchmarks, including deployment options using DirectML and TensorRT-LLM in a Windows environment.

Prerequisites

Category Details
Operating System Windows 10 or later
Python - For ORT-DML GenAI, use Python 3.11.
- For TensorRT-LLM, use Python 3.10.
- All other backends are compatible with both Python 3.10 and 3.11.
Package Manager pip
Compatible Hardware and Drivers - Ensure necessary hardware (e.g., CUDA-compatible GPU) and drivers are installed, depending on the evaluation method:
- DirectML for DirectML-based evaluation
- CUDA for TensorRT
Additional Tools - cmd: Recommended for running the provided commands.
- Tar Utility: Included in Windows 10 and later via PowerShell.
- Curl: Included in Windows 10 and later via PowerShell.

Accuracy Benchmarks

MMLU (Massive Multitask Language Understanding)

The MMLU benchmark assesses LLM performance across a wide range of tasks, producing a score between 0 and 1, where a higher score indicates better accuracy. Please refer the MMLU Paper for more details on this.

Setup

The table below lists the setup steps to prepare your environment for evaluating LLMs using the MMLU benchmark.

Step Command or Description
Open PowerShell as Administrator -
Create and Activate a Virtual Environment
(Optional but Recommended)
python -m venv llm_env
.\llm_env\Scripts\Activate.ps1
Install PyTorch and Related Packages pip install torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0 --index-url https://download.pytorch.org/whl/cu128
Install ONNX Runtime Packages pip install onnxruntime-directml==1.21.1
pip install onnxruntime-genai-directml==0.6.0
Install Benchmark Requirements pip install -r requirements.txt
Download MMLU Data mkdir data
curl.exe -L -o .\data\mmlu.tar https://huggingface.co/datasets/cais/mmlu/resolve/c30699e8356da336a370243923dbaf21066bb9fe/data.tar
tar -xf .\data\mmlu.tar -C .\data
Move-Item .\data\data .\data\mmlu

Evaluation Methods

Once the MMLU benchmark is set up, you can use the mmlu_benchmark.py script to evaluate LLMs deployed with various backends. Please refer examples below.

MMLU Benchmark with GenAI APIs for ORT-DML Deployment

To run the model with ORT-DML using GenAI, use the --ep genai_dml argument.

  • Test Suite

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep genai_dml `
        --output_file <output_log_file.json> `
        --ntrain 5
    
  • Specific Subjects

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep genai_dml `
        --output_file <output_log_file.json> `
        --subject abstract_algebra,anatomy,college_mathematics `
        --ntrain 5
    
MMLU Benchmark with ONNX Runtime APIs for DML, CUDA, or CPU Deployment

To run the model with ORT-DML, ORT-CUDA or ORT-CPU execution providers, use --ep ort_dml, --ep ort_cuda, or --ep ort_cpu respectively.

  • Test Suite

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep ort_dml `
        --output_file <output_log_file.json> `
        --ntrain 5
    
  • Specific Subjects

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep ort_dml `
        --output_file <output_log_file.json> `
        --subject abstract_algebra,anatomy,college_mathematics `
        --ntrain 5
    
MMLU Benchmark with Transformer APIs for PyTorch Hugging Face Models

To evaluate the PyTorch Hugging Face (HF) model, use the --ep pt argument.

  • Test Suite

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep pt `
        --output_file <output_log_file.json> `
        --ntrain 5 `
        --dtype <torch_dtype in model's config.json {float16|bfloat16}>
    
  • Specific Subjects

    python mmlu_benchmark.py `
        --model_name causal `
        --model_path <ONNX_model_folder> `
        --ep pt `
        --output_file <output_log_file.json> `
        --subject abstract_algebra,anatomy,college_mathematics `
        --ntrain 5 `
        --dtype <torch_dtype in model's config.json {float16|bfloat16}>
    
MMLU Benchmark with TensorRT-LLM APIs for TensorRT-LLM Deployment
  1. Install TensorRT-LLM and Compatible PyTorch

    pip install torch==2.4.0+cu121 --index-url https://download.pytorch.org/whl
    pip install tensorrt_llm==0.12.0 `
        --extra-index-url https://pypi.nvidia.com `
        --extra-index-url https://download.pytorch.org/whl/cu121/torch/
    
  2. Run the Benchmark

    • Test Suite

      python mmlu_benchmark.py `
          --model_name causal `
          --hf_model_dir <hf_model_path> `
          --engine_dir <engine_path> `
          --ep trt-llm `
          --ntrain 5 `
          --output_file result.json
      
    • Specific Subjects

      python mmlu_benchmark.py `
          --model_name causal `
          --hf_model_dir <hf_model_path> `
          --engine_dir <engine_path> `
          --ep trt-llm `
          --ntrain 5 `
          --output_file result.json `
          --subject abstract_algebra,anatomy,college_mathematics
      

Additional Metrics

Metric Directory Description
KL Divergence kl_divergence_metrics/ Measures output similarity between two models using KL divergence
Perplexity perplexity_metrics/ Evaluates language model quality using WikiText-2 perplexity
FVD fvd_metrics/ Computes Fréchet Video Distance between two sets of videos using I3D features

Each sub-directory contains its own README.md with detailed setup and usage instructions.

API changes in ONNX Runtime GenAI v0.6

In onnxruntime-genai (GenAI) v0.6, generator.compute_logits() and generator_params.input_ids are deprecated and new API generator.append_tokens(List: token_ids) is added (see GenAI PR-867 for details).

So, this MMLU script has been updated accordingly - refer following change-snippet from this MMLU script (left works with GenAI < 0.6, right works with GenAI 0.6+). Make sure to update the MMLU script accordingly (left part) for trying it with GenAI < 0.6.

alt text

Troubleshoot

  1. In case of any model specific issue (e.g. in tokenizer or in onnxruntime-genai package etc.), one can try using older GenAI e.g. export the ONNX model with onnxruntime-genai-directml 0.4 and transformers 4.44.

  2. In case of trying out MMLU run of ONNX model through GenAI, make sure that the input model is running fine with GenAI. Onnxruntime-genai has example inference scripts (e.g. see phi3 example script).