Files
Keval MorabiaandClaude Opus 5 6a2ae5a25b Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been
failing with `Failed: Timeout (>900.0s) from pytest-timeout` on
unrelated branches (runs 33027056246 and the one for `ad83a428`, while
the 2026-08-24 nightly passed). It is not the test being slow — it is
the harness deadlocking, and the deadlock also destroys the diagnostics
that would explain the underlying kill.

**Mechanism.** The traceback shows `self = <Popen: returncode: -9 args:
['scripts/huggingface_example.sh', ...]>` while still blocked in
`stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends
SIGKILL — pytest-timeout raises in the main thread, and the test's
`finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM
kill is the likely source). But `subprocess.run(..., stdout=PIPE,
stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and
a surviving grandchild (the TRT-LLM serve/build worker) still holds the
write end. EOF never arrives, so the test blocks until the 900 s alarm.
Because the pipe is never drained, **every line of child output is
discarded**, which is why the CI log says nothing about what the script
was doing when it died.

Reduced to a self-contained reproducer:

```python
script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$"
subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20)
# -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost
```

**Fix.** `_run_capturing` now starts the command in its own session,
drains its output on a reader thread (so logs stream as they arrive
instead of being buffered until the end), waits on the *process*, and
kills the process group if descendants still hold the pipe after a 30 s
grace period. A killed launcher now fails in seconds with its logs
intact instead of silently burning the test's whole timeout.

This does not fix whatever kills the script; it makes it diagnosable.
Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against
its 900 s mark** on the last green nightly, so it is fragile regardless
and may want its work trimmed or its budget raised once the logs show
where the time goes.

### Usage

```python
# unchanged public API
run_example_command(cmd_parts, example_path="llm_eval")
```

### Testing

Verified against the reproducer above and on the normal paths:

| scenario | before | after |
| --- | --- | --- |
| launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900
s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured |
| the surviving descendant | keeps running | killed with the process
group (stopped ticking, 20 -> 20 bytes) |
| normal exit | ok | `rc=0`, stdout and stderr interleaved in order |
| non-zero exit | ok | `rc=3`, output captured |

The example-test suites that use this helper run through the same code
path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised
it on nemo:26.08 in the branch this was extracted from.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `_run_capturing` keeps its
`(returncode, output)` contract; only the buffering strategy changed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — this is test
infrastructure; the scenario needs a process that outlives a SIGKILLed
parent, which is awkward to assert in CI. Verified manually with the
reproducer above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test-infrastructure fix.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved command execution reliability with real-time output capture.
* Ensured lingering child processes are cleaned up after commands exit
or are interrupted.
  * Added warnings when forced cleanup may truncate output.
  * Prevented hangs when descendant processes keep output streams open.

* **Documentation**
* Updated MMLU setup instructions to use the Hugging Face dataset
repository.
  * Improved Windows instructions by explicitly using `curl.exe`.

* **Examples**
* Improved MMLU downloads with retries, separate timeouts, resume
support, and automatic temporary-file cleanup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

### Update: the second half of the failure

With the streaming fix in place, the next CI run showed the *actual*
cause, which the old code had been hiding. That run blocked in
`process.wait()` with `<Popen: returncode: None ...>` — child alive and
working, not the previous dead-child pipe deadlock — and the now-visible
output was:

```
--2026-08-27 18:09:20--  (try: 4)  https://people.eecs.berkeley.edu/~hendrycks/data.tar
Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out.
Retrying.
--2026-08-27 18:11:40--  (try: 5)  ...
```

`huggingface_example.sh` downloads the MMLU tarball from
`people.eecs.berkeley.edu`, that host stopped answering around
2026-08-25, and wget's default retry policy (20 tries, ~2 min per
connect timeout) consumed the whole 900 s budget. Not runner-specific:
the URL also times out from a developer workstation, and the nightlies
flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage.

So this PR now carries both halves of the same failure:

1. the harness no longer deadlocks and no longer swallows the logs
(`985809cc2d`), and
2. the MMLU data comes from HuggingFace's copy of the same tarball, with
bounded retries (`40f1d89154`).

The mirror is byte-for-byte the same dataset in the same layout the
script already expects — verified by running the exact download/extract
commands:

```
https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar  ->  HTTP 200, 166 MB
data/mmlu/{dev,test,val}/  ->  57 subject CSVs each, plus auxiliary_train/
```

`wget --timeout=20 --tries=3` plus an explicit error means the next
dataset-host outage fails in about a minute with "Could not download the
MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently
eating a test's timeout. The same URL is updated in
`examples/llm_eval/README.md` so a manual run does not hit the dead host
either.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:38:25 +05:30
..
2026-04-08 19:19:13 +00:00

NVIDIA Model Optimizer - Windows

A Library to Quantize and Compress Deep Learning Models for Optimized Inference on Native Windows RTX GPUs

Documentation version license

Examples | Benchmark Results

Latest News

Table of Contents

Overview

The Model Optimizer - Windows (ModelOpt-Windows) is engineered to deliver advanced model compression techniques, including quantization, to Windows RTX PC systems. Specifically tailored to meet the needs of Windows users, ModelOpt-Windows is optimized for rapid and efficient quantization, featuring local GPU calibration, reduced system and video memory consumption, and swift processing times. The primary objective of the ModelOpt-Windows is to generate optimized, standards-compliant ONNX-format models. This makes it an ideal solution for seamless integration with ONNX Runtime (ORT) and DirectML (DML) frameworks, ensuring broad compatibility with any inference framework supporting the ONNX standard. Furthermore, ModelOpt-Windows integrates smoothly within the Windows ecosystem, with full support for tools and SDKs such as Olive and ONNX Runtime, enabling deployment of quantized models across various independent hardware vendors (IHVs) through the DML path and TensorRT path.

Model Optimizer is available for free for all developers on NVIDIA PyPI. This repository is for sharing examples and GPU-optimized recipes as well as collecting feedback from the community.

Installation

ModelOpt-Windows can be installed either as a standalone toolkit or through Microsoft's Olive.

Standalone Toolkit Installation (with CUDA 12.x)

To install ModelOpt-Windows as a standalone toolkit on CUDA 12.x systems, run the following commands:

pip install nvidia-modelopt[onnx]

Installation with Olive

The ModelOpt-Windows is integrated into Microsoft's Olive framework. Run the following commands to install ModelOpt through Olive.

pip install olive-ai[nvmo]
pip install onnxruntime-genai-cuda

For more details, or to use different ONNX Runtime Execution Providers, refer to the detailed installation instructions.

Techniques

Quantization

Quantization is an effective model optimization technique for large models. Quantization with ModelOpt-Windows can compress model size by 2x-4x, speeding up inference while preserving model quality. ModelOpt-Window enables highly performant quantization formats including INT4, FP8, INT8, etc. and supports advanced algorithms such as AWQ and SmoothQuant* focusing on post-training quantization (PTQ) for ONNX and PyTorch* models with DirectML, CUDA and TensorRT* inference backends.

For more details, please refer to the detailed quantization guide.

Getting Started

The ONNX quantization API requires a model, calibration data, along with quantization settings like algorithm, calibration-EPs etc. Here’s an example snippet to apply INT4 AWQ quantization:

from modelopt.onnx.quantization.int4 import quantize as quantize_int4
# import other packages as needed
calib_inputs = get_calib_inputs(dataset, model_name, cache_dir, calib_size, batch_size,...)
quantized_onnx_model = quantize_int4(
    onnx_path,
    calibration_method="awq_lite",
    calibration_data_reader=None if use_random_calib else calib_inputs,
    calibration_eps=["dml", "cpu"]
)
onnx.save_model(
    quantized_onnx_model,
    output_path,
    save_as_external_data=True,
    location=os.path.basename(output_path) + "_data",
    size_threshold=0,
)

Check modelopt.onnx.quantization.quantize_int4 for details about INT4 quantization API.

Refer to our Support Matrix for details about supported features and models.

To learn more about ONNX PTQ, refer to our docs.

Deployment

The quantized onnx model can be deployed using frameworks like onnxruntime. Ensure that model’s opset is 19+ for FP8 quantization, and it is 21+ for INT4 quantization. This is needed due to different opset requirements of ONNX’s Q/DQ nodes for INT4, FP8 data-types support. Refer to Apply Post Training Quantization (PTQ) for details.

# write steps (say, upgrade_opset() method) to upgrade or patch opset of the model, if needed
# the opset-upgrade, if needed, can be done on either base ONNX model or on the quantized model
# finally, save the quantized model

quantized_onnx_model = upgrade_opset(quantized_onnx_model)
onnx.save_model(
    quantized_onnx_model,
    output_path,
    save_as_external_data=True,
    location=os.path.basename(output_path) + "_data",
    size_threshold=0,
)

For detailed instructions about deployment of quantized models with ONNX Runtime, see the ONNX Runtime Deployment Guide.

Note

The ready-to-deploy optimized ONNX models from ModelOpt-Windows are available at HuggingFace NVIDIA collections.

Examples

  • Examples for Post-Training Quantization (PTQ) of ONNX models:
    • PTQ for GenAI LLMs covers how to use ONNX PTQ with ONNX Runtime GenAI built LLM ONNX models, and their deployment with DirectML.
    • PTQ for Whisper illustrates using ONNX PTQ with a Whisper ONNX model (i.e. an ASR model). It also provides example script for Optimum-ORT based inference of Whisper using CUDA EP.
    • PTQ for SAM2 illustrates using ONNX PTQ with a SAM2 ONNX model (i.e. a segmentation model).
  • Examples that demonstrate PTQ of a PyTorch model followed by ONNX export:
    • Diffusers example demonstrates how to apply PTQ to diffusion models in PyTorch format and then export the quantized models to ONNX.
  • MMLU Benchmark provides an example script for MMLU benchmarking of LLM models, and demonstrates how to run it with various popular backends like DirectML, TensorRT-LLM* and model formats like ONNX and PyTorch*.

Support Matrix

Model Type Support Matrix
Large Language Models (LLMs) View Support Matrix
Automatic Speech Recognition View Support Matrix
Segmentation Models View Support Matrix
Diffusion Models View Support Matrix

Benchmark Results

Please refer to benchmark results for performance and accuracy comparisons of popular Large Language Models (LLMs).

Collection Of Optimized ONNX Models

The ready-to-deploy optimized ONNX models from ModelOpt-Windows are available at HuggingFace NVIDIA collections. These models can be deployed using DirectML backend. Follow the instructions provided along with the published models for deployment.

Release Notes

Please refer to changelog

* Experimental support