### What does this PR do? Type of change: Bug fix `tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been failing with `Failed: Timeout (>900.0s) from pytest-timeout` on unrelated branches (runs 33027056246 and the one for `ad83a428`, while the 2026-08-24 nightly passed). It is not the test being slow — it is the harness deadlocking, and the deadlock also destroys the diagnostics that would explain the underlying kill. **Mechanism.** The traceback shows `self = <Popen: returncode: -9 args: ['scripts/huggingface_example.sh', ...]>` while still blocked in `stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends SIGKILL — pytest-timeout raises in the main thread, and the test's `finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM kill is the likely source). But `subprocess.run(..., stdout=PIPE, stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and a surviving grandchild (the TRT-LLM serve/build worker) still holds the write end. EOF never arrives, so the test blocks until the 900 s alarm. Because the pipe is never drained, **every line of child output is discarded**, which is why the CI log says nothing about what the script was doing when it died. Reduced to a self-contained reproducer: ```python script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$" subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20) # -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost ``` **Fix.** `_run_capturing` now starts the command in its own session, drains its output on a reader thread (so logs stream as they arrive instead of being buffered until the end), waits on the *process*, and kills the process group if descendants still hold the pipe after a 30 s grace period. A killed launcher now fails in seconds with its logs intact instead of silently burning the test's whole timeout. This does not fix whatever kills the script; it makes it diagnosable. Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against its 900 s mark** on the last green nightly, so it is fragile regardless and may want its work trimmed or its budget raised once the logs show where the time goes. ### Usage ```python # unchanged public API run_example_command(cmd_parts, example_path="llm_eval") ``` ### Testing Verified against the reproducer above and on the normal paths: | scenario | before | after | | --- | --- | --- | | launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900 s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured | | the surviving descendant | keeps running | killed with the process group (stopped ticking, 20 -> 20 bytes) | | normal exit | ok | `rc=0`, stdout and stderr interleaved in order | | non-zero exit | ok | `rc=3`, output captured | The example-test suites that use this helper run through the same code path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised it on nemo:26.08 in the branch this was extracted from. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `_run_capturing` keeps its `(returncode, output)` contract; only the buffering strategy changed. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ — this is test infrastructure; the scenario needs a process that outlives a SIGKILLed parent, which is awkward to assert in CI. Verified manually with the reproducer above. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test-infrastructure fix. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved command execution reliability with real-time output capture. * Ensured lingering child processes are cleaned up after commands exit or are interrupted. * Added warnings when forced cleanup may truncate output. * Prevented hangs when descendant processes keep output streams open. * **Documentation** * Updated MMLU setup instructions to use the Hugging Face dataset repository. * Improved Windows instructions by explicitly using `curl.exe`. * **Examples** * Improved MMLU downloads with retries, separate timeouts, resume support, and automatic temporary-file cleanup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ### Update: the second half of the failure With the streaming fix in place, the next CI run showed the *actual* cause, which the old code had been hiding. That run blocked in `process.wait()` with `<Popen: returncode: None ...>` — child alive and working, not the previous dead-child pipe deadlock — and the now-visible output was: ``` --2026-08-27 18:09:20-- (try: 4) https://people.eecs.berkeley.edu/~hendrycks/data.tar Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out. Retrying. --2026-08-27 18:11:40-- (try: 5) ... ``` `huggingface_example.sh` downloads the MMLU tarball from `people.eecs.berkeley.edu`, that host stopped answering around 2026-08-25, and wget's default retry policy (20 tries, ~2 min per connect timeout) consumed the whole 900 s budget. Not runner-specific: the URL also times out from a developer workstation, and the nightlies flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage. So this PR now carries both halves of the same failure: 1. the harness no longer deadlocks and no longer swallows the logs (`985809cc2d`), and 2. the MMLU data comes from HuggingFace's copy of the same tarball, with bounded retries (`40f1d89154`). The mirror is byte-for-byte the same dataset in the same layout the script already expects — verified by running the exact download/extract commands: ``` https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar -> HTTP 200, 166 MB data/mmlu/{dev,test,val}/ -> 57 subject CSVs each, plus auxiliary_train/ ``` `wget --timeout=20 --tries=3` plus an explicit error means the next dataset-host outage fails in about a minute with "Could not download the MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently eating a test's timeout. The same URL is updated in `examples/llm_eval/README.md` so a manual run does not hit the dead host either. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
NVIDIA Model Optimizer - Windows
A Library to Quantize and Compress Deep Learning Models for Optimized Inference on Native Windows RTX GPUs
Latest News
- [2024/11/19] Microsoft and NVIDIA Supercharge AI Development on RTX AI PCs
- [2024/11/18] Quantized INT4 ONNX models available on Hugging Face for download
Table of Contents
- Overview
- Installation
- Techniques
- Examples
- Support Matrix
- Benchmark Results
- Collection of Optimized ONNX Models
- Release Notes
Overview
The Model Optimizer - Windows (ModelOpt-Windows) is engineered to deliver advanced model compression techniques, including quantization, to Windows RTX PC systems. Specifically tailored to meet the needs of Windows users, ModelOpt-Windows is optimized for rapid and efficient quantization, featuring local GPU calibration, reduced system and video memory consumption, and swift processing times. The primary objective of the ModelOpt-Windows is to generate optimized, standards-compliant ONNX-format models. This makes it an ideal solution for seamless integration with ONNX Runtime (ORT) and DirectML (DML) frameworks, ensuring broad compatibility with any inference framework supporting the ONNX standard. Furthermore, ModelOpt-Windows integrates smoothly within the Windows ecosystem, with full support for tools and SDKs such as Olive and ONNX Runtime, enabling deployment of quantized models across various independent hardware vendors (IHVs) through the DML path and TensorRT path.
Model Optimizer is available for free for all developers on NVIDIA PyPI. This repository is for sharing examples and GPU-optimized recipes as well as collecting feedback from the community.
Installation
ModelOpt-Windows can be installed either as a standalone toolkit or through Microsoft's Olive.
Standalone Toolkit Installation (with CUDA 12.x)
To install ModelOpt-Windows as a standalone toolkit on CUDA 12.x systems, run the following commands:
pip install nvidia-modelopt[onnx]
Installation with Olive
The ModelOpt-Windows is integrated into Microsoft's Olive framework. Run the following commands to install ModelOpt through Olive.
pip install olive-ai[nvmo]
pip install onnxruntime-genai-cuda
For more details, or to use different ONNX Runtime Execution Providers, refer to the detailed installation instructions.
Techniques
Quantization
Quantization is an effective model optimization technique for large models. Quantization with ModelOpt-Windows can compress model size by 2x-4x, speeding up inference while preserving model quality. ModelOpt-Window enables highly performant quantization formats including INT4, FP8, INT8, etc. and supports advanced algorithms such as AWQ and SmoothQuant* focusing on post-training quantization (PTQ) for ONNX and PyTorch* models with DirectML, CUDA and TensorRT* inference backends.
For more details, please refer to the detailed quantization guide.
Getting Started
The ONNX quantization API requires a model, calibration data, along with quantization settings like algorithm, calibration-EPs etc. Here’s an example snippet to apply INT4 AWQ quantization:
from modelopt.onnx.quantization.int4 import quantize as quantize_int4
# import other packages as needed
calib_inputs = get_calib_inputs(dataset, model_name, cache_dir, calib_size, batch_size,...)
quantized_onnx_model = quantize_int4(
onnx_path,
calibration_method="awq_lite",
calibration_data_reader=None if use_random_calib else calib_inputs,
calibration_eps=["dml", "cpu"]
)
onnx.save_model(
quantized_onnx_model,
output_path,
save_as_external_data=True,
location=os.path.basename(output_path) + "_data",
size_threshold=0,
)
Check modelopt.onnx.quantization.quantize_int4 for details about INT4 quantization API.
Refer to our Support Matrix for details about supported features and models.
To learn more about ONNX PTQ, refer to our docs.
Deployment
The quantized onnx model can be deployed using frameworks like onnxruntime. Ensure that model’s opset is 19+ for FP8 quantization, and it is 21+ for INT4 quantization. This is needed due to different opset requirements of ONNX’s Q/DQ nodes for INT4, FP8 data-types support. Refer to Apply Post Training Quantization (PTQ) for details.
# write steps (say, upgrade_opset() method) to upgrade or patch opset of the model, if needed
# the opset-upgrade, if needed, can be done on either base ONNX model or on the quantized model
# finally, save the quantized model
quantized_onnx_model = upgrade_opset(quantized_onnx_model)
onnx.save_model(
quantized_onnx_model,
output_path,
save_as_external_data=True,
location=os.path.basename(output_path) + "_data",
size_threshold=0,
)
For detailed instructions about deployment of quantized models with ONNX Runtime, see the ONNX Runtime Deployment Guide.
Note
The ready-to-deploy optimized ONNX models from ModelOpt-Windows are available at HuggingFace NVIDIA collections.
Examples
- Examples for Post-Training Quantization (PTQ) of ONNX models:
- PTQ for GenAI LLMs covers how to use ONNX PTQ with ONNX Runtime GenAI built LLM ONNX models, and their deployment with DirectML.
- PTQ for Whisper illustrates using ONNX PTQ with a Whisper ONNX model (i.e. an ASR model). It also provides example script for Optimum-ORT based inference of Whisper using CUDA EP.
- PTQ for SAM2 illustrates using ONNX PTQ with a SAM2 ONNX model (i.e. a segmentation model).
- Examples that demonstrate PTQ of a PyTorch model followed by ONNX export:
- Diffusers example demonstrates how to apply PTQ to diffusion models in PyTorch format and then export the quantized models to ONNX.
- MMLU Benchmark provides an example script for MMLU benchmarking of LLM models, and demonstrates how to run it with various popular backends like DirectML, TensorRT-LLM* and model formats like ONNX and PyTorch*.
Support Matrix
| Model Type | Support Matrix |
|---|---|
| Large Language Models (LLMs) | View Support Matrix |
| Automatic Speech Recognition | View Support Matrix |
| Segmentation Models | View Support Matrix |
| Diffusion Models | View Support Matrix |
Benchmark Results
Please refer to benchmark results for performance and accuracy comparisons of popular Large Language Models (LLMs).
Collection Of Optimized ONNX Models
The ready-to-deploy optimized ONNX models from ModelOpt-Windows are available at HuggingFace NVIDIA collections. These models can be deployed using DirectML backend. Follow the instructions provided along with the published models for deployment.
Release Notes
Please refer to changelog
* Experimental support