Files
Model-Optimizer/tools
Keval MorabiaandClaude Opus 4.8 55a2101e2f Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do?

Type of change: documentation + minor example-script tweaks

Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR
was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16
tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md)
results using the **new shared calibration loop (sequence packing)** and
to **fix tool calling in evaluation**.

- Refreshed the prune → distill → eval → **FP8** results (accuracy +
vLLM throughput tables) with the new calibration loop.
- **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now
run the Python sandbox tool. The tutorial reports both **with-tools**
and **no-tools** GPQA/AIME and shows `mean ± std_dev`.
- Script tweaks: `quantize.py` calibration now uses sequence packing
(`pack=True`) which leads to slight improvement in PTQ;
`prune_minitron.py` defaults `inference_batch_size` to
`calib_batch_size`.

### Testing

Documentation + small example-script changes; tutorial relative links
resolve and the results tables / figure were verified consistent.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the main guide and evaluator instructions for prune + distill
+ FP8/NVFP4 quantization, including refreshed vLLM deployment tips,
benchmark/noise presentation, and long-context tool-calling attribution
notes.
* Refreshed README technique examples/links, reordered the model support
matrix rows, and improved pruning overview/support-matrix text.

* **Changes to Examples**
* NAS pruning now documents higher GPU memory usage vs manual pruning;
pruning batching defaults were improved.
* Quantization PTQ calibration uses packed document packing; quantized
checkpoint export messaging was streamlined.
* Updated pruning/distillation/quantization tutorial guidance,
metrics/tables, command parameters, and evaluator YAML settings
(KV-cache dtype, generation defaults, task behavior).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 18:32:13 +00:00
..