Files
Model-Optimizer/examples
ZhiyuandClaude Sonnet 5 0a8e70804a Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failure (#2480)
### What does this PR do?

Type of change: Bug fix

`post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional
post-quantization sanity-check `full_model.generate()` unguarded,
directly before `export_quantized()`. Any exception raised there aborted
the whole run and discarded a completed calibration without exporting a
checkpoint.

Root cause (traced from [NVBug
6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 /
DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with
`device_map="auto"`, relying on `accelerate`'s
`infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU
placement. On DGX Spark's unified-memory single-GPU host, that memory
probe under-reports GPU capacity, so part of even an 8B model can land
on CPU — and the existing fallback shrinks the GPU budget further (`*
gpu_mem_percentage`), compounding it. Calibration survives this because
it never invokes the real fake-quant kernel, but the post-PTQ sanity
`generate()` does, and NVFP4's dynamic-block-quantize op
(`modelopt/torch/quantization/tensor_quant.py`) hard-asserts
`amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes
there — after ~5.8 hours of calibration, before export.

This PR does not attempt to fix the underlying
`device_map`/memory-probing behavior (unverified without the actual
hardware/logs, which weren't reachable from this environment). Instead
it makes the failure mode safe: a failure in the optional sanity check
now only skips that check and warns, and export always proceeds,
regardless of why `generate()` failed.

### Usage

No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now
completes export even if the post-quantization sanity `generate()` call
raises.

### Testing

- Added
`tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`,
which drives `post_quantize()` with a `full_model.generate()` that
raises and asserts `export_quantized()` still runs.
- Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed).
- Ran `pre-commit` on the changed files (`ruff-format` reformatted line
wrapping only).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude
review` -->

### Additional Information

Fixes NVBug 6752977. Linked JIRA: OMNIML-5932.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantized checkpoint export now continues when the optional
post-quantization generation check fails.
  - A warning is shown when the generation check cannot complete.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 01:21:07 +00:00
..
2026-06-27 01:00:25 +05:30