Commit Graph
53 Commits
Author SHA1 Message Date
Chenjie Luo 1d21ab9e29 [DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary

- DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k
routing during MoE calibration. The previous all-tokens-to-all-experts
path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag.
- After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's
`input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via
`dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays
per-expert; any uncalibrated expert is filled by computing amax over the
dequantized FP8 weight.
- `mtq.print_quant_summary` is now also written to
`<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`.

## Why

Forcing all tokens through every expert doubled calibration time and
inflated `input_quantizer.amax` for cold-routing experts with outliers
they never see at inference. The new flow matches the inference
distribution, runs roughly 2x faster, and mirrors the
`layer_sync_moe_local_experts_amax` semantics that mtq runs
automatically for `QuantSequentialMLP`-derived MoEs.

## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG)

Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default):
- All 44,544 expert weight amaxes bit-identical.
- Attention, shared experts, gate: identical.
- Expert `w1.input` and `w3.input` (shared MoE block input): identical.
- Expert `w2.input` (post-SiLU gated, expert-specific): synced to
layer-wide peer max — 99.3% are larger than baseline (median 11.4x)
since peer-max captures the worst-case outlier from any expert in the
layer; 0.7% are smaller. This is the same trade-off
`set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`.

## Test plan

- [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min
(vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the
comparison above.
- [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces
`_amax_baseline/` identical (other than rounding) to the prior
`CalibMoe`-default behavior.
- [x] `.quant_summary.txt` written under `output_path` on rank 0.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--calib_all_experts` option to enable an alternate PTQ
calibration mode; default remains top-k routing with a post-calibration
per-layer peer-max synchronization and a compute fallback for
uncalibrated experts.
* **Documentation**
* Clarified default and alternate calibration behaviors and added note
about generation of a `.quant_summary.txt` summary file.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-05-04 11:08:01 -07:00
Chenjie LuoandClaude Opus 4.7 50706d1750 Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary

- New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and
`huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g.
`openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight
reconstruction for the in-range blocks.
- The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and
`_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`.
The resulting per-block scale `2^(k_j − m)` is exactly representable in
E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble
verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block
amax falls back to data-derived `max(|w_block|)`, which keeps the
post-E4M3-clamp scale close to the block's actual magnitude.

## Verification

End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only
--cast_mxfp4_to_nvfp4`:

```
[cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%)
[cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%)
```

End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200,
`--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size
4`):

```
[cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%)
[cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%)
```

Five layers fall into the OOR regime (block-spread > 17); the remaining
1,214 OOR blocks use the data-derived per-block amax fallback.
Block-level losslessness is **99.99996%** end-to-end.

Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant
(~19B elements):

| Metric | Without cast | With cast |
|---|---|---|
| Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** |
| Total RMSE | 8.67e−02 | **0** |
| max\|err\| | up to 8.0e+1 | **0** |

## Modelopt-side enablers

- `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to
`NVFP4StaticQuantizer` at the end of calibration.
- `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was
2D-only), unblocking MoE expert weights of shape `(E, F, K)`.
- BMM-experts NVFP4 export routes through
`get_weights_scaling_factor_from_quantizer` for static-mode quantizers,
so the pinned `_amax` is actually consumed.
- `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before
stacking, supporting per-block (vs scalar) static-mode amax.

## Test plan

- [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py`
(15 tests, all passing) cover: scalar/global-amax math, per-block hybrid
(in-range closed-form vs OOR data-derived), shape preservation, key
collection, and end-to-end `build_amax_map` against a synthetic
safetensors checkpoint.
- [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only`
qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100%
lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks).
- [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200,
`nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5
--calib_batch_size 4`). 67/72 layers fully lossless; 99.99996%
block-level losslessness (3,583,179,586 / 3,583,180,800).
- [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both
exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`:
- **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample:
*"Quantum computing is poised to revolutionize data analysis. However,
its potential is currently limited by quantum hardware constraints,
including error rates, qubit lifetimes, and lack of fault tolerance…"*
- **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum
computing is poised to revolutionize data storage and processing. These
rare earth-based systems could serve as robust qubits; resistant to
environmental decoherence…"*
- [x] MSE comparison script (run separately during development) confirms
per-tensor SNR=∞ across all 48 MoE expert tensors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to
enable it; helper scripts updated to expose the option.

* **Bug Fixes**
  * Fixed static NVFP4 export for expert weights.
* Improved collection/handling of quantizer amax values to avoid shape
issues.
  * Generalized FP4 kernel to accept flexible tensor dimensionality.
  * Ensured static-block NVFP4 promotion during calibration.

* **Tests**
* Added comprehensive tests for the conversion workflow, helpers, and
end-to-end application.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 19:04:04 +00:00
Chenjie LuoandKeval Morabia 168cd828c1 Add qwen3 moe experts only test (#1274)
## Summary
- Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG`
quantization config
- Verifies that `hf_quant_config.json` correctly reports `quant_algo:
NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in
`exclude_modules` while routed expert layers (`mlp.experts.*`) do not
- Reference:
https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json

Type of change: New tests

### Known issue
On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not
recognized by `get_quant_config`, causing `quant_algo=None` in the
exported config. This test currently **fails** on transformers 5.x and
is intended to be fixed by a follow-up change.

## Testing
- **transformers 4.57.6**: PASSED
- **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused
expert export gap)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models
with NVFP4 quantization.
* Verifies the exported checkpoint records the NVFP4 quantization
algorithm and that module exclusion patterns correctly exclude attention
and LM head components while not excluding routed expert paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-01 17:56:04 +00:00
Chenjie Luo 3ad4f4f093 [Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary
- `get_max_batch_size` halved `target_data_batch` on
`torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each
retry re-fed the same too-large tensor — the retry loop was effectively
a no-op.
- Refactor the expand logic into an `_expand_to(batch)` helper, rebuild
`target_input` after halving, and call `torch.cuda.empty_cache()`
between attempts.

## Test plan
- [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input`
mocks `torch.cuda.*` and asserts the second retry receives the halved
tensor (shapes seen: `[1, 10, 5]`, regulated result `4`).
- [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass
(skipping the network-only minipile test).
- [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on
commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced GPU memory management during batch size detection. When
out-of-memory errors occur during the initial probing phase, the system
now properly adapts input tensors to smaller batch sizes and clears GPU
cache before retry attempts, resulting in more reliable recovery and
stable batch sizing across diverse hardware environments.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-30 18:25:28 +00:00
Chenjie Luo 47a33db9b6 [NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary

- NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) —
`Qwen3-30B-A3B nvfp4_awq` quantization fails at export with
`AssertionError: Modules have different quantization formats`.
- Root cause: in `model_calib.awq_lite`, MoE experts that end up
disabled (NaN in act/weight scales, or no search-pass tokens) get
`max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format`
then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`.
`unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128
experts of each linear name (gate_proj/down_proj/up_proj) and calls
`preprocess_linear_fusion(..., resmooth_only=True)`, which asserts
uniform format → fires for any single mismatched expert.
- Fix: unify the disabled-expert paths in the awq_lite postprocess loop
so any expert with `is_enabled == False` (no cache hits, NaN scales, or
no search-pass tokens) receives `max_calibrate` + a neutral all-ones
`pre_quant_scale`, matching the existing behavior for `num_cache_steps
== 0`. Emit a warning so users notice that calibration coverage is
incomplete and accuracy may degrade.

## Test plan

- [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5
passed
- [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a
small calib set that leaves many experts uncalibrated:
- All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq`
(no mismatch)
  - `export_hf_checkpoint` succeeds with no `AssertionError`
- The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning
fires for each affected expert
- [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI
(cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to
fit budget) on B200:
- 36 "the second time did not forward data through
..experts.X.{gate,up,down}_proj" warnings — i.e. the exact
bug-triggering condition from the original NVBug log naturally
reproduces
- 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for
uncalibrated/disabled experts
  - 0 `AssertionError`s — export completes
- `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq`
and post-PTQ generation works

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-27 12:40:14 -07:00
Chenjie Luo fda0899e40 feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary

- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
  - `general/ptq/fp8_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.

## Motivation

Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).

## Test plan

- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.

* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.

* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 18:45:16 +00:00
Chenjie Luo 01788bb007 Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary

- Removes Mllama (Llama 3.2 Vision) model-type branches from the
`llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the
now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`.
- Drops the legacy `MllamaImageProcessor` path in
`modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF
ProcessorMixin path handles the remaining cases.
- Adds a CHANGELOG entry under 0.44 Backward Breaking Changes.

## Test plan

- [x] CI lint / unit tests pass
- [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh
--model <llm> --quant fp8`` (text-only path, non-mllama)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Removed Mllama (Llama 3.2 Vision) support from quantization examples.
This includes removal of dedicated image processor implementation,
specialized model handling, and related calibration logic.
* Updated VLM image-text calibration guidance to use
`--calib_with_images` flag with other supported VLMs instead of
Mllama-specific processing paths.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 23:24:42 +05:30
Chenjie Luo 04fcf24227 Fix LLM deploy test failure by defaulting expert parallelism to 1 (#1273)
### What does this PR do?

Type of change: Bug fix

Fixes TRT-LLM DeepEP kernel failures during LLM deployment on
unsupported GPUs (e.g. Blackwell SM 12.0) by defaulting expert
parallelism (`ep`) to 1 instead of auto-setting it to the GPU count for
MoE models.

Previously, when the model config contained expert-related keys, `ep`
was automatically set to `torch.cuda.device_count()`, which triggered
DeepEP kernel failures on GPUs that don't support it. Now `ep` defaults
to 1 while still enabling attention data parallelism for MoE models.
Expert parallelism can be enabled explicitly by the caller when the
environment is known to support it.

### Testing

- [x] Verified that the `llm_ptq` test passes with this fix on Blackwell
GPUs.
- [x] 2-gpu CI test triggered:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24495054531/job/71588037727

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-16 16:18:38 -07:00
Chenjie Luo d45219b390 Fix debugger server failing to detect editable-installed modelopt (#1270)
## Summary
- Removed `PYTHONPATH="" python -I` override in `check_modelopt_local()`
so the PYTHONPATH validation uses the actual environment instead of an
isolated one
- Moved the workdir log line earlier in `server.sh` for better debugging
visibility

## Test plan
- [x] Start `server.sh` inside a Docker container and verify it
correctly detects editable-installed modelopt
- [x] Confirm the workdir is logged before the modelopt check runs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Server startup now displays the configured work directory earlier in
the initialization process, providing improved visibility of the active
directory during server launch.
* Simplified the modelopt validation check during server initialization
while maintaining the same validation behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 22:53:46 -07:00
Chenjie Luo 7c8557158d Add job cancellation support to the debugger command relay (#1262)
## Summary

- Add a `cancel` subcommand to the client that terminates the currently
running command on the server
- Server now runs commands in the background with PID tracking, enabling
cancellation mid-execution
- Client-side timeouts automatically cancel the running command on the
server (previously the server process was left running)
- Hardened against race conditions through 4 rounds of adversarial
review (15 fixes total)

### Key changes

**server.sh:**
- Commands run in background with PID tracked in `$RELAY_DIR/running`
(atomic tmp+mv write)
- Cancel detection loop checks for `$RELAY_DIR/cancel` file with cmd_id
verification
- SIGTERM with 5s grace period, then SIGKILL escalation for stuck
processes
- `.exit` file written before `running` marker removed (ordering
guarantee)
- `set -e`-safe: `wait` uses `|| exit_code=$?` pattern; cleanup trap
fully guarded
- Stale cancel files cleared at command start; mismatched/empty signals
rejected
- Command file read into memory and removed before execution (eliminates
TOCTOU with client timeout)

**client.sh:**
- New `cancel` subcommand: writes target cmd_id to cancel file, waits
for server acknowledgment (30s timeout)
- `run` timeout now sends targeted cancel signal (verifies cmd_id match
to avoid killing wrong command)
- `run` timeout cleans up orphaned result files
- `status` shows currently running command
- `flush` rejects if a command is currently running (prevents state
corruption)
- Exit code validated as numeric before use

### Protocol additions

```
.relay/
├── running    # server writes cmd_id:pid while executing (atomic)
├── cancel     # client writes target cmd_id to request cancellation
```

## Test plan

- [ ] Start server in Docker, handshake from host
- [ ] Run a command (`client.sh run "sleep 30"`), cancel it (`client.sh
cancel`), verify exit code 130
- [ ] Run a command with short timeout (`--timeout 5 run "sleep 30"`),
verify auto-cancel
- [ ] Run a command that exits non-zero, verify server stays alive
- [ ] Run `status` during execution, verify it shows the running command
- [ ] Attempt `flush` during execution, verify it is rejected
- [ ] Cancel when nothing is running, verify clean message

🤖 Generated with [Claude Code](https://claude.com/claude-code)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (bash scripts for dev
tooling, tested manually)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal tooling)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive debug skill guide and protocol reference with
quick CLI examples and a new “Cancelling Commands” section.

* **New Features**
* Client-side `cancel` command to terminate the currently running remote
command.
  * Status now reports active command (or `(idle)`).

* **Improvements**
* Stronger startup/validation guidance, safer shutdown/cleanup,
deterministic cancel exit semantics (130), and auto-cancel on
client-side timeout.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 16:49:42 +00:00
Chenjie LuoandClaude Opus 4.6 952a62bf65 Fix missing attention_mask in calibration dataloader (#1261)
## Summary
- When `include_labels=False` (the default for PTQ calibration),
`get_dataset_dataloader` was discarding the `attention_mask` produced by
the tokenizer and only returning `input_ids`.
- Without `attention_mask`, HuggingFace models create a full causal
mask, causing padding tokens to participate in attention during
calibration and skewing quantization statistics.
- This fix includes `attention_mask` alongside `input_ids` so the model
correctly ignores padding tokens during calibration forward passes.

## Details
In `modelopt/torch/utils/dataset_utils.py`, the tokenizer call at line
387 with `padding=True` produces both `input_ids` and `attention_mask`.
The `include_labels=True` path (line 406) already preserves the full
`batch_encoded` dict including `attention_mask`. However, the
`include_labels=False` path was only keeping `input_ids` "for backward
compatibility."

During the calibration forward loop (`_forward_loop` →
`_process_batch`), the batch dict is unpacked as `**kwargs` into
`model.forward()`. Without `attention_mask`, HF models default to
attending to all positions including padding, which pollutes calibration
statistics.

**Practical impact**: With `batch_size=1` there is no padding so the bug
is invisible. With larger batch sizes and variable-length samples,
shorter sequences get padded and the effect grows.

## Test plan
- [x] Existing unit tests pass
(`tests/unit/torch/utils/test_dataset_utils.py`)
- [x] Pre-commit hooks pass
- [ ] Verify PTQ accuracy with batch_size > 1 on a padded calibration
dataset (GPU required)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 00:04:27 -07:00
Chenjie Luo f7557221e3 Add file-based command relay for remote Docker testing (#1174)
## Summary
- Adds a lightweight file-based client/server relay (`tools/debugger/`)
that enables Claude Code (or any host-side automation) to execute
commands inside a remote Docker container using only a shared filesystem
— no networking setup required.
- The server auto-detects the repo root, installs modelopt (`pip install
-e .[dev]`), sets `PYTHONPATH`, and listens for commands.
- The client supports `handshake`, `run`, `status`, and `flush`
subcommands.
- Includes `README.md` (full protocol docs) and `CLAUDE.md` (quick
reference for Claude Code).

## Tested
- Ran Qwen3.5-35B-A3B MoE PTQ with `nvfp4_experts_only` quantization via
the relay:
  ```
bash examples/llm_ptq/scripts/huggingface_example.sh --model
/hf-local/Qwen/Qwen3.5-35B-A3B/ --quant nvfp4_experts_only
  ```
  - 42,140 quantizers inserted, MTP layers correctly excluded
- Quantized checkpoint exported successfully (~208s, 147.61 GB peak GPU
memory)

## Test plan
- [x] Start `server.sh` inside a Docker container with the repo mounted
- [x] Run `client.sh handshake` from the host
- [x] Run `client.sh run "echo hello"` and verify output
- [x] Run `client.sh flush` and verify `.relay/` is cleared
- [x] Run a real PTQ workload end-to-end

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a file-based command relay system with host↔container client and
server CLIs, supporting handshake, run, status and flush workflows to
execute commands inside containers.

* **Documentation**
* Added guides describing the relay protocol, usage examples, CLI
options, lifecycle, and operational notes (workdir, timeouts, sequential
execution).

* **Chores**
  * Updated ignore rules to exclude ephemeral relay artifacts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-05 22:44:33 -07:00
Chenjie Luo 18ce04f1ce Update the hf_ptq.yaml (#1175)
### What does this PR do?

Type of change: Bug fix

Fix the CLI override example comment in
`tools/launcher/examples/Qwen/Qwen3-8B/hf_ptq.yaml` by adding a missing
`--` (double-dash separator) to the `task_0.args` override.

Without the `--` separator, the `--quant` flag would be parsed as an
argument to the launcher/download script rather than being passed
through to the PTQ script (`huggingface_example.sh`). This aligns the
comment example with the actual `task_0.args` definition in the YAML
(line 42), which already correctly includes the `--` separator.

### Testing

Verified the comment now matches the actual args format used in the YAML
config.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example configuration to correct command-line parameter
formatting in commented usage example.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-03 19:04:33 +00:00
Chenjie Luo 87ea8babe1 Add HuggingFace PTQ pipeline to launcher (#1100)
### What does this PR do?

Type of change: New feature

Adds a HuggingFace PTQ pipeline to the launcher, replacing the old
`hf_ptq.sh`/`hf_ptq_local.yaml` approach with a cleaner wrapper around
`huggingface_example.sh`.

**Key changes:**

- **New `common/hf/ptq.sh`** — wrapper script that downloads the model
via `huggingface-cli` if needed, then delegates to
`examples/llm_ptq/scripts/huggingface_example.sh`
- **New `examples/Qwen/Qwen3-8B/hf_ptq.yaml`** — example config for
Qwen3-8B nvfp4 quantization, supports both Slurm and local Docker
- **Removed `common/hf_ptq/hf_ptq.sh`** and
**`examples/Qwen/Qwen3-8B/hf_ptq_local.yaml`** — replaced by the new
unified pipeline
- **Configurable Slurm time limit** — `SlurmConfig.time` field replaces
the hardcoded `"04:00:00"` in `build_slurm_executor`
- **Configurable Slurm partition** — `slurm_factory` now reads
`SLURM_PARTITION` env var (default: `batch`)
- **`--clean` flag** — new `launch.py` option to `git clean -xdf` the
examples directory before job submission
- **Package `modelopt_recipes/`** — added to the nemo_run packager
include list

### Testing

- Tested HF PTQ pipeline on Slurm with Qwen3-8B

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Hugging Face PTQ workflow configuration for Qwen model
quantization.
* Added `clean` parameter to launcher for clearing directories before
job execution.

* **Improvements**
* Made Slurm execution time configurable per job instead of hardcoded
values.
  * Slurm partition configuration now respects environment variables.

* **Deprecated**
* Removed legacy PTQ launcher scripts, replaced with unified wrapper for
improved maintainability.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-02 18:55:38 +00:00
Chenjie Luo fcb09bf11d [NVBug: 6038899] Fix MoE export crash on meta tensors with CPU offload (#1155)
## Summary
Fixes `NotImplementedError` in `sync_moe_gate_up_amax` when quantizing
MoE models (e.g. Qwen3-30B-A3B) on a single GPU with insufficient VRAM.

When GPU memory is insufficient, ModelOpt enables CPU offload via
accelerate, leaving uncalibrated expert parameters on the `meta` device.
During export, `sync_moe_gate_up_amax` calls `torch.equal()` on these
meta tensors, which raises `NotImplementedError` because `aten::equal`
does not support meta tensors — even though calibration itself completed
successfully.

## Changes
- Add a guard in `sync_moe_gate_up_amax` to skip amax sync for meta
tensors (which have no real data to sync) and emit a warning explaining
the root cause.

Bug: https://nvbugspro.nvidia.com/bug/6038899

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Added warning messages for unsupported tensor configurations in
quantization workflows.
* Improved edge case detection to gracefully skip processing in
incompatible scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-02 06:59:57 +00:00
Chenjie LuoandClaude Opus 4.6 ada1e26ba0 [NVBug: 6000530] Fix AWQ crash for uncalibrated MoE experts (#1142)
## Summary
- Fixes NVBugs 6000530: `AttributeError: 'float' object has no attribute
'pow'` when running AWQ lite with `moe_calib_experts_ratio < 1.0` on MoE
models (e.g. Qwen3-30B-A3B).
- **Root cause**: When `moe_calib_experts_ratio=0.5`, some MoE experts
receive zero tokens during the AWQ cache phase, leaving `act_scale` as a
Python float `0.0` instead of a tensor. This causes two failures:
1. **Search phase crash**: Uncalibrated experts crash in `get_scale()`
because `float.pow()` doesn't exist.
2. **Export crash**: Calibrated experts have `pre_quant_scale` but
uncalibrated ones don't, causing `torch.stack()` to fail on mixed
`None`/tensor values in `preprocess_linear_fusion()`.
- **Fix**: Handle uncalibrated experts (`num_cache_steps == 0`) in two
stages:
1. **Before search**: Disable AWQ search (`is_enabled = False`) to
prevent `get_scale()` crash on float `act_scale`.
2. **During postprocessing**: Max calibrate weights and apply a neutral
(all-ones) `pre_quant_scale` so export can stack scaling factors
consistently across all experts. The `pre_quant_scale` buffer must be
registered outside `enable_weight_access_and_writeback` because HF
accelerate's `post_forward` hook drops newly-registered submodule
buffers.

## Test plan
- [x] Reproduce with `Qwen/Qwen3-30B-A3B`, `--qformat int4_awq`,
`--moe_calib_experts_ratio 0.5` — verify no crash during calibration and
export

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 13:55:21 -07:00
Chenjie LuoandKeval Morabia 16203d676a [NVBug 6007314] Deprecate MT-Bench support, remove openai pin, and add NeMo Evaluator reference (#1116)
### What does this PR do?

Type of change: Deprecation, Bug fix, Documentation

Removes MT-Bench (FastChat) evaluation support from `examples/llm_eval`
and `examples/llm_ptq`. Also removes the stale `openai>=0.28.1` pin from
`requirements.txt` that caused dependency conflicts with TRT-LLM (see
[NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314)). Adds a NeMo
Evaluator section to the llm_eval README as the recommended evaluation
workflow for quantized checkpoints.

**Changes:**
- Delete `examples/llm_eval/run_fastchat.sh` and
`examples/llm_eval/gen_model_answer.py`
- Remove `mtbench` task from `examples/llm_ptq/scripts/parser.sh` and
`huggingface_example.sh`
- Remove `openai` dependency from `examples/llm_eval/requirements.txt`
- Add NeMo Evaluator section to `examples/llm_eval/README.md` as the
recommended way to evaluate quantized checkpoints from llm_ptq via
TensorRT-LLM, vLLM, or SGLang
- Update README docs in both `llm_eval` and `llm_ptq`
- Add deprecation note to CHANGELOG.rst for 0.43

### Usage

N/A — this is a removal and documentation update.

### Testing

N/A — removed code paths; no new functionality introduced.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — MT-Bench evaluation via
`--tasks mtbench` is no longer supported.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information

Related: [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314) —
openai dependency conflict caused by FastChat's `llm_judge` extra
pinning `openai<1`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Deprecations**
* Removed MT-Bench (FastChat) evaluation support. NeMo Evaluator is now
the recommended approach for evaluating quantized model checkpoints
across multiple benchmarks.

* **Documentation**
* Updated evaluation guides to reflect NeMo Evaluator as the primary
evaluation method, with support for TensorRT-LLM, vLLM, and SGLang
serving backends.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-25 10:22:32 +00:00
Chenjie Luo 07bc4852a8 Default limit max hf quantzed safetensors file to 10GB (#1087)
### What does this PR do?

Type of change: New feature

Adds a `max_shard_size` parameter to `export_hf_checkpoint()` (and the
internal `_export_diffusers_checkpoint()`) that controls the maximum
size of each exported safetensors shard file. Defaults to `"10GB"`.

Previously, the shard size was not explicitly controlled, which could
result in very large single-file checkpoints [E.g. Qwen3.5] that are
difficult to handle (e.g., slow uploads, memory issues, or exceeding
file size limits on model hubs). This change ensures exported
checkpoints are automatically sharded into ≤10GB files by default, while
allowing users to customize the threshold.

### Usage

```python
import modelopt.torch.quantization as mtq
from modelopt.torch.export import export_hf_checkpoint

# Default: shards capped at 10GB
export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output")

# Custom shard size
export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output", max_shard_size="5GB")
```

### Testing

Verified that the `max_shard_size` parameter is correctly propagated to
both the HF `save_pretrained()` (for transformers models) and diffusers
component `save_pretrained()` calls.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (new optional parameter with
default matching previous behavior of HF's `save_pretrained`)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information

The 10GB default was chosen to keep exported safetensors files within
common file size limits while minimizing unnecessary sharding for most
models.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added configurable maximum shard size parameter for Hugging Face
checkpoint exports. When exporting both transformer and diffusion
models, users can now specify the maximum size of each safetensors
shard, controlling how checkpoint files are divided. The default shard
size is set to 10GB. This feature works with both quantized and
non-quantized models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-20 16:11:32 -07:00
Chenjie Luo acce79ffa3 Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?

Type of change: New feature

Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.

Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config

### Usage

```python
import modelopt.torch.quantization as mtq

model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```

Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe

recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```

### Testing

- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->

### Additional Information

The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.

* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-03-20 05:36:37 +00:00
Chenjie Luo 1dc890d971 Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0

## Summary

- **Remove `moe_count_expert_calib_tokens`** config field and the
`_moe_count_expert_calib_tokens` internal flag. Token counting is now
implicitly enabled when `moe_calib_experts_ratio` is set, removing a
redundant knob.
- **Change `--moe_calib_experts_ratio` default to `None`** in
`hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by
default; now the feature is opt-in and non-MoE models are unaffected
without any flag.
- **Disable `layer_sync_moe_local_experts_amax`** when
`moe_calib_experts_ratio` is set, since each expert is calibrated
independently with sufficient token coverage in that mode.
- **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks
on `_moe_calib_experts_ratio` inside the branch that already assumes it
is set.

## Changed files

| File | Change |
|------|--------|
| `modelopt/torch/quantization/config.py` | Remove
`moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio`
description to document amax sync behavior |
| `modelopt/torch/quantization/mode.py` | Remove
`moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` |
| `modelopt/torch/quantization/plugins/huggingface.py` | Remove
`_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify
`forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set |
| `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to
`None`; guard validation |
| `tests/unit/.../test_sparse_moe.py` | Update tests to use
`_moe_calib_experts_ratio` instead of removed flag |

## Test plan

- [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio`
(non-MoE model, default `None`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Configuration Changes**
* moe_calib_experts_ratio now defaults to None (disabled) instead of
1.0; validation only occurs when a value is provided.

* **Refactor**
* Simplified MoE calibration flow and token-counting behavior; removed a
deprecated expert-calibration configuration field.

* **Documentation**
* Changelog and docstrings updated to reflect the new default and
calibration behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-18 10:09:24 -07:00
Chenjie Luo 42482b1b0f Add nvfp4_omlp_only config and simplify the config.py (#973)
### What does this PR do?

Type of change: ? new feature

1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant
2) Add block sparse MOE to mlp only config
3) Simplfiy config.py
4) Update readme in llm_ptq mention these two configs for better
accuracy.

### Usage

huggingface_script.sh ... --quant nvfp4_omlp_only


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added nvfp4_omlp_only quantization format for NVFP4, enabling
selective quantization of MLP and output projection layers while
preserving attention QKV projection accuracy.

* **Changed**
* pass_through_bwd now defaults to True; set to False if using STE with
zeroed outlier gradients for better QAT accuracy.

* **Documentation**
* Updated post-training quantization guidance with NVFP4-specific
configuration recommendations and usage examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-06 01:03:25 +00:00
Chenjie Luo a4fde491cc Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do?

Type of change: Bug fix

Add moe expert calib ratio in huggingface_script.sh
Also fix minimax2.5 MOE detection which does not follow other HF MOE
layer convention

### Usage

scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4
--moe_calib_experts_ratio 1.0 --trust_remote_code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Configure MOE calibration experts ratio for quantization via an
environment/option, enabling finer control over calibration.

* **Bug Fixes**
* Improved detection of sparse MOE blocks to handle varying
expert/topology layouts, inferring expert counts when needed for more
reliable processing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-03 21:30:35 +00:00
03a1899dda Support force tokens to % of total experts during calibration (#910)
## What does this PR do?

**Type of change:** New feature

**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.

**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`

## Usage

Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5

Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq

quant_cfg = {
    "quant_cfg": { ... },
    "algorithm": {
        "method": "awq_lite",
        "moe_calib_experts_ratio": 0.25,  # calibrate 1/4 of experts
    },
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)

## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-02-24 12:35:19 -08:00
Chenjie Luo 9975ba1065 Fix DeepSeek PTQ script (#912)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?

Fix two bugs in the PTQ script

## Testing

Run DeepseekV3.2 PTQ and export


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced data type handling in quantization examples for bf16
operations
* Updated internal dependencies for quantization utilities to improve
modularity

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-20 19:45:37 +00:00
Chenjie Luo ac7c985d96 [NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?

**Type of change:** New feature, new tests

**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.

Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints

## Usage

Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:

import modelopt.torch.quantization as mtq

# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")

mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)

# During export, an .moe.html report with per-expert token counts is
saved automatically

## Testing
unittest, also test exporting qwen MOE

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.

* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-19 19:46:58 +00:00
Chenjie Luo 3801923e9d Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ?

Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ.

## Usage
scripts/huggingface_example.sh --model <minimax checkpoint> --quant
nvfp4 --trust_remote_code


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added MiniMax M2.1 model quantization support with nvfp4 format.
* Extended FP8 quantization capabilities with configurable dtype
parameter for enhanced precision control.

* **Improvements**
  * Enhanced detection of quantized linear module variants.
  * Improved weight unpacking for FP8-based linear modules.

* **Documentation**
  * Updated supported models table to include MiniMax M2.1.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-02-17 18:47:22 +00:00
Chenjie LuoandZhiyu 5e43b2a5f5 Support Qwen3 Next MTP load and export (#860)
## What does this PR do?

Fix MTP export for Qwen3 Next

**Overview:** ?

For Qwen3 next, the MTP weights are not stored separately in
safetensors. So we use "mtp" weights key to decide if the weights are
for MTP or not.


## Testing
Qwen3 Next PTQ and check if MTP is in the exported checkpoint.

scripts/huggingface_example.sh --model
<Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Optimized Multi-Token Prediction weight loading with improved layer
detection and handling.

* **Chores**
* Simplified status reporting to display total loaded weights and
detected layers.
  * Removed verbose per-file warnings for cleaner console output.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Zhiyu <zhiyuc@nvidia.com>
2026-02-09 22:48:15 +00:00
Chenjie Luo 615f99e746 Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support KIMI K2 Thinking PTQ from the original int4 checkpoint.
Tested with transformers  4.57.1, compressed-tensors 0.12.0

The model weights are dequantized on the fly to save GPU memory

## Usage
scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant
nvfp4_mlp_only --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Support for nvfp4_mlp_only quantization format, enabling new
layer-wise quantization options
  * Quantization support for CompressedLinear layers in quantized models

* **Improvements**
* Enhanced quantization for DeepSeek models with improved attention
configuration handling
* Optimized model loading with automatic precision configuration and
weight unpacking
* Better memory management during model export with automatic cache
cleanup
  * Conditional sample generation output controlled via verbose mode

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-21 09:32:33 +00:00
Chenjie Luo b0e7d9fd96 Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do?

**Overview:** ?

Unified the FP8 and NVFP4 kv cache scaling factor definition so the same
checkpoint can be used for both FP8 and NVFP4 kv cache quantization
deployment

## Testing
Unit test

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization,
simplifying configuration logic.

* **Chores**
  * Removed internal constants from public exports.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-20 08:34:32 +00:00
Chenjie Luo 9c24e2c08e Fix Deepseek transformers model loading (#740)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?
For Deepseek, let's force the user to apply trust_remote_code and use
AutoModelForCausalLM for loading the model.

## Testing
python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4
--export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64
--trust_remote_code --dataset cnn_dailymail

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-08 11:12:53 -08:00
Chenjie Luo d541324e84 Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do?

**Type of change:** ? Recipe improvement

**Overview:** ?

Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3
Next recipe for accuracy recovery

## Testing
Model accuracy benchmarking

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-02 11:06:25 -08:00
Chenjie Luo b1b9321877 Update llm_ptq doc (#685)
## What does this PR do?

**Type of change:** ? documentation

**Overview:** 
Update example doc

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-15 22:13:02 +00:00
Chenjie Luo bfec1b86fd Support Qwen3Next NVFP4 quantization (#681)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support Qwen3Next NVFP4 quantization. The QKV layers are not quantized
in PTQ to retain checkpoint accuracy across benchmarks.

## Testing
Checkpoint export and benchmarking using nemo-evaluator

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-15 21:14:04 +00:00
Chenjie Luo e4c5a68bd8 [NVBug: 5707914] Update DS V3 repo commit id (#635)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** 

DS V3 repo commit id update so the DS R1/V3.1 models can be correctly
loaded.

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-02 14:32:08 -08:00
Chenjie Luo 5ade7b03f8 [NNBUG: 5701866] Update DS V3.2 PTQ code (#630)
## What does this PR do?

**Type of change:** ?  Bug fix

**Overview:** 

1) Update the DS V3.2 repo code reference to the latest version
2) The new DS V3.2 model now includes fp32 layers. We cast it down to
match the checkpoint format during loading
3) Fix get_quant_config API change.

## Testing
Generate the deepseek-ai/DeepSeek-V3.2 checkpoint

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-02 09:36:55 -08:00
Chenjie Luo e20d218b43 [OMNIML-2857] [Experimental] Support the DeepSeek V3.2 model (#435)
## What does this PR do?

**Type of change:** ? New model support

**Overview:** ?

## Usage
Please see examples/deepseek/README.md


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* New Features
* Support for DeepSeek V3.2 quantization and automatic detection of
available DeepSeek versions.
* Triton-backed weight dequantization utility and MoE-aware calibration
mode to improve calibration fidelity.

* Documentation
* DeepSeek examples README expanded with setup, conversion, calibration,
and FP8→FP4 quantization workflows for R1, V3, and V3.2.

* Bug Fixes
* More robust, failure-tolerant copying of auxiliary files/assets during
quantization.

* Chores
  * Updated changelog and lint/ignore rules for example artifacts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-17 16:55:16 +00:00
Chenjie Luo ce8ce2229c [NVBUG: 5617733] Update LLM generate API for modelopt LLM eval (#498)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?

1) Remove kv_cache_config in the generate API. It's no longer used in
the code as well. We just estimate KV cache usage from other parameters
2) Add max_seq_len in the generate API to better estimate the real KV
cache usage.
3) Assume default lm_eval max input sequence length to be 4096

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-11-04 11:08:03 -08:00
Chenjie Luo 9a85e4922b [NVBUG: 5612606] Clear GPU cache for large models layer quantization during export (#497)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** ?

For large models like llama4 maverick, the stacked weights to fp8
conversion might hit OOM. This change aim to fix that.

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-04 18:46:58 +00:00
Chenjie Luo 5f0ef3b310 [NVBUG: 5608888] Update link in vlm_ptq README for support matrix details (#499)
## What does this PR do?

**Type of change:** ? documentation

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-04 22:52:06 +05:30
Chenjie Luo 4b522e0303 Do not modify num calib data samples to batch boundary (#483)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-10-30 15:18:42 -07:00
Chenjie Luo 32d168c522 [OMNIML-2791] Use nemotron post training dataset for calibration (#420)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-10-14 11:42:14 -07:00
Chenjie Luo 340eb7a756 Remove Qwen tokenizer modification (#390)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-10-04 17:41:08 +05:30
Chenjie Luo 17439e653d Move phi4_mm warning to above (#389)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-09-30 18:24:35 +00:00
Chenjie LuoandKeval Morabia ad091e884e Improve realquant gemm impl (#368)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-26 09:37:51 +05:30
Chenjie Luo b15a2b84ae [NVBUG: 5472822] Skip memory monitoring if not available (#374)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-09-26 09:24:27 +05:30
Chenjie Luo 598b9ce6d9 [NVBUG: 5535437] Refactor engine_dir to checkpoint_dir in PTQ examples. (#365)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-09-25 12:07:06 -07:00
Chenjie Luo 5a3fd29fb1 Add int8_sq back to auto_quant support list (#345)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-09-19 06:04:19 -07:00
Chenjie Luo b7ed8cd7e6 [NVBug: 5525758] Update VLM-PTQ readme (#339)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-09-18 16:51:00 +00:00
Chenjie Luo 912f3dcbd5 Reinstate int8_sq support for vlm_example. (#333)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-09-18 10:43:01 +05:30
146d1d9d40 Deprecate TRTLLM-build in examples (#297)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Chenjie Luo <chenjie@omniml.ai>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-18 03:22:37 +05:30
Chenjie Luo 8d0e40f9b5 Upgrade TensorRT-LLM docker to 1.1.0RC2 (#327)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-09-16 22:42:36 +00:00
Chenjie Luo 512dbb7a08 Avoid squeezing original weight tensors with leading dim 1 (#294)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-09-09 00:12:11 +00:00
Chenjie Luo 2b52759bcd Fix numerical stability of test_gemm_common.py (#283)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-09-05 09:03:17 -07:00