mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
1d21ab9e298592ad65bb8b24e21b080393a5e25b
53
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d21ab9e29 |
[DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary - DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k routing during MoE calibration. The previous all-tokens-to-all-experts path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag. - After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's `input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via `dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays per-expert; any uncalibrated expert is filled by computing amax over the dequantized FP8 weight. - `mtq.print_quant_summary` is now also written to `<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`. ## Why Forcing all tokens through every expert doubled calibration time and inflated `input_quantizer.amax` for cold-routing experts with outliers they never see at inference. The new flow matches the inference distribution, runs roughly 2x faster, and mirrors the `layer_sync_moe_local_experts_amax` semantics that mtq runs automatically for `QuantSequentialMLP`-derived MoEs. ## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG) Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default): - All 44,544 expert weight amaxes bit-identical. - Attention, shared experts, gate: identical. - Expert `w1.input` and `w3.input` (shared MoE block input): identical. - Expert `w2.input` (post-SiLU gated, expert-specific): synced to layer-wide peer max — 99.3% are larger than baseline (median 11.4x) since peer-max captures the worst-case outlier from any expert in the layer; 0.7% are smaller. This is the same trade-off `set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`. ## Test plan - [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min (vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the comparison above. - [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces `_amax_baseline/` identical (other than rounding) to the prior `CalibMoe`-default behavior. - [x] `.quant_summary.txt` written under `output_path` on rank 0. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--calib_all_experts` option to enable an alternate PTQ calibration mode; default remains top-k routing with a post-calibration per-layer peer-max synchronization and a compute fallback for uncalibrated experts. * **Documentation** * Clarified default and alternate calibration behaviors and added note about generation of a `.quant_summary.txt` summary file. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
50706d1750 |
Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary - New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and `huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g. `openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight reconstruction for the in-range blocks. - The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and `_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`. The resulting per-block scale `2^(k_j − m)` is exactly representable in E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block amax falls back to data-derived `max(|w_block|)`, which keeps the post-E4M3-clamp scale close to the block's actual magnitude. ## Verification End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only --cast_mxfp4_to_nvfp4`: ``` [cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%) [cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%) ``` End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`): ``` [cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers [cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%) [cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%) ``` Five layers fall into the OOR regime (block-spread > 17); the remaining 1,214 OOR blocks use the data-derived per-block amax fallback. Block-level losslessness is **99.99996%** end-to-end. Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant (~19B elements): | Metric | Without cast | With cast | |---|---|---| | Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** | | Total RMSE | 8.67e−02 | **0** | | max\|err\| | up to 8.0e+1 | **0** | ## Modelopt-side enablers - `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to `NVFP4StaticQuantizer` at the end of calibration. - `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was 2D-only), unblocking MoE expert weights of shape `(E, F, K)`. - BMM-experts NVFP4 export routes through `get_weights_scaling_factor_from_quantizer` for static-mode quantizers, so the pinned `_amax` is actually consumed. - `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before stacking, supporting per-block (vs scalar) static-mode amax. ## Test plan - [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` (15 tests, all passing) cover: scalar/global-amax math, per-block hybrid (in-range closed-form vs OOR data-derived), shape preservation, key collection, and end-to-end `build_amax_map` against a synthetic safetensors checkpoint. - [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only` qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100% lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks). - [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200, `nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size 4`). 67/72 layers fully lossless; 99.99996% block-level losslessness (3,583,179,586 / 3,583,180,800). - [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`: - **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample: *"Quantum computing is poised to revolutionize data analysis. However, its potential is currently limited by quantum hardware constraints, including error rates, qubit lifetimes, and lack of fault tolerance…"* - **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum computing is poised to revolutionize data storage and processing. These rare earth-based systems could serve as robust qubits; resistant to environmental decoherence…"* - [x] MSE comparison script (run separately during development) confirms per-tensor SNR=∞ across all 48 MoE expert tensors. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to enable it; helper scripts updated to expose the option. * **Bug Fixes** * Fixed static NVFP4 export for expert weights. * Improved collection/handling of quantizer amax values to avoid shape issues. * Generalized FP4 kernel to accept flexible tensor dimensionality. * Ensured static-block NVFP4 promotion during calibration. * **Tests** * Added comprehensive tests for the conversion workflow, helpers, and end-to-end application. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
168cd828c1 |
Add qwen3 moe experts only test (#1274)
## Summary - Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG` quantization config - Verifies that `hf_quant_config.json` correctly reports `quant_algo: NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in `exclude_modules` while routed expert layers (`mlp.experts.*`) do not - Reference: https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json Type of change: New tests ### Known issue On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not recognized by `get_quant_config`, causing `quant_algo=None` in the exported config. This test currently **fails** on transformers 5.x and is intended to be fixed by a follow-up change. ## Testing - **transformers 4.57.6**: PASSED - **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused expert export gap) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models with NVFP4 quantization. * Verifies the exported checkpoint records the NVFP4 quantization algorithm and that module exclusion patterns correctly exclude attention and LM head components while not excluding routed expert paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
3ad4f4f093 |
[Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary - `get_max_batch_size` halved `target_data_batch` on `torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each retry re-fed the same too-large tensor — the retry loop was effectively a no-op. - Refactor the expand logic into an `_expand_to(batch)` helper, rebuild `target_input` after halving, and call `torch.cuda.empty_cache()` between attempts. ## Test plan - [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input` mocks `torch.cuda.*` and asserts the second retry receives the halved tensor (shapes seen: `[1, 10, 5]`, regulated result `4`). - [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass (skipping the network-only minipile test). - [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on commit. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced GPU memory management during batch size detection. When out-of-memory errors occur during the initial probing phase, the system now properly adapts input tensors to smaller batch sizes and clears GPU cache before retry attempts, resulting in more reliable recovery and stable batch sizing across diverse hardware environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
47a33db9b6 |
[NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary - NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) — `Qwen3-30B-A3B nvfp4_awq` quantization fails at export with `AssertionError: Modules have different quantization formats`. - Root cause: in `model_calib.awq_lite`, MoE experts that end up disabled (NaN in act/weight scales, or no search-pass tokens) get `max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format` then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`. `unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128 experts of each linear name (gate_proj/down_proj/up_proj) and calls `preprocess_linear_fusion(..., resmooth_only=True)`, which asserts uniform format → fires for any single mismatched expert. - Fix: unify the disabled-expert paths in the awq_lite postprocess loop so any expert with `is_enabled == False` (no cache hits, NaN scales, or no search-pass tokens) receives `max_calibrate` + a neutral all-ones `pre_quant_scale`, matching the existing behavior for `num_cache_steps == 0`. Emit a warning so users notice that calibration coverage is incomplete and accuracy may degrade. ## Test plan - [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5 passed - [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a small calib set that leaves many experts uncalibrated: - All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq` (no mismatch) - `export_hf_checkpoint` succeeds with no `AssertionError` - The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning fires for each affected expert - [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI (cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to fit budget) on B200: - 36 "the second time did not forward data through ..experts.X.{gate,up,down}_proj" warnings — i.e. the exact bug-triggering condition from the original NVBug log naturally reproduces - 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for uncalibrated/disabled experts - 0 `AssertionError`s — export completes - `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq` and post-PTQ generation works --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
fda0899e40 |
feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary
- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
- `general/ptq/fp8_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-fp8_cast_kv`
- `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.
## Motivation
Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).
## Test plan
- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.
* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.
* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
01788bb007 |
Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary - Removes Mllama (Llama 3.2 Vision) model-type branches from the `llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`. - Drops the legacy `MllamaImageProcessor` path in `modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF ProcessorMixin path handles the remaining cases. - Adds a CHANGELOG entry under 0.44 Backward Breaking Changes. ## Test plan - [x] CI lint / unit tests pass - [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh --model <llm> --quant fp8`` (text-only path, non-mllama) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed Mllama (Llama 3.2 Vision) support from quantization examples. This includes removal of dedicated image processor implementation, specialized model handling, and related calibration logic. * Updated VLM image-text calibration guidance to use `--calib_with_images` flag with other supported VLMs instead of Mllama-specific processing paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
04fcf24227 |
Fix LLM deploy test failure by defaulting expert parallelism to 1 (#1273)
### What does this PR do? Type of change: Bug fix Fixes TRT-LLM DeepEP kernel failures during LLM deployment on unsupported GPUs (e.g. Blackwell SM 12.0) by defaulting expert parallelism (`ep`) to 1 instead of auto-setting it to the GPU count for MoE models. Previously, when the model config contained expert-related keys, `ep` was automatically set to `torch.cuda.device_count()`, which triggered DeepEP kernel failures on GPUs that don't support it. Now `ep` defaults to 1 while still enabling attention data parallelism for MoE models. Expert parallelism can be enabled explicitly by the caller when the environment is known to support it. ### Testing - [x] Verified that the `llm_ptq` test passes with this fix on Blackwell GPUs. - [x] 2-gpu CI test triggered: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24495054531/job/71588037727 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
d45219b390 |
Fix debugger server failing to detect editable-installed modelopt (#1270)
## Summary - Removed `PYTHONPATH="" python -I` override in `check_modelopt_local()` so the PYTHONPATH validation uses the actual environment instead of an isolated one - Moved the workdir log line earlier in `server.sh` for better debugging visibility ## Test plan - [x] Start `server.sh` inside a Docker container and verify it correctly detects editable-installed modelopt - [x] Confirm the workdir is logged before the modelopt check runs 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Server startup now displays the configured work directory earlier in the initialization process, providing improved visibility of the active directory during server launch. * Simplified the modelopt validation check during server initialization while maintaining the same validation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
7c8557158d |
Add job cancellation support to the debugger command relay (#1262)
## Summary - Add a `cancel` subcommand to the client that terminates the currently running command on the server - Server now runs commands in the background with PID tracking, enabling cancellation mid-execution - Client-side timeouts automatically cancel the running command on the server (previously the server process was left running) - Hardened against race conditions through 4 rounds of adversarial review (15 fixes total) ### Key changes **server.sh:** - Commands run in background with PID tracked in `$RELAY_DIR/running` (atomic tmp+mv write) - Cancel detection loop checks for `$RELAY_DIR/cancel` file with cmd_id verification - SIGTERM with 5s grace period, then SIGKILL escalation for stuck processes - `.exit` file written before `running` marker removed (ordering guarantee) - `set -e`-safe: `wait` uses `|| exit_code=$?` pattern; cleanup trap fully guarded - Stale cancel files cleared at command start; mismatched/empty signals rejected - Command file read into memory and removed before execution (eliminates TOCTOU with client timeout) **client.sh:** - New `cancel` subcommand: writes target cmd_id to cancel file, waits for server acknowledgment (30s timeout) - `run` timeout now sends targeted cancel signal (verifies cmd_id match to avoid killing wrong command) - `run` timeout cleans up orphaned result files - `status` shows currently running command - `flush` rejects if a command is currently running (prevents state corruption) - Exit code validated as numeric before use ### Protocol additions ``` .relay/ ├── running # server writes cmd_id:pid while executing (atomic) ├── cancel # client writes target cmd_id to request cancellation ``` ## Test plan - [ ] Start server in Docker, handshake from host - [ ] Run a command (`client.sh run "sleep 30"`), cancel it (`client.sh cancel`), verify exit code 130 - [ ] Run a command with short timeout (`--timeout 5 run "sleep 30"`), verify auto-cancel - [ ] Run a command that exits non-zero, verify server stays alive - [ ] Run `status` during execution, verify it shows the running command - [ ] Attempt `flush` during execution, verify it is rejected - [ ] Cancel when nothing is running, verify clean message 🤖 Generated with [Claude Code](https://claude.com/claude-code) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (bash scripts for dev tooling, tested manually) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A (internal tooling) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a comprehensive debug skill guide and protocol reference with quick CLI examples and a new “Cancelling Commands” section. * **New Features** * Client-side `cancel` command to terminate the currently running remote command. * Status now reports active command (or `(idle)`). * **Improvements** * Stronger startup/validation guidance, safer shutdown/cleanup, deterministic cancel exit semantics (130), and auto-cancel on client-side timeout. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
952a62bf65 |
Fix missing attention_mask in calibration dataloader (#1261)
## Summary - When `include_labels=False` (the default for PTQ calibration), `get_dataset_dataloader` was discarding the `attention_mask` produced by the tokenizer and only returning `input_ids`. - Without `attention_mask`, HuggingFace models create a full causal mask, causing padding tokens to participate in attention during calibration and skewing quantization statistics. - This fix includes `attention_mask` alongside `input_ids` so the model correctly ignores padding tokens during calibration forward passes. ## Details In `modelopt/torch/utils/dataset_utils.py`, the tokenizer call at line 387 with `padding=True` produces both `input_ids` and `attention_mask`. The `include_labels=True` path (line 406) already preserves the full `batch_encoded` dict including `attention_mask`. However, the `include_labels=False` path was only keeping `input_ids` "for backward compatibility." During the calibration forward loop (`_forward_loop` → `_process_batch`), the batch dict is unpacked as `**kwargs` into `model.forward()`. Without `attention_mask`, HF models default to attending to all positions including padding, which pollutes calibration statistics. **Practical impact**: With `batch_size=1` there is no padding so the bug is invisible. With larger batch sizes and variable-length samples, shorter sequences get padded and the effect grows. ## Test plan - [x] Existing unit tests pass (`tests/unit/torch/utils/test_dataset_utils.py`) - [x] Pre-commit hooks pass - [ ] Verify PTQ accuracy with batch_size > 1 on a padded calibration dataset (GPU required) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f7557221e3 |
Add file-based command relay for remote Docker testing (#1174)
## Summary - Adds a lightweight file-based client/server relay (`tools/debugger/`) that enables Claude Code (or any host-side automation) to execute commands inside a remote Docker container using only a shared filesystem — no networking setup required. - The server auto-detects the repo root, installs modelopt (`pip install -e .[dev]`), sets `PYTHONPATH`, and listens for commands. - The client supports `handshake`, `run`, `status`, and `flush` subcommands. - Includes `README.md` (full protocol docs) and `CLAUDE.md` (quick reference for Claude Code). ## Tested - Ran Qwen3.5-35B-A3B MoE PTQ with `nvfp4_experts_only` quantization via the relay: ``` bash examples/llm_ptq/scripts/huggingface_example.sh --model /hf-local/Qwen/Qwen3.5-35B-A3B/ --quant nvfp4_experts_only ``` - 42,140 quantizers inserted, MTP layers correctly excluded - Quantized checkpoint exported successfully (~208s, 147.61 GB peak GPU memory) ## Test plan - [x] Start `server.sh` inside a Docker container with the repo mounted - [x] Run `client.sh handshake` from the host - [x] Run `client.sh run "echo hello"` and verify output - [x] Run `client.sh flush` and verify `.relay/` is cleared - [x] Run a real PTQ workload end-to-end 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a file-based command relay system with host↔container client and server CLIs, supporting handshake, run, status and flush workflows to execute commands inside containers. * **Documentation** * Added guides describing the relay protocol, usage examples, CLI options, lifecycle, and operational notes (workdir, timeouts, sequential execution). * **Chores** * Updated ignore rules to exclude ephemeral relay artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
18ce04f1ce |
Update the hf_ptq.yaml (#1175)
### What does this PR do? Type of change: Bug fix Fix the CLI override example comment in `tools/launcher/examples/Qwen/Qwen3-8B/hf_ptq.yaml` by adding a missing `--` (double-dash separator) to the `task_0.args` override. Without the `--` separator, the `--quant` flag would be parsed as an argument to the launcher/download script rather than being passed through to the PTQ script (`huggingface_example.sh`). This aligns the comment example with the actual `task_0.args` definition in the YAML (line 42), which already correctly includes the `--` separator. ### Testing Verified the comment now matches the actual args format used in the YAML config. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated example configuration to correct command-line parameter formatting in commented usage example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
87ea8babe1 |
Add HuggingFace PTQ pipeline to launcher (#1100)
### What does this PR do? Type of change: New feature Adds a HuggingFace PTQ pipeline to the launcher, replacing the old `hf_ptq.sh`/`hf_ptq_local.yaml` approach with a cleaner wrapper around `huggingface_example.sh`. **Key changes:** - **New `common/hf/ptq.sh`** — wrapper script that downloads the model via `huggingface-cli` if needed, then delegates to `examples/llm_ptq/scripts/huggingface_example.sh` - **New `examples/Qwen/Qwen3-8B/hf_ptq.yaml`** — example config for Qwen3-8B nvfp4 quantization, supports both Slurm and local Docker - **Removed `common/hf_ptq/hf_ptq.sh`** and **`examples/Qwen/Qwen3-8B/hf_ptq_local.yaml`** — replaced by the new unified pipeline - **Configurable Slurm time limit** — `SlurmConfig.time` field replaces the hardcoded `"04:00:00"` in `build_slurm_executor` - **Configurable Slurm partition** — `slurm_factory` now reads `SLURM_PARTITION` env var (default: `batch`) - **`--clean` flag** — new `launch.py` option to `git clean -xdf` the examples directory before job submission - **Package `modelopt_recipes/`** — added to the nemo_run packager include list ### Testing - Tested HF PTQ pipeline on Slurm with Qwen3-8B ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Hugging Face PTQ workflow configuration for Qwen model quantization. * Added `clean` parameter to launcher for clearing directories before job execution. * **Improvements** * Made Slurm execution time configurable per job instead of hardcoded values. * Slurm partition configuration now respects environment variables. * **Deprecated** * Removed legacy PTQ launcher scripts, replaced with unified wrapper for improved maintainability. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
fcb09bf11d |
[NVBug: 6038899] Fix MoE export crash on meta tensors with CPU offload (#1155)
## Summary Fixes `NotImplementedError` in `sync_moe_gate_up_amax` when quantizing MoE models (e.g. Qwen3-30B-A3B) on a single GPU with insufficient VRAM. When GPU memory is insufficient, ModelOpt enables CPU offload via accelerate, leaving uncalibrated expert parameters on the `meta` device. During export, `sync_moe_gate_up_amax` calls `torch.equal()` on these meta tensors, which raises `NotImplementedError` because `aten::equal` does not support meta tensors — even though calibration itself completed successfully. ## Changes - Add a guard in `sync_moe_gate_up_amax` to skip amax sync for meta tensors (which have no real data to sync) and emit a warning explaining the root cause. Bug: https://nvbugspro.nvidia.com/bug/6038899 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Added warning messages for unsupported tensor configurations in quantization workflows. * Improved edge case detection to gracefully skip processing in incompatible scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
ada1e26ba0 |
[NVBug: 6000530] Fix AWQ crash for uncalibrated MoE experts (#1142)
## Summary - Fixes NVBugs 6000530: `AttributeError: 'float' object has no attribute 'pow'` when running AWQ lite with `moe_calib_experts_ratio < 1.0` on MoE models (e.g. Qwen3-30B-A3B). - **Root cause**: When `moe_calib_experts_ratio=0.5`, some MoE experts receive zero tokens during the AWQ cache phase, leaving `act_scale` as a Python float `0.0` instead of a tensor. This causes two failures: 1. **Search phase crash**: Uncalibrated experts crash in `get_scale()` because `float.pow()` doesn't exist. 2. **Export crash**: Calibrated experts have `pre_quant_scale` but uncalibrated ones don't, causing `torch.stack()` to fail on mixed `None`/tensor values in `preprocess_linear_fusion()`. - **Fix**: Handle uncalibrated experts (`num_cache_steps == 0`) in two stages: 1. **Before search**: Disable AWQ search (`is_enabled = False`) to prevent `get_scale()` crash on float `act_scale`. 2. **During postprocessing**: Max calibrate weights and apply a neutral (all-ones) `pre_quant_scale` so export can stack scaling factors consistently across all experts. The `pre_quant_scale` buffer must be registered outside `enable_weight_access_and_writeback` because HF accelerate's `post_forward` hook drops newly-registered submodule buffers. ## Test plan - [x] Reproduce with `Qwen/Qwen3-30B-A3B`, `--qformat int4_awq`, `--moe_calib_experts_ratio 0.5` — verify no crash during calibration and export 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
16203d676a |
[NVBug 6007314] Deprecate MT-Bench support, remove openai pin, and add NeMo Evaluator reference (#1116)
### What does this PR do? Type of change: Deprecation, Bug fix, Documentation Removes MT-Bench (FastChat) evaluation support from `examples/llm_eval` and `examples/llm_ptq`. Also removes the stale `openai>=0.28.1` pin from `requirements.txt` that caused dependency conflicts with TRT-LLM (see [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314)). Adds a NeMo Evaluator section to the llm_eval README as the recommended evaluation workflow for quantized checkpoints. **Changes:** - Delete `examples/llm_eval/run_fastchat.sh` and `examples/llm_eval/gen_model_answer.py` - Remove `mtbench` task from `examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh` - Remove `openai` dependency from `examples/llm_eval/requirements.txt` - Add NeMo Evaluator section to `examples/llm_eval/README.md` as the recommended way to evaluate quantized checkpoints from llm_ptq via TensorRT-LLM, vLLM, or SGLang - Update README docs in both `llm_eval` and `llm_ptq` - Add deprecation note to CHANGELOG.rst for 0.43 ### Usage N/A — this is a removal and documentation update. ### Testing N/A — removed code paths; no new functionality introduced. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ — MT-Bench evaluation via `--tasks mtbench` is no longer supported. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional Information Related: [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314) — openai dependency conflict caused by FastChat's `llm_judge` extra pinning `openai<1`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecations** * Removed MT-Bench (FastChat) evaluation support. NeMo Evaluator is now the recommended approach for evaluating quantized model checkpoints across multiple benchmarks. * **Documentation** * Updated evaluation guides to reflect NeMo Evaluator as the primary evaluation method, with support for TensorRT-LLM, vLLM, and SGLang serving backends. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
07bc4852a8 |
Default limit max hf quantzed safetensors file to 10GB (#1087)
### What does this PR do? Type of change: New feature Adds a `max_shard_size` parameter to `export_hf_checkpoint()` (and the internal `_export_diffusers_checkpoint()`) that controls the maximum size of each exported safetensors shard file. Defaults to `"10GB"`. Previously, the shard size was not explicitly controlled, which could result in very large single-file checkpoints [E.g. Qwen3.5] that are difficult to handle (e.g., slow uploads, memory issues, or exceeding file size limits on model hubs). This change ensures exported checkpoints are automatically sharded into ≤10GB files by default, while allowing users to customize the threshold. ### Usage ```python import modelopt.torch.quantization as mtq from modelopt.torch.export import export_hf_checkpoint # Default: shards capped at 10GB export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output") # Custom shard size export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output", max_shard_size="5GB") ``` ### Testing Verified that the `max_shard_size` parameter is correctly propagated to both the HF `save_pretrained()` (for transformers models) and diffusers component `save_pretrained()` calls. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ (new optional parameter with default matching previous behavior of HF's `save_pretrained`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ❌ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ ### Additional Information The 10GB default was chosen to keep exported safetensors files within common file size limits while minimizing unnecessary sharding for most models. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added configurable maximum shard size parameter for Hugging Face checkpoint exports. When exporting both transformer and diffusion models, users can now specify the maximum size of each safetensors shard, controlling how checkpoint files are divided. The default shard size is set to 10GB. This feature works with both quantized and non-quantized models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
acce79ffa3 |
Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?
Type of change: New feature
Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.
Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config
### Usage
```python
import modelopt.torch.quantization as mtq
model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```
Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe
recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```
### Testing
- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->
### Additional Information
The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.
* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
|
||
|
|
1dc890d971 |
Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0 ## Summary - **Remove `moe_count_expert_calib_tokens`** config field and the `_moe_count_expert_calib_tokens` internal flag. Token counting is now implicitly enabled when `moe_calib_experts_ratio` is set, removing a redundant knob. - **Change `--moe_calib_experts_ratio` default to `None`** in `hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by default; now the feature is opt-in and non-MoE models are unaffected without any flag. - **Disable `layer_sync_moe_local_experts_amax`** when `moe_calib_experts_ratio` is set, since each expert is calibrated independently with sufficient token coverage in that mode. - **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks on `_moe_calib_experts_ratio` inside the branch that already assumes it is set. ## Changed files | File | Change | |------|--------| | `modelopt/torch/quantization/config.py` | Remove `moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio` description to document amax sync behavior | | `modelopt/torch/quantization/mode.py` | Remove `moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` | | `modelopt/torch/quantization/plugins/huggingface.py` | Remove `_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify `forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set | | `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to `None`; guard validation | | `tests/unit/.../test_sparse_moe.py` | Update tests to use `_moe_calib_experts_ratio` instead of removed flag | ## Test plan - [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio` (non-MoE model, default `None`) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Configuration Changes** * moe_calib_experts_ratio now defaults to None (disabled) instead of 1.0; validation only occurs when a value is provided. * **Refactor** * Simplified MoE calibration flow and token-counting behavior; removed a deprecated expert-calibration configuration field. * **Documentation** * Changelog and docstrings updated to reflect the new default and calibration behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
42482b1b0f |
Add nvfp4_omlp_only config and simplify the config.py (#973)
### What does this PR do? Type of change: ? new feature 1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant 2) Add block sparse MOE to mlp only config 3) Simplfiy config.py 4) Update readme in llm_ptq mention these two configs for better accuracy. ### Usage huggingface_script.sh ... --quant nvfp4_omlp_only ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added nvfp4_omlp_only quantization format for NVFP4, enabling selective quantization of MLP and output projection layers while preserving attention QKV projection accuracy. * **Changed** * pass_through_bwd now defaults to True; set to False if using STE with zeroed outlier gradients for better QAT accuracy. * **Documentation** * Updated post-training quantization guidance with NVFP4-specific configuration recommendations and usage examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
a4fde491cc |
Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do? Type of change: Bug fix Add moe expert calib ratio in huggingface_script.sh Also fix minimax2.5 MOE detection which does not follow other HF MOE layer convention ### Usage scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4 --moe_calib_experts_ratio 1.0 --trust_remote_code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Configure MOE calibration experts ratio for quantization via an environment/option, enabling finer control over calibration. * **Bug Fixes** * Improved detection of sparse MOE blocks to handle varying expert/topology layouts, inferring expert counts when needed for more reliable processing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
03a1899dda |
Support force tokens to % of total experts during calibration (#910)
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.
**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`
## Usage
Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5
Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq
quant_cfg = {
"quant_cfg": { ... },
"algorithm": {
"method": "awq_lite",
"moe_calib_experts_ratio": 0.25, # calibrate 1/4 of experts
},
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)
## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
||
|
|
9975ba1065 |
Fix DeepSeek PTQ script (#912)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? Fix two bugs in the PTQ script ## Testing Run DeepseekV3.2 PTQ and export <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced data type handling in quantization examples for bf16 operations * Updated internal dependencies for quantization utilities to improve modularity <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
ac7c985d96 |
[NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?
**Type of change:** New feature, new tests
**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.
Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints
## Usage
Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:
import modelopt.torch.quantization as mtq
# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")
mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)
# During export, an .moe.html report with per-expert token counts is
saved automatically
## Testing
unittest, also test exporting qwen MOE
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.
* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
3801923e9d |
Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ. ## Usage scripts/huggingface_example.sh --model <minimax checkpoint> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added MiniMax M2.1 model quantization support with nvfp4 format. * Extended FP8 quantization capabilities with configurable dtype parameter for enhanced precision control. * **Improvements** * Enhanced detection of quantized linear module variants. * Improved weight unpacking for FP8-based linear modules. * **Documentation** * Updated supported models table to include MiniMax M2.1. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
5e43b2a5f5 |
Support Qwen3 Next MTP load and export (#860)
## What does this PR do? Fix MTP export for Qwen3 Next **Overview:** ? For Qwen3 next, the MTP weights are not stored separately in safetensors. So we use "mtp" weights key to decide if the weights are for MTP or not. ## Testing Qwen3 Next PTQ and check if MTP is in the exported checkpoint. scripts/huggingface_example.sh --model <Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Optimized Multi-Token Prediction weight loading with improved layer detection and handling. * **Chores** * Simplified status reporting to display total loaded weights and detected layers. * Removed verbose per-file warnings for cleaner console output. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Zhiyu <zhiyuc@nvidia.com> |
||
|
|
615f99e746 |
Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support KIMI K2 Thinking PTQ from the original int4 checkpoint. Tested with transformers 4.57.1, compressed-tensors 0.12.0 The model weights are dequantized on the fly to save GPU memory ## Usage scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant nvfp4_mlp_only --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support for nvfp4_mlp_only quantization format, enabling new layer-wise quantization options * Quantization support for CompressedLinear layers in quantized models * **Improvements** * Enhanced quantization for DeepSeek models with improved attention configuration handling * Optimized model loading with automatic precision configuration and weight unpacking * Better memory management during model export with automatic cache cleanup * Conditional sample generation output controlled via verbose mode <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b0e7d9fd96 |
Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do? **Overview:** ? Unified the FP8 and NVFP4 kv cache scaling factor definition so the same checkpoint can be used for both FP8 and NVFP4 kv cache quantization deployment ## Testing Unit test ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization, simplifying configuration logic. * **Chores** * Removed internal constants from public exports. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
9c24e2c08e |
Fix Deepseek transformers model loading (#740)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? For Deepseek, let's force the user to apply trust_remote_code and use AutoModelForCausalLM for loading the model. ## Testing python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4 --export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64 --trust_remote_code --dataset cnn_dailymail ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
d541324e84 |
Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do? **Type of change:** ? Recipe improvement **Overview:** ? Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3 Next recipe for accuracy recovery ## Testing Model accuracy benchmarking Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b1b9321877 |
Update llm_ptq doc (#685)
## What does this PR do? **Type of change:** ? documentation **Overview:** Update example doc Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
bfec1b86fd |
Support Qwen3Next NVFP4 quantization (#681)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support Qwen3Next NVFP4 quantization. The QKV layers are not quantized in PTQ to retain checkpoint accuracy across benchmarks. ## Testing Checkpoint export and benchmarking using nemo-evaluator ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
e4c5a68bd8 |
[NVBug: 5707914] Update DS V3 repo commit id (#635)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** DS V3 repo commit id update so the DS R1/V3.1 models can be correctly loaded. Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
5ade7b03f8 |
[NNBUG: 5701866] Update DS V3.2 PTQ code (#630)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** 1) Update the DS V3.2 repo code reference to the latest version 2) The new DS V3.2 model now includes fp32 layers. We cast it down to match the checkpoint format during loading 3) Fix get_quant_config API change. ## Testing Generate the deepseek-ai/DeepSeek-V3.2 checkpoint ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
e20d218b43 |
[OMNIML-2857] [Experimental] Support the DeepSeek V3.2 model (#435)
## What does this PR do? **Type of change:** ? New model support **Overview:** ? ## Usage Please see examples/deepseek/README.md <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * New Features * Support for DeepSeek V3.2 quantization and automatic detection of available DeepSeek versions. * Triton-backed weight dequantization utility and MoE-aware calibration mode to improve calibration fidelity. * Documentation * DeepSeek examples README expanded with setup, conversion, calibration, and FP8→FP4 quantization workflows for R1, V3, and V3.2. * Bug Fixes * More robust, failure-tolerant copying of auxiliary files/assets during quantization. * Chores * Updated changelog and lint/ignore rules for example artifacts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ce8ce2229c |
[NVBUG: 5617733] Update LLM generate API for modelopt LLM eval (#498)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? 1) Remove kv_cache_config in the generate API. It's no longer used in the code as well. We just estimate KV cache usage from other parameters 2) Add max_seq_len in the generate API to better estimate the real KV cache usage. 3) Assume default lm_eval max input sequence length to be 4096 Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
9a85e4922b |
[NVBUG: 5612606] Clear GPU cache for large models layer quantization during export (#497)
## What does this PR do? **Type of change:** Bug fix **Overview:** ? For large models like llama4 maverick, the stacked weights to fp8 conversion might hit OOM. This change aim to fix that. --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
5f0ef3b310 |
[NVBUG: 5608888] Update link in vlm_ptq README for support matrix details (#499)
## What does this PR do? **Type of change:** ? documentation --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
4b522e0303 |
Do not modify num calib data samples to batch boundary (#483)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
32d168c522 |
[OMNIML-2791] Use nemotron post training dataset for calibration (#420)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
340eb7a756 |
Remove Qwen tokenizer modification (#390)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
17439e653d |
Move phi4_mm warning to above (#389)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ad091e884e |
Improve realquant gemm impl (#368)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
b15a2b84ae |
[NVBUG: 5472822] Skip memory monitoring if not available (#374)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
598b9ce6d9 |
[NVBUG: 5535437] Refactor engine_dir to checkpoint_dir in PTQ examples. (#365)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
5a3fd29fb1 |
Add int8_sq back to auto_quant support list (#345)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
b7ed8cd7e6 |
[NVBug: 5525758] Update VLM-PTQ readme (#339)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
912f3dcbd5 |
Reinstate int8_sq support for vlm_example. (#333)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
146d1d9d40 |
Deprecate TRTLLM-build in examples (#297)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Chenjie Luo <chenjie@omniml.ai> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
8d0e40f9b5 |
Upgrade TensorRT-LLM docker to 1.1.0RC2 (#327)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
512dbb7a08 |
Avoid squeezing original weight tensors with leading dim 1 (#294)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
2b52759bcd |
Fix numerical stability of test_gemm_common.py (#283)
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |