mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ee5c256204 |
Noeyy/fix bug 6701777 (#2402)
### What does this PR do? Type of change: Bug fix: 6701777 Regression source: "[OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT" (#1289), merged into 0.44.0rc3 via the batch cherry-pick #1350. This PR: 1. Registers nn.LayerNorm as a QuantModule for the first time (modelopt/torch/quantization/nn/modules/quant_layernorm.py), intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output quantizer rules apply to ViT. 2. Removes the prior forced Cast-alignment logic in export_onnx.py that used to normalize Q/DQ node dtypes to trt_high_precision_dtype. Root Cause: Once nn.LayerNorm became a registered QuantModule, these wildcards started unintentionally matching norm1.norm inside FLUX's AdaLayerNormZero block — an elementwise_affine=False LayerNorm with no learnable weight/bias. Its input got routed through NVFP4 Q/DQ (emitted as Float32) while its synthesized affine scale remained native BFloat16, producing the dtype mismatch. Chosen fix: Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets rather than touching the global QuantModuleRegistry (which ViT FP8 MHA still needs). Add, in both modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules (list order matters — later entries override earlier ones): - parent_class: 'nn.LayerNorm' quantizer_name: '*' enable: false ### Usage ``` python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16 trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 ``` ### Testing The above test commands. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by excluding LayerNorm modules from quantization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
a1bcda4727 |
Fix protobuf size-check failures in ONNX deployment (#2403)
### What does this PR do? Type of change: Bug fix:6701737 The ONNX deployment path assumed that ModelProto.ByteSize() would always return a valid size. With newer protobuf versions, querying the size of a model exceeding the protobuf serialization limit can itself raise EncodeError: Failed to serialize proto. Replaced both direct size checks with the existing is_model_too_large_for_protobuf() helper. This helper handles size-query failures conservatively and checks the protobuf size limit: Shape inference now selects the external-data/file-based path when ByteSize() fails or the model is too large. Metadata creation uses the same safe check instead of raising another serialization error. The unused TWO_GB constant was also removed. ### Usage ``` python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image ``` ### Testing the above test case pass on B100 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved ONNX model size detection during shape inference and export processing. * Ensured large models consistently use the appropriate external-data handling path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
5d2d5a5d15 |
deprecate trtllm-build in weight_sparsity (#2371)
### What does this PR do?
Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190
<!-- Details about the change. -->
### Usage
```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt --calib_size 128 --output_dir Llama-3.1-8B-Instruct_pts
python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts
trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
--tp_size 1 \
--pp_size 1 \
--host 0.0.0.0 \
--port 8000
```
### Testing
PTS and SAT tested
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
* Corrected the PTS model restoration path.
* **Bug Fixes**
* Model export now saves the tokenizer alongside the checkpoint.
* Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
|
||
|
|
f10d5183ec |
launcher: bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq… (#1982)
bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq_len for Qwen3.5-4B ### What does this PR do? Type of change: Bug fix, new feature - Upgrade TRT-LLM container from 1.3.0rc10 to 1.3.0rc20 across all Qwen3-8B, Qwen3-30B-A3B, Kimi-K2.5, and gpt-oss-20b launcher configs. - Replace Kimi-K2.5 aarch64-specific vLLM image (v0.22.0-aarch64) with the multi-arch v0.22.0 tag, which resolves to amd64/arm64 automatically. - Fix Qwen3.5-4B throughput_32k runs: raise max_seq_len from 40960 to 65536 to accommodate outlier prompts (~46.6K tokens) that caused VLLMValidationError and aborted the entire benchmark run. - Fix specdec_bench_mtp_vllm.yaml: remove reference to non-existent runtime_params_throughput_32k.yaml; use --max_seq_len 65536 instead ### Usage ``` cd Model-Optimizer/tools/launcher uv run launch.py --yaml examples/Qwen/Qwen3-8B/megatron_lm_ptq_local.yaml hf_local=/mnt/hf-local --yes ``` ### Testing N/A - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?:N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Updated launcher example pipelines and the default launcher Slurm container to the newer TensorRT-LLM `1.3.0rc20` image. * Updated Kimi-K2.5 workflows to use the multi-architecture vLLM `0.22.0` image (removing prior architecture-specific variants). * Increased the Qwen3.5-4B long-context benchmark maximum sequence length to 65,536 tokens and simplified the corresponding throughput configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
31f0783e35 |
add new cases for ckpts on HF (#938)
## What does this PR do? **Type of change:** new tests **Overview:** Add new deployment tests for newly added checkpoints on HF ## Usage pytest test_deploy.py --run-release ```python None ``` ## Testing None ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for new model deployments: Qwen3 variants (480B, 235B, 397B), Kimi K2.5, and additional Mixtral configurations. * **Improvements** * Expanded backend support for select models, including TensorRT-LLM integration for Nemotron models. * **Chores** * Renamed environment variable from MODELOPT_LOCAL_MODEL_ROOT to MODELOPT_LOCAL_EAGLE_MODEL for improved clarity. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
9de9b8f20c |
Noeyy/add test cases for the newly added checkpoints on HF (#827)
## What does this PR do? **Type of change:** new tests **Overview:** Add new test cases for the newly added checkpoints on HuggingFace. ## Usage pytest test_deploy.py --run-release ```python None ``` ## Testing None ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information None <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for NVFP4 model variants across multiple model families (DeepSeek, Llama, Qwen, and others). * **Improvements** * Enhanced backend availability detection to automatically identify and manage supported deployment backends at runtime. * **Tests** * Improved test infrastructure for better reproducibility and backend compatibility handling. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
883c8731aa |
Noeyy/add new cases for newly added checkpoints on HF (#728)
## What does this PR do? **Type of change:** Add TRT LLM/vLLM/SGLang functional test cases for newly added checkpoints on HF **Overview:** 1.Since the speculative draft model only supports loading from a local path, we should set the MODELOPT_LOCAL_MODEL_ROOT environment variable. If we don't set it, these test cases will be skipped. 2. Newly added checkpoints: - nvidia/gpt-oss-120b-Eagle3-short-context - nvidia/gpt-oss-120b-Eagle3-throughput - nvidia/EAGLE3-NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 ## Usage ```python pytest tests/examples/llm_ptq/test_deploy.py --run-release ``` ## Testing Run release testing ## Before your PR is "*Ready for review*" Ready for review - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information N/A --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
c77eebcacc |
Noeyy/add new ckpts test cases (#650)
## What does this PR do? **Type of change:** new tests **Overview:** Add new checkpoints test cases. ## Usage pytest test_deploy.py --run-release ```python # Add a code snippet demonstrating how to use this ``` ## Testing N//A ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information N/A --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
d311edaf5f |
Add functional test cases for published checkpoints on HF (#455)
## What does this PR do? **Type of change:** new tests **Overview:** Add vLLM/SGLang/TRT LLM deployment tests for published checkpoints on HF. Add e2e test case for gpt-oss ## Usage ```python # pytest test_deploy.py -k "vllm" ``` ## Testing ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added ModelDeployer and ModelDeployerList test utilities with COMMON_PROMPTS to run multi-backend deployment and inference scenarios. * Introduced an end-to-end GPT-OSS QAT pipeline test (SFT → QAT → MXFP4 conversion) with optional deployment/benchmarking. * Added a broad parametric LLM deployment test suite covering many models/backends with readable test IDs. * Added automatic Hugging Face cache cleanup after tests to keep runs isolated. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: neoyy-mino <174223378+noeyy-mino@users.noreply.github.com> Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
970044c02e |
update model_type of Qwen (#477)
Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |