10 Commits
Author SHA1 Message Date
noeyy-mino ee5c256204 Noeyy/fix bug 6701777 (#2402)
### What does this PR do?

Type of change: Bug fix: 6701777

Regression source: "[OMNIML-3349] Add FP8 MHA
quantization support for HuggingFace ViT" (#1289), merged into
0.44.0rc3 via the batch cherry-pick #1350. This PR:
  1. Registers nn.LayerNorm as a QuantModule for the first time
     (modelopt/torch/quantization/nn/modules/quant_layernorm.py),
     intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output
     quantizer rules apply to ViT.
  2. Removes the prior forced Cast-alignment logic in export_onnx.py
that used to normalize Q/DQ node dtypes to trt_high_precision_dtype.

Root Cause:
Once nn.LayerNorm became a registered QuantModule, these wildcards
started unintentionally matching norm1.norm inside FLUX's
AdaLayerNormZero block — an elementwise_affine=False LayerNorm with
no learnable weight/bias. Its input got routed through NVFP4 Q/DQ
(emitted as Float32) while its synthesized affine scale remained
native BFloat16, producing the dtype mismatch. 

Chosen fix:
Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets
rather than touching the global QuantModuleRegistry (which ViT FP8
MHA still needs). Add, in both
modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and
nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules
(list order matters — later entries override earlier ones):

  - parent_class: 'nn.LayerNorm'
    quantizer_name: '*'
    enable: false

### Usage

```
python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16

trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1
```

### Testing
 The above test commands.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by
excluding LayerNorm modules from quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-09-22 00:55:33 +00:00
noeyy-mino a1bcda4727 Fix protobuf size-check failures in ONNX deployment (#2403)
### What does this PR do?

Type of change:  Bug fix:6701737

The ONNX deployment path assumed that ModelProto.ByteSize() would always
return a valid size. With newer protobuf versions, querying the size of
a model exceeding the protobuf serialization limit can itself raise
EncodeError: Failed to serialize proto.

Replaced both direct size checks with the existing
is_model_too_large_for_protobuf() helper. This helper handles size-query
failures conservatively and checks the protobuf size limit:

Shape inference now selects the external-data/file-based path when
ByteSize() fails or the model is too large.
Metadata creation uses the same safe check instead of raising another
serialization error.
The unused TWO_GB constant was also removed.

### Usage

```
python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image
```

### Testing
the above test case pass on B100

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved ONNX model size detection during shape inference and export
processing.
* Ensured large models consistently use the appropriate external-data
handling path.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
2026-09-11 10:46:10 -04:00
noeyy-mino 5d2d5a5d15 deprecate trtllm-build in weight_sparsity (#2371)
### What does this PR do?

Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190

<!-- Details about the change. -->

### Usage

```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt  --calib_size 128  --output_dir Llama-3.1-8B-Instruct_pts

python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts

trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
    --tp_size 1 \
    --pp_size 1 \
    --host 0.0.0.0 \
    --port 8000

```

### Testing
PTS and SAT tested

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
  * Corrected the PTS model restoration path.

* **Bug Fixes**
  * Model export now saves the tokenizer alongside the checkpoint.
  * Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
2026-09-11 16:51:23 +05:30
noeyy-mino f10d5183ec launcher: bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq… (#1982)
bump TRT-LLM to 1.3.0rc20, pin vLLM to v0.22.0, fix max_seq_len for
Qwen3.5-4B

### What does this PR do?

Type of change: Bug fix, new feature

- Upgrade TRT-LLM container from 1.3.0rc10 to 1.3.0rc20 across all
  Qwen3-8B, Qwen3-30B-A3B, Kimi-K2.5, and gpt-oss-20b launcher configs.
- Replace Kimi-K2.5 aarch64-specific vLLM image (v0.22.0-aarch64) with
the multi-arch v0.22.0 tag, which resolves to amd64/arm64 automatically.
- Fix Qwen3.5-4B throughput_32k runs: raise max_seq_len from 40960 to
  65536 to accommodate outlier prompts (~46.6K tokens) that caused
  VLLMValidationError and aborted the entire benchmark run.
- Fix specdec_bench_mtp_vllm.yaml: remove reference to non-existent
  runtime_params_throughput_32k.yaml; use --max_seq_len 65536 instead

### Usage

```
cd Model-Optimizer/tools/launcher
uv run launch.py --yaml examples/Qwen/Qwen3-8B/megatron_lm_ptq_local.yaml hf_local=/mnt/hf-local --yes
```

### Testing
N/A


- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Updated launcher example pipelines and the default launcher Slurm
container to the newer TensorRT-LLM `1.3.0rc20` image.
* Updated Kimi-K2.5 workflows to use the multi-architecture vLLM
`0.22.0` image (removing prior architecture-specific variants).
* Increased the Qwen3.5-4B long-context benchmark maximum sequence
length to 65,536 tokens and simplified the corresponding throughput
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2026-07-22 06:57:50 -07:00
noeyy-mino 31f0783e35 add new cases for ckpts on HF (#938)
## What does this PR do?

**Type of change:** new tests

**Overview:** 
Add new deployment tests for newly added checkpoints on HF

## Usage
pytest test_deploy.py --run-release

```python
None
```

## Testing
None

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for new model deployments: Qwen3 variants (480B, 235B,
397B), Kimi K2.5, and additional Mixtral configurations.

* **Improvements**
* Expanded backend support for select models, including TensorRT-LLM
integration for Nemotron models.

* **Chores**
* Renamed environment variable from MODELOPT_LOCAL_MODEL_ROOT to
MODELOPT_LOCAL_EAGLE_MODEL for improved clarity.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2026-03-05 06:55:23 +00:00
noeyy-mino 9de9b8f20c Noeyy/add test cases for the newly added checkpoints on HF (#827)
## What does this PR do?

**Type of change:** new tests

**Overview:** Add new test cases for the newly added checkpoints on
HuggingFace.

## Usage
pytest test_deploy.py --run-release

```python
None
```

## Testing
None

## Before your PR is "*Ready for review*"


- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
None


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for NVFP4 model variants across multiple model families
(DeepSeek, Llama, Qwen, and others).

* **Improvements**
* Enhanced backend availability detection to automatically identify and
manage supported deployment backends at runtime.

* **Tests**
* Improved test infrastructure for better reproducibility and backend
compatibility handling.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2026-02-04 09:27:49 +08:00
noeyy-mino 883c8731aa Noeyy/add new cases for newly added checkpoints on HF (#728)
## What does this PR do?

**Type of change:** Add TRT LLM/vLLM/SGLang functional test cases for
newly added checkpoints on HF

**Overview:** 
1.Since the speculative draft model only supports loading from a local
path, we should set the MODELOPT_LOCAL_MODEL_ROOT environment variable.
If we don't set it, these test cases will be skipped.
2. Newly added checkpoints:

- nvidia/gpt-oss-120b-Eagle3-short-context
- nvidia/gpt-oss-120b-Eagle3-throughput
- nvidia/EAGLE3-NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8

## Usage


```python
pytest tests/examples/llm_ptq/test_deploy.py --run-release
```

## Testing
Run release testing

## Before your PR is "*Ready for review*"
Ready for review

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
N/A

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2025-12-30 10:52:11 +05:30
noeyy-mino c77eebcacc Noeyy/add new ckpts test cases (#650)
## What does this PR do?

**Type of change:** new tests

**Overview:**  Add new checkpoints test cases.

## Usage
pytest test_deploy.py --run-release

```python
# Add a code snippet demonstrating how to use this
```

## Testing
N//A

## Before your PR is "*Ready for review*"


- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
N/A

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2025-12-11 00:22:45 +05:30
noeyy-minoandcoderabbitai[bot] d311edaf5f Add functional test cases for published checkpoints on HF (#455)
## What does this PR do?

**Type of change:** new tests

**Overview:** 
Add vLLM/SGLang/TRT LLM deployment tests for published checkpoints on
HF.
Add e2e test case for gpt-oss

## Usage


```python
# pytest test_deploy.py -k "vllm"
```

## Testing


## Before your PR is "*Ready for review*"


- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added ModelDeployer and ModelDeployerList test utilities with
COMMON_PROMPTS to run multi-backend deployment and inference scenarios.
* Introduced an end-to-end GPT-OSS QAT pipeline test (SFT → QAT → MXFP4
conversion) with optional deployment/benchmarking.
* Added a broad parametric LLM deployment test suite covering many
models/backends with readable test IDs.
* Added automatic Hugging Face cache cleanup after tests to keep runs
isolated.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: neoyy-mino <174223378+noeyy-mino@users.noreply.github.com>
Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-11-14 16:15:38 +05:30
noeyy-mino 970044c02e update model_type of Qwen (#477)
Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2025-11-04 11:32:21 +05:30