mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What does this PR do? **Type of change:** New feature **Overview:** - Implemented the NVFP4QuantExporter - Deprecated fp4qdq_to_2dq - Updated tests ## Usage ```python python torch_quant_to_onnx.py --quantize_mode=nvfp4 \ --onnx_save_path=vit_base_patch16_224.nvfp4.onnx \ --calibration_data_size 64 \ --batch_size 128 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ``` python evaluate.py --onnx_path=vit_base_patch16_224.nvfp4.onnx \ --model_name=vit_base_patch16_224 \ --results_path=./results.txt \ --batch_size 128 ``` Results: ``` The top1 accuracy of the model is 84.39% The top5 accuracy of the model is 97.312% Inference latency of the model is 7.22412 ms ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No - Deprecated fp4qdq_to_2dq - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>