### What does this PR do? - Add experimental support for transformers >=5.0 and remove deprecated usages: https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md - ⚠️ For accelerate examples that used `--warmup-ratio: float` (deprecated in 5.x), we now change it to `--warmup-steps: float | int` which works as ratio if float but only for 5.x. For 4.x, it will error out if float and prompt user to change back to `--warmup-ratio` or pass an int absolute step count. - ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints may not work for some models with transformers>=5.0 yet as it requires a lot of fixes (e.g. change in how MoE experts are organized) - ~Add Workaround for TRT-LLM's import of deprecated transformers functions so trt-llm based gpu unit tests work fine. Still deployment for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq example tests still run with transformers 4.57~ - Everything except PTQ and Export (mainly MoE) should work fine with transformers>=5.0 - Bump min torch to 2.8 and enable 2.11 cicd testing - NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3 ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD tests passing - [x] Manually tested unit tests, gpu tests with transformers 4.56 and 5.4 - [x] Manually tested example tests (except trt-llm container tests) with transformers 4.56 and 5.4 - [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540), [example tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Make remote-code usage opt-in via a configurable --trust_remote_code flag across examples and tools. * **Bug Fixes** * Improve checkpoint/resume detection and related training guidance to avoid erroneous errors. * **Refactor** * Consolidate dtype/config naming, switch warmup settings from ratio → steps, and unify tokenizer invocation patterns. * **Documentation** * Simplify changelog title and add misc notes for release 0.44. * **Chores** * Remove scheduled PR-branch cleanup workflow and relax/remove several transformers version pins. * **Tests** * Adjust test gates, skips, and structures to align with updated deps and behaviors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
KL Divergence Model Validation Toolkit
This toolkit provides comprehensive model validation capabilities using KL divergence metrics to compare two models. It's designed to evaluate the similarity between model outputs across different optimization techniques, frameworks, and hardware backends.
Overview
The toolkit measures output similarity between models using KL (Kullback-Leibler) divergence, which quantifies how one probability distribution differs from another. Lower KL divergence values indicate more similar model outputs.
Primary Use Cases:
- Model Optimization Validation - Verify that optimized models (quantization, pruning) maintain output quality
- Framework Comparison - Compare Hugging Face models vs ONNX Runtime GenAI models
- Precision Analysis - Evaluate FP16 vs INT4 vs INT8 model outputs
- Execution Provider Testing - Test different EP implementations (CUDA, DirectML, CPU, TensorRT)
Key Components
Main Script
| Script | Purpose | Comparison Modes |
|---|---|---|
compute_kl_divergence.py |
Two-model sequential comparison | • HF vs GenAI • GenAI vs GenAI (same EP) • GenAI vs HF • HF vs HF |
Datasets Used
- Wikitext-2 test split for consistent evaluation across all models
- Automatic dataset loading and preprocessing via HuggingFace datasets
Installation
1. Install Base Requirements
pip install -r requirements.txt
Note: Install torch with CUDA for faster inference: "pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu129"
2. Install ONNX Runtime GenAI Package
Install one of the following based on your hardware:
# For CUDA
pip install onnxruntime-genai-cuda
# For DirectML support
pip install onnxruntime-genai-directml
# For CPU
pip install onnxruntime-genai
Usage Examples
Quick Start
Compare HF vs GenAI Model
python compute_kl_divergence.py \
--model1 "meta-llama/Llama-3.1-8B-Instruct" --model1_type hf \
--model2 "G:\models\genai_model" --model2_type genai \
--device cuda \
--output results.json
Compare Two GenAI Models (Same EP)
python compute_kl_divergence.py \
--model1 "G:\models\genai_fp16" --model1_type genai \
--model2 "G:\models\genai_int4" --model2_type genai \
--output fp16_vs_int4.json
Advanced Options
Enable Debug Output
python compute_kl_divergence.py \
--model1 "meta-llama/Llama-3.1-8B-Instruct" --model1_type hf \
--model2 "G:\models\genai_model" --model2_type genai \
--device cuda \
--output results.json \
--debug # Enables verbose logging
Configuration Parameters
compute_kl_divergence.py
Required Parameters:
| Parameter | Description | Values |
|---|---|---|
--model1 |
Path to first model | Local path or HF Hub identifier |
--model1_type |
Type of first model | hf, genai |
--model2 |
Path to second model | Local path or HF Hub identifier |
--model2_type |
Type of second model | hf, genai |
Optional Parameters:
| Parameter | Description | Default |
|---|---|---|
--device |
Device for HF model inference | cuda |
--output |
Output JSON file path | None (prints to console) |
--debug |
Enable verbose debug output | False |
Model Path Formats:
- HF models:
- Hub identifier:
meta-llama/Llama-3.1-8B-Instruct - Local path:
F:\shared\Llama-3.1-8B-Instruct
- Hub identifier:
- GenAI models:
- Local path only:
G:\models\genai_model
- Local path only:
Key Insights
- Lower is better: Smaller KL divergence = more similar outputs
- Relative comparison: Compare against baseline (e.g., HF FP32)
Troubleshooting
Common Issues and Solutions
1. CUDA Out of Memory
Error:
RuntimeError: CUDA out of memory
Solutions:
- Use CPU for HF model:
--device cpu - Close other applications using GPU
- Try smaller batch size (modify code if needed)
- Ensure only one model loads at a time (script should handle this)
2. Execution Provider Mismatch
Error:
[INFO] Comparing two GenAI models (same execution provider)
Note: This is informational. GenAI vs GenAI comparisons require same EP.
Solution: Ensure both models were created for the same execution provider.