### What does this PR do? - Add experimental support for transformers >=5.0 and remove deprecated usages: https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md - ⚠️ For accelerate examples that used `--warmup-ratio: float` (deprecated in 5.x), we now change it to `--warmup-steps: float | int` which works as ratio if float but only for 5.x. For 4.x, it will error out if float and prompt user to change back to `--warmup-ratio` or pass an int absolute step count. - ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints may not work for some models with transformers>=5.0 yet as it requires a lot of fixes (e.g. change in how MoE experts are organized) - ~Add Workaround for TRT-LLM's import of deprecated transformers functions so trt-llm based gpu unit tests work fine. Still deployment for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq example tests still run with transformers 4.57~ - Everything except PTQ and Export (mainly MoE) should work fine with transformers>=5.0 - Bump min torch to 2.8 and enable 2.11 cicd testing - NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3 ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD tests passing - [x] Manually tested unit tests, gpu tests with transformers 4.56 and 5.4 - [x] Manually tested example tests (except trt-llm container tests) with transformers 4.56 and 5.4 - [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540), [example tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643) ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Make remote-code usage opt-in via a configurable --trust_remote_code flag across examples and tools. * **Bug Fixes** * Improve checkpoint/resume detection and related training guidance to avoid erroneous errors. * **Refactor** * Consolidate dtype/config naming, switch warmup settings from ratio → steps, and unify tokenizer invocation patterns. * **Documentation** * Simplify changelog title and add misc notes for release 0.44. * **Chores** * Remove scheduled PR-branch cleanup workflow and relax/remove several transformers version pins. * **Tests** * Adjust test gates, skips, and structures to align with updated deps and behaviors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Table of Contents
- Overview
- Prerequisites
- Accuracy Benchmarks
- Additional Metrics
- API changes in ONNX Runtime GenAI v0.6
- Troubleshoot
Overview
This repository provides scripts, popular third-party benchmarks, and instructions for evaluating the accuracy of Large Language Models (LLMs). It demonstrates how to use a ModelOpt quantized LLM with various established benchmarks, including deployment options using DirectML and TensorRT-LLM in a Windows environment.
Prerequisites
| Category | Details |
|---|---|
| Operating System | Windows 10 or later |
| Python | - For ORT-DML GenAI, use Python 3.11. - For TensorRT-LLM, use Python 3.10. - All other backends are compatible with both Python 3.10 and 3.11. |
| Package Manager | pip |
| Compatible Hardware and Drivers | - Ensure necessary hardware (e.g., CUDA-compatible GPU) and drivers are installed, depending on the evaluation method: - DirectML for DirectML-based evaluation - CUDA for TensorRT |
| Additional Tools | - cmd: Recommended for running the provided commands. - Tar Utility: Included in Windows 10 and later via PowerShell. - Curl: Included in Windows 10 and later via PowerShell. |
Accuracy Benchmarks
MMLU (Massive Multitask Language Understanding)
The MMLU benchmark assesses LLM performance across a wide range of tasks, producing a score between 0 and 1, where a higher score indicates better accuracy. Please refer the MMLU Paper for more details on this.
Setup
The table below lists the setup steps to prepare your environment for evaluating LLMs using the MMLU benchmark.
| Step | Command or Description |
|---|---|
| Open PowerShell as Administrator | - |
| Create and Activate a Virtual Environment (Optional but Recommended) |
python -m venv llm_env .\llm_env\Scripts\Activate.ps1 |
| Install PyTorch and Related Packages | pip install torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0 --index-url https://download.pytorch.org/whl/cu128 |
| Install ONNX Runtime Packages | pip install onnxruntime-directml==1.21.1 pip install onnxruntime-genai-directml==0.6.0 |
| Install Benchmark Requirements | pip install -r requirements.txt |
| Download MMLU Data | mkdir data curl -o .\data\mmlu.tar https://people.eecs.berkeley.edu/~hendrycks/data.tar tar -xf .\data\mmlu.tar -C .\data Move-Item .\data\data .\data\mmlu |
Evaluation Methods
Once the MMLU benchmark is set up, you can use the mmlu_benchmark.py script to evaluate LLMs deployed with various backends. Please refer examples below.
MMLU Benchmark with GenAI APIs for ORT-DML Deployment
To run the model with ORT-DML using GenAI, use the --ep genai_dml argument.
-
Test Suite
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep genai_dml ` --output_file <output_log_file.json> ` --ntrain 5 -
Specific Subjects
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep genai_dml ` --output_file <output_log_file.json> ` --subject abstract_algebra,anatomy,college_mathematics ` --ntrain 5
MMLU Benchmark with ONNX Runtime APIs for DML, CUDA, or CPU Deployment
To run the model with ORT-DML, ORT-CUDA or ORT-CPU execution providers, use --ep ort_dml, --ep ort_cuda, or --ep ort_cpu respectively.
-
Test Suite
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep ort_dml ` --output_file <output_log_file.json> ` --ntrain 5 -
Specific Subjects
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep ort_dml ` --output_file <output_log_file.json> ` --subject abstract_algebra,anatomy,college_mathematics ` --ntrain 5
MMLU Benchmark with Transformer APIs for PyTorch Hugging Face Models
To evaluate the PyTorch Hugging Face (HF) model, use the --ep pt argument.
-
Test Suite
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep pt ` --output_file <output_log_file.json> ` --ntrain 5 ` --dtype <torch_dtype in model's config.json {float16|bfloat16}> -
Specific Subjects
python mmlu_benchmark.py ` --model_name causal ` --model_path <ONNX_model_folder> ` --ep pt ` --output_file <output_log_file.json> ` --subject abstract_algebra,anatomy,college_mathematics ` --ntrain 5 ` --dtype <torch_dtype in model's config.json {float16|bfloat16}>
MMLU Benchmark with TensorRT-LLM APIs for TensorRT-LLM Deployment
-
Install TensorRT-LLM and Compatible PyTorch
pip install torch==2.4.0+cu121 --index-url https://download.pytorch.org/whl pip install tensorrt_llm==0.12.0 ` --extra-index-url https://pypi.nvidia.com ` --extra-index-url https://download.pytorch.org/whl/cu121/torch/ -
Run the Benchmark
-
Test Suite
python mmlu_benchmark.py ` --model_name causal ` --hf_model_dir <hf_model_path> ` --engine_dir <engine_path> ` --ep trt-llm ` --ntrain 5 ` --output_file result.json -
Specific Subjects
python mmlu_benchmark.py ` --model_name causal ` --hf_model_dir <hf_model_path> ` --engine_dir <engine_path> ` --ep trt-llm ` --ntrain 5 ` --output_file result.json ` --subject abstract_algebra,anatomy,college_mathematics
-
Additional Metrics
| Metric | Directory | Description |
|---|---|---|
| KL Divergence | kl_divergence_metrics/ |
Measures output similarity between two models using KL divergence |
| Perplexity | perplexity_metrics/ |
Evaluates language model quality using WikiText-2 perplexity |
| FVD | fvd_metrics/ |
Computes Fréchet Video Distance between two sets of videos using I3D features |
Each sub-directory contains its own README.md with detailed setup and usage instructions.
API changes in ONNX Runtime GenAI v0.6
In onnxruntime-genai (GenAI) v0.6, generator.compute_logits() and generator_params.input_ids are deprecated and new API generator.append_tokens(List: token_ids) is added (see GenAI PR-867 for details).
So, this MMLU script has been updated accordingly - refer following change-snippet from this MMLU script (left works with GenAI < 0.6, right works with GenAI 0.6+). Make sure to update the MMLU script accordingly (left part) for trying it with GenAI < 0.6.
Troubleshoot
-
In case of any model specific issue (e.g. in tokenizer or in onnxruntime-genai package etc.), one can try using older GenAI e.g. export the ONNX model with
onnxruntime-genai-directml0.4 andtransformers4.44. -
In case of trying out MMLU run of ONNX model through GenAI, make sure that the input model is running fine with GenAI. Onnxruntime-genai has example inference scripts (e.g. see phi3 example script).
