Files
Model-Optimizer/examples/gpt-oss/qat-finetune-transformers.ipynb
T
Keval Morabia 04cd596d79 Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do?

- Add experimental support for transformers >=5.0 and remove deprecated
usages:
https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- ⚠️ For accelerate examples that used `--warmup-ratio: float`
(deprecated in 5.x), we now change it to `--warmup-steps: float | int`
which works as ratio if float but only for 5.x. For 4.x, it will error
out if float and prompt user to change back to `--warmup-ratio` or pass
an int absolute step count.
- ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints
may not work for some models with transformers>=5.0 yet as it requires a
lot of fixes (e.g. change in how MoE experts are organized)
- ~Add Workaround for TRT-LLM's import of deprecated transformers
functions so trt-llm based gpu unit tests work fine. Still deployment
for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq
example tests still run with transformers 4.57~
- Everything except PTQ and Export (mainly MoE) should work fine with
transformers>=5.0
- Bump min torch to 2.8 and enable 2.11 cicd testing
- NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3

### Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] CI/CD tests passing
- [x] Manually tested unit tests, gpu tests with transformers 4.56 and
5.4
- [x] Manually tested example tests (except trt-llm container tests)
with transformers 4.56 and 5.4
- [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540),
[example
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Make remote-code usage opt-in via a configurable --trust_remote_code
flag across examples and tools.

* **Bug Fixes**
* Improve checkpoint/resume detection and related training guidance to
avoid erroneous errors.

* **Refactor**
* Consolidate dtype/config naming, switch warmup settings from ratio →
steps, and unify tokenizer invocation patterns.

* **Documentation**
  * Simplify changelog title and add misc notes for release 0.44.

* **Chores**
* Remove scheduled PR-branch cleanup workflow and relax/remove several
transformers version pins.

* **Tests**
* Adjust test gates, skips, and structures to align with updated deps
and behaviors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-09 09:59:37 +05:30

9.5 KiB

Quantization-Aware Fine-Tuning for GPT-OSS

This notebook demonstrates a complete workflow for fine-tuning language models with Quantization-Aware Training (QAT) using modelopt and SFTTrainer for gpt-oss models.

Overview

The workflow includes:

• Model and tokenizer loading

• Dataset preparation

• Training configuration setup

• Model quantization

• Quantization aware training

• Model saving and checkpointing

Setup Environment

In [ ]:
%pip install --upgrade transformers trl
In [ ]:
import modelopt.torch.opt as mto

# Enable automatic save/load of modelopt state huggingface checkpointing
# modelopt state will be saved automatically to "modelopt_state.pth"
mto.enable_huggingface_checkpointing()

Model Configuration

Configure the model parameters including the model path, attention implementation, and data type. Set up the model configuration and prepare the model loading arguments.

In [ ]:
from transformers import AutoConfig, Mxfp4Config
from trl import ModelConfig

model_args = ModelConfig(
    model_name_or_path="openai/gpt-oss-20b",
    attn_implementation="eager",
    torch_dtype="bfloat16",
)
model_kwargs = {
    "revision": model_args.model_revision,
    "trust_remote_code": model_args.trust_remote_code,
    "attn_implementation": model_args.attn_implementation,
    "torch_dtype": model_args.torch_dtype,
    "use_cache": False,
    "device_map": "auto",
}

# Dequantize if the model is in MXFP4 format
config = AutoConfig.from_pretrained(model_args.model_name_or_path)
if (
    getattr(config, "quantization_config", {})
    and config.quantization_config.get("quant_method", None) == "mxfp4"
):
    model_kwargs["quantization_config"] = Mxfp4Config(dequantize=True)

Load the Model and Tokenizer

Load the pre-trained model and tokenizer with the specified configuration.

In [ ]:
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(model_args.model_name_or_path, **model_kwargs)

tokenizer = AutoTokenizer.from_pretrained(
    model_args.model_name_or_path,
)

Dataset Configuration

Set up the dataset parameters for training and evaluation. This includes specifying the dataset name, train/test splits, and test size ratio.

In [ ]:
from trl import ScriptArguments

script_args = ScriptArguments(
    dataset_name="HuggingFaceH4/Multilingual-Thinking",
    dataset_train_split="train",
    dataset_test_split="test",
)
test_size = 0.1

Load and Prepare Dataset

Load the dataset and split it into training and evaluation sets. The dataset is split with the specified test size ratio and random seed for reproducibility.

In [ ]:
from datasets import load_dataset

dataset = load_dataset(script_args.dataset_name)
# split the dataset into train and test
dataset = dataset[script_args.dataset_train_split].train_test_split(test_size=test_size, seed=42)
train_dataset = dataset[script_args.dataset_train_split]
eval_dataset = dataset[script_args.dataset_test_split]

Training Configuration

Configure the training parameters including epochs, batch sizes, learning rate, gradient accumulation, and evaluation strategy. This sets up the SFT configuration for supervised fine-tuning.

In [ ]:
from trl import SFTConfig

training_args = SFTConfig(
    output_dir="gpt-oss-20b-multilingual-reasoner",
    num_train_epochs=0.1,
    learning_rate=2e-5,
    per_device_train_batch_size=1,
    per_device_eval_batch_size=1,
    gradient_accumulation_steps=2,
    max_length=4096,
    warmup_steps=0.03,  # use warmup_ratio if using transformers<5.0
    eval_strategy="steps",
    eval_on_start=True,
    logging_steps=10,
    save_steps=50,
    eval_steps=10,
    save_total_limit=2,
)

Initialize Trainer

Set up the SFT trainer with the model, dataset, and training configuration.

In [ ]:
from trl import SFTTrainer

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset[script_args.dataset_train_split],
    eval_dataset=dataset[script_args.dataset_test_split],
    processing_class=tokenizer,
)

Quantization aware Training

Configure the quantization parameters and prepare the calibration dataset. This step sets up the quantization configuration, creates a calibration subset from the evaluation dataset, and defines a forward loop function for model calibration. The calibration process helps determine optimal quantization scales for the model weights and activations.

In [ ]:
import torch

import modelopt.torch.quantization as mtq

# MXFP4_MLP_WEIGHT_ONLY_CFG doesn't need calibration, but other quantization configurations may require it.
quantization_config = mtq.MXFP4_MLP_WEIGHT_ONLY_CFG
calib_size = 128

dataset = torch.utils.data.Subset(
    trainer.eval_dataset, list(range(min(len(trainer.eval_dataset), calib_size)))
)
data_loader = trainer.get_eval_dataloader(dataset)


def forward_loop(model):
    for data in data_loader:
        model(**data)

Apply quantization to the model using the prepared configuration and calibration data.

In [ ]:
mtq.quantize(model, quantization_config, forward_loop)

Start the quantization-aware training.

In [ ]:
trainer.train()

Model Saving and Checkpointing

Save the trained and quantized model with HuggingFace checkpointing enabled to store the modelopt state automatically.

In [ ]:
model.save_pretrained(training_args.output_dir)