mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Add `examples/alpamayo/qad.py` which distills the quantized Alpamayo
checkpoint's VLM against the FP16 VLM of the original with ModelOpt's
QADTrainer, and shards student and teacher with FSDP2 for multi-GPU
runs. Only the VLM is trained; the action expert stays frozen.
Type of change: new example
<!-- Details about the change. -->
### Usage
```
torchrun --standalone --nproc_per_node 8 qad.py \
--student_ckpt ./alpamayo-auto \
--output_dir ./alpamayo-auto-qad \
--parquet ./train_clips.parquet \
--max_steps 500 --fsdp2 --grad_ckpt --export
```
### Testing
Tested end-to-end on public Alpamayo-1 checkpoint
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added a quantization-aware distillation workflow for AlpamayoR1
models.
- Supports prompt-only or rollout-inclusive distillation, FSDP2
training, dataset slicing, revision pinning, and checkpoint resumption.
- Supports exporting trained models as complete, reloadable AlpamayoR1
checkpoints.
- Added optional vision-parameter freezing, trajectory-history fusion,
gradient checkpointing, evaluation, and synchronized training cadence.
- **Documentation**
- Expanded the Alpamayo guide with setup, training, dataset, and export
instructions.
- Clarified sensitivity-based quantization behavior.
- Added a version 0.47 quantization changelog entry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>