mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new tests Runs the VLM case of `test_qad` under context parallelism, so QAD on a Qwen3-VL model is covered on the path that until now could not run at all. `Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the batch has to hand it full-length `position_ids`. Megatron-Bridge's `get_batch` was sharding them too, leaving the rotary embedding at `seq / cp**2` against hidden states at `seq / cp`. The fix is upstream in [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243); this PR is the coverage that would have caught it. The VLM case moves from tensor to context parallelism and the PTQ step is sized to the same TP, so QAD still loads a matching checkpoint. The LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and `test_distill_vlm` still covers TP for a VLM, so nothing loses coverage. ### Usage ```bash # Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243 # it now works for VLMs, where it previously died in the rotary embedding. python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ... ``` ### Testing On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on `PYTHONPATH`: - `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD across 2 CP ranks, export; quantizers survive and the vision tower is byte-identical. **1 passed (183 s).** Without the upstream fix the same run dies with `AttributeError: 'NoneType' object has no attribute 'ndim'` in `rope.py:175`. - `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).** - Gate check: on today's `nemo:26.08` (no #6243) the probe resolves `False` and the VLM case stays at `cp_size=1`, byte-identical to current CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a 1-GPU runner it degenerates to today's config either way. - `pre-commit run --files ...` clean (ruff check, ruff format, mypy, bandit, markdownlint). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — existing tests extended rather than new ones added. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test coverage and one doc line; no feature, break, deprecation, or fix for a released bug. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information - Depends on [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243). Safe to merge before it lands: the gate keeps the VLM case at `cp_size=1` until a container ships the fix, at which point the coverage switches on by itself. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Qwen3.6 QAD instructions to keep tensor and pipeline parallelism set to 1, while allowing context parallelism to increase for longer sequences with the `nemo:26.10` container. * **Tests** * QAD validation now selects parallelism settings based on whether the Megatron-Bridge context-parallel fix is available, and reports when multi-GPU VLM coverage is reduced. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Megatron-Bridge Tutorials
End-to-end tutorials that combine ModelOpt optimization techniques on NVIDIA Megatron-Bridge models.
Each one walks through a complete workflow using the scripts in examples/megatron_bridge (prune_minitron.py, distill.py, quantize.py, export_quantized_megatron_to_hf.py, export_distilled_megatron_to_hf.py).
Available tutorials
| Tutorial | What it covers |
|---|---|
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | End-to-end optimization of the Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) model: Minitron structured pruning (31.6B/A3.6B → 22B/A3.0B) → two-phase knowledge distillation (100B tokens, 8K then 32K seq length) → quantization → vLLM deployment. Includes data-blend preparation, evaluation setup, and detailed pruning / data-blend / long-context ablations. |
| Qwen3.6-35B-A3B | End-to-end W4A4 NVFP4 optimization of the Qwen3.6-35B-A3B (MoE + hybrid linear/full attention VLM) model: NVFP4 W4A4 post-training quantization → quantization-aware distillation (QAD) to recover the accuracy W4A4 costs → evaluation → vLLM deployment. |