Files
Keval MorabiaandClaude Opus 5 9c1cf80f1b test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do?

Type of change: new tests

Runs the VLM case of `test_qad` under context parallelism, so QAD on a
Qwen3-VL model is covered on the path that until now could not run at
all.

`Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the
batch has to hand it full-length `position_ids`. Megatron-Bridge's
`get_batch` was sharding them too, leaving the rotary embedding at `seq
/ cp**2` against hidden states at `seq / cp`. The fix is upstream in
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243);
this PR is the coverage that would have caught it.

The VLM case moves from tensor to context parallelism and the PTQ step
is sized to the same TP, so QAD still loads a matching checkpoint. The
LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and
`test_distill_vlm` still covers TP for a VLM, so nothing loses coverage.

### Usage

```bash
# Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243
# it now works for VLMs, where it previously died in the rotary embedding.
python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ...
```

### Testing

On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on
`PYTHONPATH`:

- `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD
across 2 CP ranks, export; quantizers survive and the vision tower is
byte-identical. **1 passed (183 s).** Without the upstream fix the same
run dies with `AttributeError: 'NoneType' object has no attribute
'ndim'` in `rope.py:175`.
- `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).**
- Gate check: on today's `nemo:26.08` (no #6243) the probe resolves
`False` and the VLM case stays at `cp_size=1`, byte-identical to current
CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a
1-GPU runner it degenerates to today's config either way.
- `pre-commit run --files ...` clean (ruff check, ruff format, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests extended
rather than new ones added.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test coverage and one doc line; no feature, break, deprecation, or
fix for a released bug.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

- Depends on
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243).
Safe to merge before it lands: the gate keeps the VLM case at
`cp_size=1` until a container ships the fix, at which point the coverage
switches on by itself.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the Qwen3.6 QAD instructions to keep tensor and pipeline
parallelism set to 1, while allowing context parallelism to increase for
longer sequences with the `nemo:26.10` container.
* **Tests**
* QAD validation now selects parallelism settings based on whether the
Megatron-Bridge context-parallel fix is available, and reports when
multi-GPU VLM coverage is reduced.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-01 23:05:55 +05:30
..

Megatron-Bridge Tutorials

End-to-end tutorials that combine ModelOpt optimization techniques on NVIDIA Megatron-Bridge models. Each one walks through a complete workflow using the scripts in examples/megatron_bridge (prune_minitron.py, distill.py, quantize.py, export_quantized_megatron_to_hf.py, export_distilled_megatron_to_hf.py).

Available tutorials

Tutorial What it covers
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 End-to-end optimization of the Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) model: Minitron structured pruning (31.6B/A3.6B → 22B/A3.0B) → two-phase knowledge distillation (100B tokens, 8K then 32K seq length) → quantization → vLLM deployment. Includes data-blend preparation, evaluation setup, and detailed pruning / data-blend / long-context ablations.
Qwen3.6-35B-A3B End-to-end W4A4 NVFP4 optimization of the Qwen3.6-35B-A3B (MoE + hybrid linear/full attention VLM) model: NVFP4 W4A4 post-training quantization → quantization-aware distillation (QAD) to recover the accuracy W4A4 costs → evaluation → vLLM deployment.