mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
b16356776b8dd329059985dea6fc98d26bf07d94
36
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d73278808b |
Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do? Type of change: Bug fix Bumps the Megatron-Bridge examples, tests and launcher configs to `nemo:26.08` and removes the version-gated fallbacks they carried, plus the fixes needed to make the suites green on that container. **26.08 bump and shim removal** - Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher configs move to `nemo:26.08`. - `examples/megatron_bridge/_distillation_provider.py` is deleted — 26.08's Megatron-Bridge ships `convert_to_distillation_provider(..., distill_submodule=...)` natively, so `distill.py` imports it directly. - `prune_minitron.py` drops the `AutoBridge.from_hf_config` / config-only-export probing; `--no_moe_grouped_gemm` is no longer needed in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native MoE expert mappings are in 26.08). - `_DynamicMambaMixer` targets only the raw `conv1d_weight` / `conv1d_bias` parameters that replaced the `conv1d` module in Megatron-Core. **MambaModel / MambaModelProvider removal** Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is a deprecated subclass that shares its `forward`, so `DMRegistry` resolves those instances to the `HybridModel` registration and the separate entry is redundant. Same for `MambaModelProvider` vs `HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` / `ExtendedRMSNorm` are untouched — the layers still exist. The deprecated `get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`. **Bug fix: compressed output_layer extra state** `mtq.compress` converts even a *disabled* `output_layer` into a `RealQuantLinear` (its weight is left uncompressed, since `pack_real_quantize_weight` skips disabled quantizers). The guard added in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra state and every worker died in `GPTModel.sharded_state_dict`: ``` RuntimeError: Boolean value of Tensor with more than one value is ambiguous megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict output_extra_state and output_extra_state.data ``` The guard now keys off whether the weight was actually compressed (`QTensorWrapper`) instead of the class. This took out all 12 `test_homogeneous_compressed_sharded_state_dict` params, and the crashed workers poisoned the pool, which surfaced as unrelated timeouts and NCCL errors in `test_layer_sync_moe_local_experts_amax`, `test_kv_cache_quant`, `test_kv_cache_amax_sync`, `test_convert_mcore_te_gpt_model` and `test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it; `test_output_layer_extra_state_empty_when_nothing_quantized` now asserts the contract directly and is not skipped. **Checkpoint import entry point** 26.08 replaced `examples/conversion/convert_checkpoints.py` with `scripts/conversion/convert.sh`, so `tools/launcher/common/megatron_bridge/import/import.sh` and the three README snippets are retargeted. `import.sh` uses the distributed GPU backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs. **Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still dispatches between the `conv1d` module (26.06 and earlier) and the raw parameters (26.08+), so `import_mcore_gpt_from_hf` / `export_mcore_gpt_to_hf` handle NemotronH on both. Only the Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models require 26.08. **Test consolidation** `test_export_distilled_megatron_to_hf.py` is merged into `test_distill.py`: `test_distill_llm` becomes `test_distill_llm_hf_export` and covers the standalone `--export_iterations all` run on the checkpoints it already produces, saving one full distillation (~185 s of CI time). The two mamba-named gpu test files are renamed to `hybrid`. ### Usage ```bash # HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \ --executor local \ --device gpu \ --gpus-per-node 8 \ --hf-model Qwen/Qwen3-8B \ --megatron-path /tmp/Qwen3-8B-megatron ``` ### Testing All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout overrides: - `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since `qwen3_5_moe_vl` covers the VLM QAD path. - `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`, `peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed. - `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch change: 27 passed. - The 21 previously failing/hanging quantization tests: 21 passed (12 + 9). - `tests/gpu_megatron/torch/{nas,prune}`: verified separately. `import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight tensors after flattening each dist checkpoint with `dcp_to_torch_save` — the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh` end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device cpu`. `nemo:26.06` compatibility was checked directly in that image: `megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and `hybrid_layer_pattern` are all present, while `megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are not. The NemotronH round-trip test failed there before the conv1d dispatch was restored and the dispatch is back in place; per project convention the suites themselves only run on 26.08. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus Minitron pruning of Mamba/hybrid models now require `nemo:26.08`. Megatron-LM quantization and checkpoint export still run on `nemo:26.06`. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — `test_output_layer_extra_state_empty_when_nothing_quantized` for the compress fix; existing tests extended for the merged export coverage. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for importing Hugging Face checkpoints into Megatron distributed format. * Expanded distillation workflows to export selected or all checkpoint iterations. * **Improvements** * Expanded Hybrid model support across Megatron workflows. * Updated distributed import tooling with GPU and parallelism options. * Updated supported environments and examples to NVIDIA NeMo 26.08. * **Bug Fixes** * Corrected output-layer quantization state handling when quantization is disabled. * **Documentation** * Added compatibility guidance for current and legacy NeMo containers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2fc352be2d |
Add VLM pruning and PTQ with image-text calibration (Megatron-Bridge) (#1792)
### What does this PR do?
Type of change: New feature
Adds **vision-language model (VLM) support** to the Megatron-Bridge
examples for both **Minitron pruning** (`prune_minitron.py`) and **PTQ**
(`quantize.py`). Only the **language model** is pruned/quantized — the
vision tower and vision→language projector are left in full precision —
and the full VLM is saved back. `hidden_size` is skipped for pruning
when it is shared with the vision→LM projector.
Supported VLMs (tested e2e): **Qwen{3,3.5}-VL** (dense; hybrid
GatedDeltaNet + gated attention) and **Gemma3-VL** (sliding/full
attention).
### Calibration (image-text)
Calibration is conditioned on real **image-text** data so the language
model's pruning importance / quantizer statistics see vision-conditioned
activations. The modality is inferred from `--calib_dataset_name`:
- an **image-text** dataset (default for VLMs,
`nemotron_vlm_dataset_v2`) drives the **full VLM forward**;
- a **text** dataset runs text-only calibration of the language model
(for text-vs-image ablations).
A shared `get_megatron_vlm_calibration_forward_loop` (built on
`megatron_prefill`) drives the full VLM forward over image-text pairs
from `vlm_dataset_utils` (`scienceqa`, `nemotron_vlm_dataset_v2`, with
config-driven subset/shard caps to bound downloads). It shards across
**data-parallel (DP)** ranks like the text loop (#1804); **context
parallelism (CP)** applies to text-only VLM calibration (the shared text
loop), not the multimodal forward — splitting the sequence would
misalign the merged vision embeddings.
### Results - Cosmos-Reason2-2B
Validated end-to-end on **Cosmos-Reason2-2B** (Qwen3-VL). Minitron NAS
prunes the language-model tower **1.72B → ~1.59B** (vision encoder +
projector frozen), top_k=1. Calibration data drives pruning importance;
image-text calibration runs the full VLM forward.
| Model | Calibration | MMLU | BLINK Rel-Depth | RealWorldQA |
|---|---|---|---|---|
| Baseline (1.72B) | — | 0.58 | 0.76 | 0.61 |
| Pruned (1.59B) | text (`nemotron-post-training-dataset-v2`) | 0.51\* |
~0.69 | ~0.57 |
| Pruned (1.59B) | image+text (`nemotron_vlm_dataset_v2`) | 0.49\* |
**0.77** | **0.61** |
\* Pruned MMLU on the 10% split (the pruning score function); baseline
MMLU is the full set. The VLM-benchmark numbers for the text row were
measured with a different text calibration set and are expected to be
similar for `nemotron-post-training-dataset-v2` (marked `~`).
> [!NOTE]
> These numbers come from short single runs on small eval splits — read
them for **high-level trends only**, not as exact values.
Takeaways: pruning the LM tower of a VLM works end-to-end. **Image-text
calibration** (this PR's feature) preserves the VLM benchmarks better
than text-only — BLINK Rel-Depth ~0.77 vs ~0.69 and RealWorldQA ~0.61 vs
~0.57, both close to the unpruned baseline (0.76 / 0.61) — which is the
motivation for calibrating on vision-conditioned activations.
### Results - Qwen3.5-9B
| Model | MMLU | MMStar |
|----------------------------|:------:|:------:|
| Qwen3.5-9B | 0.7003 | 0.6117 |
| Pruned-7B (text calib) | 0.5527 | 0.4411 |
| Pruned-7B (image+text calib) | 0.5107 | 0.3941 |
### Key changes
- `quantize.py`: quantizes the **root** model with non-LM (vision)
quantizers disabled, so the ModelOpt state lives on the root (required
by the Megatron save) while only the language model is quantized.
- `prune_minitron.py`: image-text (or text) calibration for VLM pruning
importance.
- Shared VLM calibration forward loop (`megatron_prefill`-based, unwraps
tuple outputs, DP-sharded) + `vlm_dataset_utils`.
- Tiny VLM test fixtures (Qwen3.5-VL, Gemma3-VL) with vision tokens
derived dynamically from the reference processor; VLM prune + quantize
example tests.
- README + CHANGELOG.
### Usage
```bash
# Prune the language model of a VLM (image-text calibration by default)
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path <vlm> \
--prune_target_params 3e9 \
--output_hf_path /tmp/vlm-pruned
# PTQ the language model of a VLM
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path <vlm> \
--quant_cfg fp8 \
--export_megatron_path /tmp/vlm-fp8-megatron
```
### Testing
- `test_prune_minitron.py::test_prune_minitron_vlm` — Gemma3-VL,
image-text (ScienceQA) calibration; full load → prune (depth + ffn) →
save → reload.
- `test_quantize_export.py::test_quantize_vlm` — Qwen3.5-VL, text
calibration; quantize LM → save Megatron checkpoint.
- LM regression tests (`test_prune_minitron`,
`test_quantize_and_export`) unchanged and passing.
### Not in scope
- **HF unified export of a quantized VLM** is not yet supported;
`export.py` saves the Megatron checkpoint only for VLMs (tracked by a
TODO in `export.py`). The recommended path is to route the megatron→HF
quant export through Megatron-Bridge's
`AutoBridge.export_hf_weights_quant(quantization_checker, quant_fn,
quant_block_size)`, which reuses the bridge's per-model mcore↔HF mapping
— covering Qwen3.5-VL / Gemma3-VL and the vision tower/projector (left
full precision) for free — so modelopt supplies only the checker +
pack/scale fn + `hf_quant_config` (KV-cache scales need a separate
path). This avoids re-authoring per-model mappings in modelopt (cf.
#1482's Qwen3-VL-only `mcore_qwen3vl.py`).
> [!NOTE]
> Qwen3.5-VL **MoE** is not tested e2e: the Megatron-Bridge weight
conversion expects packed (`gate_up_proj`) experts that transformers'
tiny checkpoint doesn't emit. MoE pruning itself is covered by
`test_mcore_qwen35_gdn_moe_pruning`.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
### Additional Information
Follow-up to the GatedDeltaNet/MLA/latent-MoE pruning PR (#1747).
Rebased on `main` to pick up CP/DP calibration (#1804); the VLM
calibration loop now shards across DP ranks the same way. `hidden_size`
pruning for VLMs (requires resizing the vision projector) is left for a
future PR.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added VLM-aware Minitron pruning and post-training quantization that
target only the language-model portion, keeping the vision
tower/projector in full precision.
* Calibration now auto-selects text vs image-text datasets based on
model type, with modality validation.
* Expanded Megatron-Core CP/DP guidance and introduced a `--cp_size`
flag in quantization examples.
* **Bug Fixes**
* Improved VLM generation/prefill output handling and made vocabulary
sizing more robust for VLM wrappers.
* **Tests / Documentation**
* Updated pruning/quantization docs and refreshed/added VLM-focused
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b6bf6b7997 |
Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
87f1a4f496 |
Add Minitron hidden_size pruning support for GatedDeltaNet, MLA, and latent MoE (#1747)
### What does this PR do?
Type of change: New feature
Adds Minitron (`mcore_minitron`) pruning support for Megatron-Core
models with attention and MoE variants that were previously unsupported.
For all of these, only ``hidden_size`` is pruned (alongside the usual
`ffn_hidden_size` / `num_layers` / MoE dimensions); the variant-internal
dimensions are kept:
- **GatedDeltaNet** (linear attention) and **gated attention**
(`attention_output_gate`) — e.g. Qwen3.5 (hybrid GatedDeltaNet +
gated-attention), including MoE variants. Attention / linear-attention
heads are not pruned.
- **Multi-Latent Attention (MLA)** — e.g. DeepSeek. MLA latent ranks are
not pruned.
- **Latent MoE** (`moe_latent_size`) — e.g. Nemotron-3. `hidden_size`
pruning resizes the `fc1`/`fc2` latent projections while the experts
stay in the (static) latent dim.
Introduces a `_DynamicAttention` base class whose policy decides whether
only `hidden_size` is pruned or attention heads too —
`_DynamicSelfAttention`, `_DynamicGatedDeltaNet` and
`_DynamicMLASelfAttention` build on it (extensible for future attention
types). A TODO documents how to also prune `moe_latent_size` as a
follow-up.
Note that the new `megatron.core` imports are not guarded because they
are available in all containers we support (e.g. `nemo:26.02` onwards)
### Usage
```python
import modelopt.torch.prune as mtp
# Qwen3.5 (GatedDeltaNet), DeepSeek (MLA), or Nemotron-3 (latent MoE) GPTModel
model, _ = mtp.prune(
model,
mode="mcore_minitron",
constraints={"export_config": {"hidden_size": 2048, "ffn_hidden_size": 8192}},
dummy_input=None,
config={"forward_loop": forward_loop},
)
```
### Testing
New GPU tests in
`tests/gpu_megatron/.../test_mcore_gpt_minitron_pruning.py` —
`test_mcore_qwen35_gdn_moe_pruning`, `test_mcore_mla_pruning`,
`test_mcore_latent_moe_pruning` — run across available GPUs (covers
pipeline parallel). Existing GPT/MoE prune + parameter-sorting tests
still pass; the file's tests were refactored to share scaffolding
(`_build_and_prune_variant`, `_sort_and_capture`, `_make_forward_loop`,
`_assert_reprune_matches`). `pre-commit` (ruff, mypy) clean.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Expanded Minitron pruning for Megatron-Core to support GatedDeltaNet
(including gated-attention MoE variants) and Multi-Latent Attention
(MLA).
* Added optional latent MoE pruning behavior: pruning targets
hidden-size projections while keeping the MoE latent dimension
unchanged.
* **Documentation**
* Updated the pruning support matrix to clarify attention types that
don’t support `num_attention_heads` pruning.
* **Bug Fixes**
* Improved attention importance/activation capture so pruning hooks only
run when attention-head pruning is actually enabled.
* **Tests**
* Added/expanded GPU coverage for gated-delta-net MoE, MLA, and latent
MoE pruning, with shape/config and inference checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
55a2101e2f |
Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do? Type of change: documentation + minor example-script tweaks Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16 tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) results using the **new shared calibration loop (sequence packing)** and to **fix tool calling in evaluation**. - Refreshed the prune → distill → eval → **FP8** results (accuracy + vLLM throughput tables) with the new calibration loop. - **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now run the Python sandbox tool. The tutorial reports both **with-tools** and **no-tools** GPQA/AIME and shows `mean ± std_dev`. - Script tweaks: `quantize.py` calibration now uses sequence packing (`pack=True`) which leads to slight improvement in PTQ; `prune_minitron.py` defaults `inference_batch_size` to `calib_batch_size`. ### Testing Documentation + small example-script changes; tutorial relative links resolve and the results tables / figure were verified consistent. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the main guide and evaluator instructions for prune + distill + FP8/NVFP4 quantization, including refreshed vLLM deployment tips, benchmark/noise presentation, and long-context tool-calling attribution notes. * Refreshed README technique examples/links, reordered the model support matrix rows, and improved pruning overview/support-matrix text. * **Changes to Examples** * NAS pruning now documents higher GPU memory usage vs manual pruning; pruning batching defaults were improved. * Quantization PTQ calibration uses packed document packing; quantized checkpoint export messaging was streamlined. * Updated pruning/distillation/quantization tutorial guidance, metrics/tables, command parameters, and evaluator YAML settings (KV-cache dtype, generation defaults, task behavior). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5584ce4558 |
Migrate Nemotron-3-Nano tutorial PTQ to MBridge scripts and move under examples/megatron_bridge (#1601)
### What does this PR do? Type of change: documentation (+ minor test fixes) Migrates the Nemotron-3-Nano-30B-A3B-BF16 tutorial quantization step from `examples/llm_ptq/hf_ptq.py` to the Megatron-Bridge quantize + export, and relocates the tutorial next to the scripts it now uses. Now that the whole tutorial is Megatron-Bridge based, it lives under `examples/megatron_bridge/`. - **Quantization migration:** replace the single `hf_ptq.py` call with `examples/megatron_bridge/quantize.py` (calibrate + save a Megatron checkpoint) → `examples/megatron_bridge/export.py` (deployable unified HF checkpoint). The FP8 results table is refreshed with the `quantize.py` numbers (same defaults, slightly better on average). - **Relocation:** moved `examples/pruning/minitron/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/` → `examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`. A **redirect-stub `README.md`** remains at the old path (a directory symlink isn't traversable in the GitHub web UI), and all in-repo references (root README, CHANGELOG, pruning READMEs, megatron_bridge README) plus the tutorial's own relative links are updated. - **Evaluation:** per-format vLLM benchmark commands (BF16 / FP8), FP8 deployment notes documented in `nemo_evaluator.yaml`, reduced LiveCodeBench/AIME `num_repeats` (were too slow), and bumped the `nemo-evaluator-launcher` pin. - **Misc:** drop the `examples/megatron_bridge/requirements.txt` `transformers<5` pin in favor of an inline "downgrade `transformers<5` to save pruned Nemotron checkpoints" note; guard the hybrid Mamba-MoE sharded-state-dict test behind `HAS_MAMBA` (requires `mamba_ssm`); shrink the tiny Gemma3 test fixture's attention heads. > **Note:** the **NVFP4 + QAD** experiments (formerly the focus of this PR) are split out — their accuracy/throughput results are still in progress — and will follow in a separate PR on top of this one. ### Testing Docs-only + test-guard changes. Pre-commit hooks (markdownlint, RST checks, ruff, mypy) pass. The tutorial's relative links and the old-path redirect stub were verified to resolve to real files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (old tutorial path still resolves via a redirect-stub README; `quantize.py`/`export.py` already exist in `examples/megatron_bridge`) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (adjusts/guards existing tests only) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (existing tutorial entry updated to the new path) - Did you get Claude approval on this PR?: ✅ ### Additional Information Supersedes the previous "Part 3 of 4 (NVFP4 + QAD docs)" scope of this PR; the NVFP4 + QAD tutorial additions will land in a follow-up. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Moved the Nemotron-3-Nano-30B-A3B tutorial into the Megatron-Bridge tutorials and replaced the old file with a pointer to the new location. * Updated vLLM throughput numbers to 2.6× and expanded results/throughput tables. * Reworked the FP8 quantization/export workflow and added a note to use transformers<5 when saving pruned models. * Added a tutorials index and adjusted evaluator launcher pin and repeat counts. * **Tests** * Tests now detect optional Mamba support and skip related tests when unavailable. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
999c99913e |
Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment (#1376)
### What does this PR do? Type of change: example/tutorial <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment <img width="2079" height="1613" alt="image" src="https://github.com/user-attachments/assets/19b6ab82-7f01-45df-a0a5-d1c3282b384a" /> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * End-to-end Nemotron-3-Nano-30B tutorial: pruning, two‑phase distillation, FP8 PTQ, evaluation, and vLLM deployment; new ablations and long‑context analyses. * Distillation CLI: configurable seed and activation‑recomputation options. * **Bug Fixes** * Preprocessing hardened to skip malformed JSONL and normalize tool‑call argument formats. * **Documentation** * Many README/examples/evaluator docs updated (news list, tokenization guides, tutorials, configs, and deployment notes). * **Tests** * Added test verifying preprocessing handles stringified tool‑call arguments. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1376?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
5d0441ae3d |
Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary
Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.
The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.
Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)
Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.
Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.
## Experimental results
Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.
**Three calibration data shapes compared:**
- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.
| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |
### Key findings
- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
2ce745a92e |
Deprecate gradnas pruning and bert example (#1427)
### What does this PR do? Type of change: Deprecation of dead code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Deprecation warning already added in 0.44 as per 1-release deprecation policy GradNAS only works for Bert and GPT-J and we dont actively maintain it or test it. Keeping it creates an expectation that it works plus it adds one more option for user to choose from. We already have much better pruning algorithms (Minitron and Puzzletron) for LLM pruning already hence removing GradNas. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ No but we dont have any users of this feature either <!--- If ❌, explain why. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecation** * GradNAS pruning algorithm deprecated; related examples removed. * **Documentation** * Pruning and NAS guides and changelog updated to focus on Minitron and FastNAS; GradNAS references removed. * **Chores** * Chained-optimizations example and scripts removed. * Ownership mappings updated for README and examples; license insertion now applies to a previously excluded example file. * **Tests** * Multiple unit tests and test utilities related to GradNAS/transformer NAS removed. [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1427) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> |
||
|
|
f0eaa198df |
Enable active-param and memory based Minitron pruning constraint (#1377)
### What does this PR do?
Type of change: New feature, new tests, documentation.
OMNIML-4108: Extends the Minitron NAS pruner to support pruning by
**active parameter count** (`active_params`) and **memory footprint**
(`memory_mb`) in addition to the existing total parameter count
(`params`) constraint. Also adds standalone utilities for analytical
model stats.
#### Changes
**New pruning constraint keys**
- `active_params`: prune to a target number of active (routed) params —
useful for MoE models where total ≫ active; when present,
`active_params` is the **primary sort/display metric** for candidates
(priority: `active_params` > `params` > `memory_mb`)
- `memory_mb`: prune to fit a memory budget (BF16 weights + KV-cache +
Mamba state at a given sequence length and batch size)
- Constraints can be combined (AND logic): e.g. `{"params": 6e9,
"memory_mb": 12288}`
**New standalone utilities**
(`modelopt.torch.nas.plugins.megatron_model_stats`)
- `mcore_param_count`: analytically computes total and active parameter
counts for GPT and Mamba/hybrid MCore models
- `mcore_memory_footprint_mb`: estimates memory in MB (weights +
KV-cache + Mamba state)
- `print_mcore_model_stats`: rich-formatted model stats panel
**Rich-formatted pruning logs** — search space, top-k candidate tables,
and best subnet panel printed on rank 0
**`prune_score_func` format update** — now `mmlu_<N>pct_bs<bs>` (e.g.
`mmlu_10pct_bs32`) to explicitly control batch size for MMLU evaluation;
old `mmlu_<N>pct` format removed
**Infrastructure**
- NeMo container bumped to `nvcr.io/nvidia/nemo:26.04` in CI and docs
- Added `examples/megatron_bridge/requirements.txt` with
`transformers<5.0` (required for saving some Nemotron-3-Nano models)
### Usage
```python
# Prune to 3B active params (MoE-aware) — active_params is the primary sort metric
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"active_params": 3e9}, config=pruning_config)
# Prune to fit a 12 GB memory budget
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"memory_mb": 12288}, config=pruning_config)
```
### Testing
Pruned Nemotron-3-Nano-30B-A3B (31.6B, A3.6B) --> A3.0B. Takes <1hr on
8x H100 (more details in #1376)
```bash
torchrun --nproc_per_node 8 examples/megatron_bridge/prune_minitron.py \
--pp_size 8 \
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--trust_remote_code \
--prune_target_params 28e9 \
--prune_target_active_params 3e9 \
--hparams_to_skip num_attention_heads \
--seq_length 8192 \
--output_hf_path pruned/Nemotron-3-Nano-30B-A3B-Pruned-28B-A3B-top20-max15depth-max30width-mmlu_10pct_bs32 \
--top_k 20 \
--max_depth_pruning 0.15 \
--max_width_pruning 0.30 \
--prune_score_func mmlu_10pct_bs32 \
--num_layers_in_first_pipeline_stage 5 \
--num_layers_in_last_pipeline_stage 5
```
```
╭──────────────────────────────────────────────────── Original Model Stats ─────────────────────────────────────────────────────╮
│ Total Parameters 31.58B │
│ Active Parameters 3.58B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 60230.1 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 60301.9 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
Top 20 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 120, │ 3.00B │ 27.06B │ 0.3399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 2 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.4650 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 3 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2343 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.3762 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 7 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.4783 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 8 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2420 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 9 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2399 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 10 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112, │ 3.00B │ 26.17B │ 0.2601 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 11 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 112, │ 3.00B │ 25.37B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 12 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.4329 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 13 │ {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 128, │ 3.00B │ 26.17B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 2816} │ │ │ │
│ 14 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2336 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 15 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2559 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 16 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.4608 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
│ 17 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2455 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 18 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104, │ 3.00B │ 24.42B │ 0.2503 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} │ │ │ │
│ 19 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 120, │ 3.00B │ 27.92B │ 0.2587 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} │ │ │ │
│ 20 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.2469 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072} │ │ │ │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘
╭──────────────────────────────────────────────────────────────────────── Best Subnet ─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.4783 │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭───────────────────────────────────────────────────── Pruned Model Stats ──────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=1) weights: 42489.7 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 42561.6 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
bb08094ff1 |
Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how ModelOpt techniques stack: - **Pruning** — Minitron structured pruning 9B → 7B - **Distillation** — Megatron-Bridge knowledge distillation up to 80B tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math 500, IFEval, SciCode - **Evaluation** - using nemo-evaluator - **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from \`config.json\`) - **vLLM Throughput** — BF16 vs FP8 benchmark on single H100 <img width="2085" height="1740" alt="image" src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87" /> <img width="2085" height="810" alt="image" src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843" /> ### Files changed - `examples/pruning/minitron/README.md` — index page for Minitron end-to-end tutorials - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` — full repro doc with 6 sections: data prep, pruning, distillation, evaluation, FP8 quantization, vLLM benchmarking - `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml` — NeMo Evaluator config used for all benchmark numbers - `examples/pruning/puzzletron/README.md` — index page for Puzzletron distillation results - `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron distillation results (renamed from puzzletron.md) - `examples/pruning/README.md` — updated Results section with direct links to new locations - `examples/megatron_bridge/README.md` — updated results link to point to `examples/pruning/` - `examples/puzzletron/README.md` — updated distillation results link - `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for all datasets used in the data blend 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Documentation * **New end-to-end tutorial** for model optimization covering Minitron pruning, knowledge distillation, FP8 quantization, and vLLM deployment with reproducibility steps and benchmark results * **Dataset preparation guide** with ready-to-run tokenization templates for Nemotron HuggingFace datasets * **Evaluation configuration** and results documentation including ablation studies across multiple benchmarks * **Updated navigation** across pruning, distillation, and dataset examples to streamline user workflows <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
e4b054bf32 |
Fix and Speedup megatron_mmlu by >10x via prefill scoring and global batching (#1280)
### What does this PR do?
Type of change: new feature + bug fix
Two improvements to Megatron inference utilities:
**1. Pipeline Parallel (PP) correctness fixes**
PP inference was producing garbage output (MMLU ~0.24, random chance).
Two root causes:
- `megatron_generate` / `megatron_prefill` used
`get_forward_backward_func()` (the training pipeline scheduler), which
is not designed for inference. Rewrote both functions to use explicit
P2P communication via `recv_from_prev_pipeline_rank_` /
`send_to_next_pipeline_rank`, matching the `run_mcore_inference`
pattern.
- `import_mcore_gpt_from_hf` loads HF weights into stage 0's embedding
but never updates the output_layer on the last PP stage when
`share_embeddings_and_output_weights=True`. At model init,
`setup_embeddings_and_output_layer()` all-reduces from stage 0 to sync
the output layer; after importing HF weights that all-reduce is stale.
Fix: call `model.setup_embeddings_and_output_layer()` again after
import.
**2. `megatron_mmlu` speedup (~6x)**
Replaces the `megatron_mmlu` implementation with a significantly faster
approach that matches how `lm-evaluation-harness` scores multiple-choice
questions.
**Before:** autoregressive generation (`megatron_generate`, `osl=2`) per
example, 114 separate `load_dataset` calls, batch_size=1 — 260s for 5%
data.
**After:** single prefill forward pass + argmax over {A,B,C,D} logits, 2
`load_dataset` calls, configurable batch_size — 18s for 5% data (~6x
faster).
### Changes
**PP fixes:**
- `megatron_generate` / `megatron_prefill`: replace
`get_forward_backward_func` with explicit P2P
(`recv_from_prev_pipeline_rank_` / `send_to_next_pipeline_rank`)
- `import_mcore_gpt_from_hf`: call
`model.setup_embeddings_and_output_layer()` after HF weight import when
PP>1 and `share_embeddings_and_output_weights=True`
- `megatron_prefill`: add `skip_return_logits` param and VLM support
(needed for PP non-last stages)
**MMLU speedup:**
- **Log-likelihood scoring**: replace `megatron_generate` with
`megatron_prefill` — one forward pass per batch, no autoregressive
decode loop
- **Global batching**: collect all examples across all subjects, sort by
descending sequence length, run in `batch_size` chunks
- **2 dataset loads** instead of 114: use `load_dataset("cais/mmlu",
"all")` with per-subject grouping; skip dev load when `few_shots=0`
- **`percentage` → `fraction`** parameter rename for clarity
- **tqdm progress bar** (rank-0 only)
### Testing
- `test_megatron_generate_and_mmlu` parametrized over `tp` and `pp`.
Accuracy assertion: `0.36 < score < 0.39`. Manually checked generated
text is coherent.
- Re-ran M-Bridge Minitron MMLU based pruning for Nano v2 9B -> 7B and
all top 10 candidate's MMLU numbers are ballpark similar as before
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ — `percentage` parameter
renamed to `fraction`; `enable_kv_cache` removed from `megatron_mmlu`
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing test updated and
parametrized for TP+PP
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
🤖 Generated with [Claude Code](https://claude.ai/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved pipeline-parallel generation and MMLU evaluation reliability;
fixed output-layer synchronization in shared-embedding + pipeline
setups.
* **New Features**
* MMLU scoring now uses batched prefill logit scoring for faster,
batched evaluation.
* **Behavior Changes**
* Default MMLU sampling increased from 5% to 10%; calibration batch
sizing adjusted and related CLI/help text updated.
* **Tests**
* Distributed tests cover tensor- and pipeline-parallel modes and
tighten MMLU validation ranges.
* **Documentation**
* Updated pruning example and benchmark timing to reflect new sampling
and speedup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
361f7e391b |
Merge puzzletron compression algorithm (#1121)
### What does this PR do? Implement puzzletron compression algorithm based on Puzzle paper (https://arxiv.org/abs/2411.19146) <details> <summary> Th list of reviewed and merged MRs that resulted in the feature/puzzletron branch</summary> Merging dkorzekwa/any_model to feature/puzzletron [Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull Request #974 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974) - merged [Draft: anymodel activation scoring by danielkorzekwa · Pull Request #989 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989) - merged [Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/) - merged [Draft: Merging anymodel:build_library_and_stats by danielkorzekwa · Pull Request #993 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993) - merged [Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull Request #994 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994) - merged [Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull Request #995 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995) - merged [Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request #1007 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/) - merged PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged [Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020) - merged [Merge any_model tutorial by danielkorzekwa · Pull Request #1035 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035) - merged [Merge mbridge distillation for any_model by danielkorzekwa · Pull Request #1036 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036) - merged [MR branch for the remaining difference between dkorzekwa/any_model an… by danielkorzekwa · Pull Request #1047 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047) - merged [Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071) - merged [Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request #1073 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073) - merged [Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request #1085 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085) - merged [Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull Request #1102 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102) - merged [Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull Request #1103 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103) - merged [code clean up by danielkorzekwa · Pull Request #1110 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110) - merged Merging into main: [Activation hooks redesign (reuse hooks component across both minitron and puzzletron) by danielkorzekwa · Pull Request #1022 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022) - merged [Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa · Pull Request #1115 · NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115) - merged </details> <!-- Details about the change. --> ### Usage Puzzletron tutorial: https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron ### Testing The main e2e test for compressing 9 models with Puzzletron: https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py 2-gpu nightly tests: - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203 - https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952 ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with AnyModel support, example pipelines, deployment and evaluation utilities, and tools for converting/pruning and exporting compressed checkpoints. * **Documentation** * Comprehensive Puzzletron tutorials, model-specific guides, evaluator instructions, example configs, and changelog entry. * **Chores** * CI/workflow updates (extras installation, longer GPU test timeout), pre-commit hook exclusion updated, and CODEOWNERS entries added. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com> Signed-off-by: jrausch <jrausch@nvidia.com> Signed-off-by: root <root@pool0-00848.cm.cluster> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com> Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com> Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
d0bf0bef96 |
Remove deprecated Nemo 2.0 references / examples (#1098)
### What does this PR do? - Remove `examples/nemo_run` and other deprecated Nemo 2.0 references - Add Megatron-Bridge example links where missing <!-- Details about the change. --> ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Deprecations** * Removed deprecated NeMo 2.0 support and related example flows and utilities. * **Documentation** * Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM and refreshed technique/deployment guidance and links. * **New Features** * Added CLI options for additional parallelism (context/expert tensor/expert model) in Megatron-Bridge distillation. * **Chores** * Removed legacy CI configs and refreshed container image tags across examples and docs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
b61fb4e16d |
Allow searcher ckpt dir for per-rank ckpt files (#1091)
### What does this PR do? Type of change: minor improvement <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Currently we modify user's `path.pth` to `path<rank>.pth` which may be confusing. Instead we now add alternative to provide `path/` and store `path/rank<rank>.pth` per-rank searcher state files ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated pruning examples and CLI help to reference directory-based checkpoint paths for intermediate pruning scores and clarified default wording. * **Improvements** * Checkpoint handling now accepts directory paths for intermediate score storage, produces per-rank files automatically, and resolves to a directory default when unset. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f22f4f5522 |
Add Full TE Spec support for Megatron Pruning DynamicModules + MoE bug fixes (#1024)
### What does this PR do? Type of change: Improvement + Bug Fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Quantization recently added support for Full TE spec. Adding same for Pruning as well so we can retire ModelOpt spec and just use standard TE spec. **NOTE: We still dont support TEGroupedGemm and instead use TE SequentialMLP for now (but this can be configured in standard TE Spec so we dont need modelopt spec)** Note that this does not affect the usage of the pruning workflow but makes pruning slightly faster and may result in slightly different pruned model because of different kernel and numerics. [Bug fix]: Previously NAS-based pruning for MoE models would hang when evaluating MMLU for pruned candidate models because of a bug. Fixed in this PR as well [Bug fix]: Previously hidden size importance hooks were not applied to pre_mlp_layernorm for MoE layers. Fixed in this PR as well resulting in a significant improvement in MMLU for Qwen3-30B-A3B ### Testing <!-- Mention how have you tested your change if applicable. --> - [x] Unit tests updated and passing - [x] Compare pruning results for Qwen3-8B -> 6B. ⚠️ Difference in MMLU scores resulting in a different best picked model. But scores more or less in similar range - difference may be because of different kernel for TE layers ``` Least important 6 layers: ModelOpt Spec: 27, 28, 29, 31, 32, 33 TE Spec: 27, 28, 30, 31, 32, 33 Top 10 pruned candidates: | num_layers | hidden_size | ffn_hidden_size | Params (B) | MMLU (ModelOpt Spec) | MMLU (TE Spec) | |------------|-------------|-----------------|------------|----------------------|----------------| | 34 | 3328 | 11264 | 5.99 | 0.390 | 0.397 | | 30 | 3584 | 11776 | 5.99 | 0.572 [BEST] | 0.575 | | 36 | 3840 | 8192 | 5.98 | 0.511 | 0.511 | | 36 | 3584 | 9216 | 5.98 | 0.477 | 0.497 | | 36 | 3072 | 11776 | 5.97 | 0.278 | 0.252 | | 32 | 3584 | 10752 | 5.96 | 0.542 | 0.541 | | 36 | 3328 | 10240 | 5.92 | 0.365 | 0.412 | | 34 | 3840 | 8704 | 5.91 | 0.537 | 0.539 | | 30 | 4096 | 9216 | 5.90 | 0.566 | 0.591 [BEST] | | 34 | 3584 | 9728 | 5.89 | 0.499 | 0.510 | ``` - [x] Compare pruning results for Nemotron-Nano-9B-v2 -> 7B. MMLU scores slight difference but best pruned model selection same ``` Least important 8 layers (Before and After): [43, 44, 45, 46, 47, 48, 50, 52] Top 10 pruned candidates: | num_layers | hidden_size | mamba_num_heads | mamba_head_dim | ffn_hidden_size | Params (B) | MMLU (ModelOpt Spec) | MMLU (TE Spec) | |------------|-------------|------------------|---------------|-----------------|------------|----------------------|----------------| | 50 | 4480 | 128 | 56 | 15680 | 7.00 | 0.211 | 0.202 | | 56 | 4096 | 96 | 80 | 14336 | 7.00 | 0.438 | 0.436 | | 48 | 4352 | 120 | 80 | 13824 | 7.00 | 0.679 [BEST] | 0.679 [BEST] | | 56 | 4352 | 112 | 80 | 10240 | 7.00 | 0.516 | 0.520 | | 54 | 4480 | 104 | 80 | 11264 | 7.00 | 0.263 | 0.262 | | 46 | 4480 | 128 | 72 | 14848 | 7.00 | 0.610 | 0.617 | | 50 | 4480 | 112 | 64 | 15680 | 7.00 | 0.426 | 0.421 | | 54 | 4096 | 112 | 80 | 13312 | 7.00 | 0.579 | 0.589 | | 56 | 4352 | 120 | 72 | 10752 | 7.00 | 0.466 | 0.469 | | 52 | 4352 | 120 | 72 | 12800 | 7.00 | 0.561 | 0.560 | ``` - [x] Compare pruning results for Qwen3-30B-A3B -> 24B. Previously there was a bug in hooks added so now we see a big improvement ``` Top 10 pruned candidates (~1 hour per candidate MMLU computation so skipped after 3): | num_layers | hidden_size | num_attention_heads | num_moe_experts | Params (B)| MMLU (ModelOpt Spec) | MMLU (TE Spec) | |------------|-------------|---------------------|-----------------|-----------|----------------------|----------------| | 46 | 2048 | 28 | 104 | 23.98B | 0.663 | 0.698 | | 40 | 2048 | 28 | 120 | 23.95B | 0.577 | 0.668 | | 46 | 1792 | 24 | 120 | 23.94B | 0.435 | 0.500 | | 46 | 2048 | 24 | 104 | 23.88B | | | | 40 | 2048 | 24 | 120 | 23.87B | | | | 46 | 1792 | 20 | 120 | 23.85B | | | | 40 | 2048 | 20 | 120 | 23.78B | | | | 46 | 2048 | 20 | 104 | 23.78B | | | | 42 | 2048 | 32 | 112 | 23.62B | | | | 48 | 1792 | 32 | 112 | 23.54B | | | ``` - [x] Run pruning experiments for gptoss-20b (21B actually) -> 18B with TESpec. Seems like GPTOSS MMLU is dropping steeply even in 15% pruning ``` "Only considering atmost 40% for width and 20% for depth pruning hparams Skipping hparams_to_skip=['num_attention_heads'] during search space generation... Search space for num_layers: [20, 22, 24] Search space for hidden_size: [2048, 2304, 2560, 2816, 2880] Search space for num_moe_experts: [24, 32] Search space for moe_ffn_hidden_size: [2048, 2304, 2560, 2816, 2880] Total search space in consideration: 150 Top 10 candidates with scores: {'num_layers': 20, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.62B params, 0.3780 score {'num_layers': 22, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2560} -> 17.32B params, 0.4160 score [BEST SUBNET] {'num_layers': 20, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 17.27B params, 0.3523 score {'num_layers': 20, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.23B params, 0.3848 score {'num_layers': 22, 'hidden_size': 2560, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.13B params, 0.3062 score {'num_layers': 24, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2304} -> 17.09B params, 0.3984 score {'num_layers': 22, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2560} -> 16.94B params, 0.3957 score {'num_layers': 20, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 16.88B params, 0.3835 score {'num_layers': 22, 'hidden_size': 2560, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 16.78B params, 0.2154 score {'num_layers': 24, 'hidden_size': 2304, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 16.73B params, 0.0014 score" ``` - [x] Run pruning experiments for Nemotron-3-Nano-30B-A3B (31.5B actually) -> 24B with TESpec ``` Top 10 candidates with scores: {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} -> 24.00B params, 0.0000 score {'num_layers': 52, 'hidden_size': 2048, 'mamba_num_heads': 64, 'num_moe_experts': 128, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 24.00B params, 0.2764 score {'num_layers': 48, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} -> 24.00B params, 0.6098 score {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328} -> 24.00B params, 0.6233 score {'num_layers': 48, 'hidden_size': 2688, 'mamba_num_heads': 64, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 24.00B params, 0.6301 score [BEST SUBNET] {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 2560} -> 23.99B params, 0.6125 score {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} -> 23.99B params, 0.4255 score {'num_layers': 50, 'hidden_size': 2048, 'mamba_num_heads': 64, 'num_moe_experts': 128, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3584} -> 23.99B params, 0.2859 score {'num_layers': 50, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 23.99B params, 0.6125 score {'num_layers': 42, 'hidden_size': 2304, 'mamba_num_heads': 40, 'num_moe_experts': 120, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3584} -> 23.99B params, 0.0366 score ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ⚠️ TE has different kernels so pruned model may be slightly different because of different numerics <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information OMNIML-3504 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Full Transformer Engine support for Minitron pruning; no custom model spec required. * **Bug Fixes** * Resolved pruning hang on MoE models by correcting importance-hook behavior. * **Documentation** * Updated changelog and example README; bumped recommended container tag and expanded Docker run/mount guidance; adjusted release dates. * **Improvements** * Per-rank local activation handling, broader candidate caching, runtime router scoring mitigation, clearer pruning status messages. * **Tests** * Updated tests to validate Transformer Engine backend and adjusted dynamic-module expectations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
fb5923c89e |
Add Megatron-Bridge pruning example scripts (#800)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
Megatron-Bridge pruning example scripts (HF input, HF / Megatron
output). Also defined some utility functions we can reuse for adding
examples for quantization or other optimizations:
- `modelopt.torch.utils.plugins.mbridge.load_mbridge_model_from_hf`:
Load HF to MBridge with ModelOpt spec in desired TP/PP/etc configuration
-
`modelopt.torch.utils.plugins.mbridge.get_hf_mbridge_calibration_loop`:
Create `forward_loop` for calibration on a HF dataset
- Supports all datasets available in
`modelopt.torch.utils.dataset_utils` (`cnn_dailymail`,
`nemotron-post-training-dataset-v2`, etc)
- Support applying chat template for chat-based data
## Usage
<!-- You can potentially add a usage example below. -->
From `nvcr.io/nvidian/nemo:26.02.rc1` container (mount latest code to
`/opt/Megatron-Bridge` and `/opt/Model-Optimizer`)
```python
torchrun --nproc_per_node 2 /opt/Model-Optimizer/examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_target_params 6e9 \
--hparams_to_skip num_attention_heads \
--output_hf_path /tmp/Qwen3-8B-Pruned-6B
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- [x] Manually ran pruning script in nemo:25.11 container (plus modelopt
and mbridge mounted to latest) for Qwen3-8B and Nemotron-Nano-9B-v2 with
PP=8 and PP=4
- [ ] Added per-PR CI/CD test for example script
Results when pruning Qwen3 8B -> 6B (10 different configurations) with
and without chat template on the dataset samples. Perhaps MMLU is not
the right metric to look at.
| Layers | Hidden Size | FFN Hidden Size | Params | MMLU (Concatenated
messages) | MMLU (Applied Chat Template) |
|--------|-------------|-----------------|--------|------------------|----------------------|
| 34 | 3328 | 11264 | 5.99B | 0.401 | 0.393 |
| 30 | 3584 | 11776 | 5.99B | 0.588 | 0.576 |
| 36 | 3840 | 8192 | 5.98B | 0.507 | 0.518 |
| 36 | 3584 | 9216 | 5.98B | 0.477 | 0.469 |
| 36 | 3072 | 11776 | 5.97B | 0.255 | 0.249 |
| 32 | 3584 | 10752 | 5.96B | 0.554 | 0.549 |
| 28 | 4096 | 10240 | 5.94B | 0.400 | 0.438 |
| 36 | 4096 | 7168 | 5.93B | 0.461 | 0.438 |
| 36 | 3328 | 10240 | 5.92B | 0.362 | 0.359 |
| 34 | 3840 | 8704 | 5.91B | 0.515 | 0.546 |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: ‼️ TODO
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
**New Features**
* Added new Megatron-Bridge pruning example demonstrating Minitron-based
model optimization with advanced pruning configurations.
**Documentation**
* Updated core project documentation to highlight Megatron-Bridge as a
supported optimization framework.
* Added comprehensive example documentation for Megatron-Bridge
workflows including pruning, distillation, and quantization.
* Updated pruning guides with Megatron-Bridge integration examples and
best practices.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
563a1e09c6 |
Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?
**Type of change:** New feature
- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
- Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.
## Usage
<!-- You can potentially add a usage example below. -->
Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
model,
mode="mcore_minitron",
constraints=constraints,
dummy_input=None, # Not used
config=config,
)
```
1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```
2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
return megatron_mmlu(m, tokenizer, percentage=0.05) # 5% sampled data for faster eval
config["score_func"] = score_func
# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune
mtp.prune(...)
```
To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
hidden_size_divisor=256,
ffn_hidden_size_divisor=512,
mamba_head_dim_divisor=8,
num_moe_experts_divisor=8,
num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```
## Testing
**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```
**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads': 88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads': 96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads': 80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```
## Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
OMNIML-3043
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
* Introduced parameter counting capabilities for architecture search
* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
* Updated configuration documentation with granular divisor parameters
* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
10d0fec907 |
Minitron pruning refactor [1/2]: Remove num_query_groups pruning and simplify megatron dynamic modules for follow-up PRs (#690)
## What does this PR do? - Deprecate `num_query_groups` parameter pruning support in `mcore_minitron` mode. This parameter pruning was not part of the paper but added hacky support in the past which makes refactoring current implementation a bit challenging hence removed its support. - Besides, `num_query_groups` should be multiple of TP so its less likely to be pruned. - Users can still prune it if they want to by using ModelOpt 0.40.0 or earlier - Going forward, `num_query_groups` if passed in pruning `export_config` should be same as original else raises an Exception. - As a follow-up PR, I will de-couple DynamicModule with forward hooks / hparam importance logic so its easy for users to change to different scoring logic ## Testing <!-- Mention how have you tested your change if applicable. --> - [x] CI/CD Tests pass - [x] Compare MMLU for pruning Qwen3-8B (nmm-sanbox) with previous and current implementation ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No. See reasoning above <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Deprecations** * num_query_groups parameter in Minitron pruning is deprecated. Use ModelOpt 0.40.0 or earlier if pruning with this parameter is needed. * **Documentation** * Updated NAS search parameters documentation * Refined pruning scope documentation for Transformer, MoE, and Mamba models * Updated example configurations and support matrix reflecting current supported pruning parameters <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
53a2ddebab |
Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer (OMNIML-3033) - [x] Mention in Latest News section with date on the date of merging this PR (12/08) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
be64f6b1d5 |
Deprecate examples/megatron-lm in favor of M-LM repo examples folder (#567)
- Deprecate `examples/megatron-lm` and move missing readme sections to https://github.com/NVIDIA/Megatron-LM/tree/main/examples/post_training/modelopt to keep only one copy of this doc. - Linked PR: https://github.com/NVIDIA/Megatron-LM/pull/2273 Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
6a8d6dac26 |
Enable Yarn RoPE in minitron pruning for gpt-oss support (#530)
## What does this PR do? **Type of change:** Improve existing feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** GPT-OSS model has Yarn RoPE which adds additional nn.Embedding modules that need to be enabled in DynamicModule for Minitron pruning ## Testing <!-- Mention how have you tested your change if applicable. --> - gpt-oss-20b pruned using M-LM pruning example and conf scripts. Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
41c66bb6d2 |
Add MoE (e.g. Qwen3-30B-A3B, Mamba hybrid) pruning support in Minitron (#467)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Support pruning `num_moe_experts`, `moe_ffn_hidden_size`, and `moe_shared_expert_intermediate_size` in `mcore_minitron` pruning ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added Mixture of Experts (MoE) pruning support with new configurable dimensions for expert count and intermediate sizes * Extended NAS architecture search capabilities to include MoE model parameters * **Documentation** * Updated support matrix and pruning documentation for MoE-compatible models * Clarified available pruning dimensions and parameters for MoE architectures <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
90b1e68fcf |
Update HF nvidia collection links in docs (#475)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
6dffcd0ff1 |
Add minitron pruning and distillation guidelines in pruning readme (#419)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
cb44c556d8 |
Minor pruning update (#397)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
b4d6cedb4a |
Support checkpointing Minitron pruning scores / prune without sorting (#361)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c0590b0255 |
Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c60baae5fa |
Add Megatron-LM pruning example link (#344)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
103b1bb137 |
Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4d1eb0caf5 |
Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
dcdc2df084 |
Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
2017cd9063 | Update for 0.25.0 release | ||
|
|
5c9390c215 | 0.23.1 Release - fix for torch 2.6 | ||
|
|
73d6af785f | Update 0.23.0 - OSS release |