Commit Graph
3 Commits
Author SHA1 Message Date
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00
realAsma 4b04e732ed Fix prequant layernorm export without scales (#1838)
## Summary
- Skip pre-quant LayerNorm fusion when the representative input
quantizer has no `_pre_quant_scale` buffer.
- Add CPU unit coverage for the no-scale weight-only path and the
existing fusion/removal behavior when pre-quant scales are present.
- Compose the public `general/ptq/int4_blockwise_weight_only` recipe
from the shared recipe units so it stays aligned with
`INT4_BLOCKWISE_WEIGHT_ONLY_CFG`.

## Context
NVBug: https://nvbugspro.nvidia.com/bug/6311597

NVBug 6311597 reports a deterministic HF export crash for
`general/ptq/int4_blockwise_weight_only` on Llama-3.1-8B. That
weight-only recipe disables input quantization and skips calibration, so
`_pre_quant_scale` is not registered, but export still reaches
`fuse_prequant_layernorm` through the AWQ-like format path.

This PR also carries the recipe-sync diff that was previously in draft
PR #1836; PR #1836 has been closed after moving that diff here.

## Tests
- `pre-commit run --files modelopt/torch/export/quant_utils.py
tests/unit/torch/export/test_unified_export_hf.py
modelopt_recipes/general/ptq/int4_blockwise_weight_only.yaml
tests/unit/recipe/test_loader.py`
- `pytest_pwd
tests/unit/torch/export/test_unified_export_hf.py::test_fuse_prequant_layernorm_skips_modules_without_pre_quant_scale
tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins
tests/unit/recipe/test_loader.py::test_general_ptq_yaml_matches_config_dicts
tests/unit/recipe/test_presets.py -q` -> 29 passed

Full GB10/Llama repro was not run locally.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved handling of layer norm fusion so it safely skips cases where
pre-quantization scale data is unavailable, avoiding unexpected errors.
* Updated the layer norm fusion flow to correctly apply scaling when
pre-quantization data is present.

* **Tests**
* Expanded coverage for export and recipe loading scenarios, including
cases with and without pre-quantization scale data.
  * Added validation for the new int4 blockwise weight-only recipe.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-06-26 17:34:47 -07:00
realAsmaandClaude Opus 4.8 196c091027 [1/N] Refactor llm_qat example: YAML configs + ModelOptArgParser (#1172)
### What does this PR do?

Type of change: new example

Refactors `examples/llm_qat` from a monolithic launch/script flow into a
modular, config-driven Hugging Face QAT/QAD workflow. The high-level
user flow is now:

1. Quantize a base model with a ModelOpt PTQ recipe.
2. Train or evaluate the quantized checkpoint with QAT, QAD, LoRA QAT,
QLoRA, or fine-tuning configs.
3. Export the trained checkpoint for deployment.

Highlights:

- Replaces the legacy `examples/llm_qat/launch.sh` +
`examples/llm_qat/main.py` path with separate `quantize.py` and
`train.py` entrypoints.
- Adds YAML-driven argument parsing through `ModelOptArgParser`,
including `--config <yaml>` defaults, CLI overrides, and generated
`examples/llm_qat/ARGUMENTS.md`.
- Adds declarative configs under `examples/llm_qat/configs/` for
training modes, dataset blends, and Accelerate backends.
- Adds `dataset_utils.py` for weighted multi-source dataset blending,
Hugging Face streaming, local dataset loading, distributed rank-aware
loading, tokenization caching, pre-tokenization, chat templating, and
assistant-token label masking.
- Moves the Hugging Face QAD flow into the shared trainer path through
`DistillArguments`, `QADTrainer`, and teacher-model distillation kwargs.
- Updates `QuantizationArguments` to prefer recipe paths via `--recipe`,
while keeping legacy `--quant_cfg` available with deprecation warnings
for in-trainer quantization.
- Updates `QATTrainer` handling for pre-quantized checkpoints, FSDP2
TensorQuantizer buffers, and recipe-resolved PTQ configs.
- Adds the `general/ptq/int4_blockwise_weight_only` recipe and refreshes
docs for NVFP4, FP8, INT4, FSDP2, DDP, DeepSpeed, QLoRA, and
LLaMA-Factory integration.
- Adds focused parser, dataset tokenization, assistant-mask, and example
workflow coverage.

### Usage

From the repo root:

```sh
cd examples/llm_qat

# 1. Quantize
python quantize.py \
  --model_name_or_path Qwen/Qwen3-8B \
  --dataset_config configs/dataset/blend.yaml \
  --recipe general/ptq/nvfp4_default-kv_fp8 \
  --output_dir qwen3-8b-quantized

# 2. QAT train
accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
  --config configs/train/qat_nvfp4.yaml \
  --model_name_or_path qwen3-8b-quantized \
  --output_dir qwen3-8b-qat-nvfp4

# 3. QAD train
accelerate launch --config-file configs/accelerate/fsdp2.yaml train.py \
  --config configs/train/qad_nvfp4.yaml \
  --model_name_or_path qwen3-8b-quantized \
  --teacher_model Qwen/Qwen3-8B \
  --output_dir qwen3-8b-qad-nvfp4
```

Dataset blends can be pre-tokenized and cached before training:

```sh
cd examples/llm_qat
python dataset_utils.py \
  --dataset_config configs/dataset/blend.yaml \
  --model_name_or_path Qwen/Qwen3-8B
```

### Testing

Focused coverage added or updated:

- `tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py`
- `tests/examples/llm_qat/test_dataset_tokenization.py`
- `tests/examples/llm_qat/test_assistant_mask.py`
- `tests/examples/llm_qat/test_llm_qat.py`

Recorded validation:

- [x] `pytest tests/unit/torch/opt/plugins/test_modelopt_arg_parser.py`
- [x] `pytest
tests/examples/llm_qat/test_llm_qat.py::test_dataset_utils_pretokenize`
- [x] `pre-commit run --all-files`
- [ ] `pytest tests/examples/llm_qat/test_llm_qat.py` full GPU/backend
suite

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: No for the legacy
`examples/llm_qat` CLI/file layout (`main.py`, `launch.sh`, and FSDP1
config are removed); library `quant_cfg` usage remains available but is
deprecated for this workflow.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: Yes.
`train.py`/`simple_qat_train.py` retain the upstream Alpaca attribution
where applicable; `examples/llm_qat/requirements.txt` switches
`tensorboardX` to `tensorboard`.
- Did you write any new necessary tests?: Yes. Parser, dataset
tokenization, assistant masking, pre-tokenization, and QAT/QAD workflow
tests were added or updated.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes.
- Did you get Claude approval on this PR?: N/A in this description
update.

### Additional Information

This PR intentionally changes the `llm_qat` example surface. Existing
users should move from `examples/llm_qat/main.py` and
`examples/llm_qat/launch.sh` to `examples/llm_qat/quantize.py`,
`examples/llm_qat/train.py`, and the YAML configs under
`examples/llm_qat/configs/`.

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 01:25:02 +00:00