mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Move Mistral Medium recipe under canonical base model (#2150)
### What does this PR do?
Type of change: Bug fix
Moves the Mistral Medium 3.5 checkpoint-mirror PTQ recipe under the
canonical Hugging Face base model, `mistralai/Mistral-Medium-3.5-128B`.
Updates the recipe catalog, loader smoke-test path, and changelog.
### Usage
```bash
python examples/hf_ptq/hf_ptq.py \
--model mistralai/Mistral-Medium-3.5-128B \
--recipe huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib
```
### Testing
- `uvx --from pre-commit pre-commit run --files CHANGELOG.rst
modelopt_recipes/huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml
modelopt_recipes/ptq.md tests/unit/recipe/test_loader.py`
- Recipe documentation consistency checks (all passed)
- Focused pytest was unavailable locally because the checkout
environment lacks pytest/Torch; the dedicated pre-commit recipe
validator passed.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ❌ The built-in recipe path
changes; the changelog documents the replacement path.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ Updated the built-in recipe
loader smoke test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A
### Additional Information
The recipe targets the vendor's canonical base checkpoint,
[`mistralai/Mistral-Medium-3.5-128B`](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B),
and reproduces NVIDIA's
[`Mistral-Medium-3.5-128B-NVFP4`](https://huggingface.co/nvidia/Mistral-Medium-3.5-128B-NVFP4)
quantization map.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added a Mistral Medium 3.5 128B post-training quantization recipe
supporting NVFP4 and FP8 configurations with max calibration.
* **Documentation**
* Updated checkpoint-mirror guidance to reflect the new model-based
recipe location.
* **Breaking Changes**
* Saved references to the previous recipe path must be updated to the
new Mistral Medium 3.5 recipe path.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
This commit is contained in:
@@ -22,6 +22,7 @@ Changelog
|
||||
|
||||
**Backward Breaking Changes**
|
||||
|
||||
- Move the Mistral Medium 3.5 checkpoint-mirror recipe from ``huggingface/models/nvidia/Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib`` to ``huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib``, keying it by the canonical Hugging Face base model. Update any saved ``--recipe`` paths to the new location.
|
||||
- Transformer Engine ``TEGroupedMLP`` (fused MoE experts) now uses **per-expert** weight quantization (one ``amax`` per expert) instead of a single shared ``amax``, so ModelOpt checkpoints containing quantized ``TEGroupedMLP`` modules saved before 0.47 are **not compatible** with 0.47. Re-run PTQ to regenerate compatible checkpoints.
|
||||
|
||||
**Deprecations**
|
||||
|
||||
@@ -235,7 +235,7 @@ that baseline. The deviations come in four kinds:
|
||||
| **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama` |
|
||||
| **Algorithm override** | Same numerics & scope, but the *calibration algorithm* is tweaked because the default breaks or regresses | `gemma`, `gemma4`, `mpt` |
|
||||
| **Extra exclusions** | Adds disabled-quantizer patterns so non-language branches stay full precision | `nemotron_vl`, `diffusion_gemma` |
|
||||
| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/Nemotron-3-*`, `models/nvidia/Mistral-Medium-3.5-128B-NVFP4` |
|
||||
| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/Nemotron-3-*`, `models/mistralai/Mistral-Medium-3.5-128B` |
|
||||
|
||||
The numerics and standard exclusions are still inherited from `configs/`
|
||||
wherever possible — the model folder captures *only* the delta. Each `<task>/`
|
||||
@@ -335,12 +335,12 @@ encoders (or the never-calibrated self-conditioning branch), regressing those
|
||||
modalities or crashing export. The extra patterns keep them in full precision;
|
||||
everything else matches the general recipe.
|
||||
|
||||
### Checkpoint mirrors — `models/nvidia/<checkpoint>`
|
||||
### Checkpoint mirrors — `models/<org>/<checkpoint>`
|
||||
|
||||
The `huggingface/models/` tier reproduces a **single published (or planned)
|
||||
checkpoint's** quant config verbatim:
|
||||
|
||||
- **`Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib`** mirrors
|
||||
- **`models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib`** mirrors
|
||||
`nvidia/Mistral-Medium-3.5-128B-NVFP4`: decoder MLP layers 4–86 use NVFP4
|
||||
W4A4, edge MLP layers 0–3 and 87 use FP8 W8A8, and all attention projections
|
||||
and the KV cache use FP8. It uses max calibration.
|
||||
|
||||
@@ -167,7 +167,7 @@ _BUILTIN_PTQ_RECIPES = [
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8",
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8_layerwise",
|
||||
"huggingface/models/nvidia/Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib",
|
||||
"huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib",
|
||||
"general/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
"general/ptq/nvfp4_mlp_only-novit-kv_fp8",
|
||||
"general/ptq/nvfp4_mlp_only-kv_fp8_cast",
|
||||
|
||||
Reference in New Issue
Block a user