3 Commits
Author SHA1 Message Date
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00
Keval MorabiaandClaude Opus 4.8 42458def24 ci: fix torch_trt on torch 2.13; default unit tests to torch 2.13; add example allow-failure hatch (#1951)
### What does this PR do?

Type of change: Bug fix + CI / infra

torch 2.13 + torchvision 0.28 were published to PyPI on 2026-07-08 and
broke the `onnx (torch_trt)` example job (which had passed the day
before). This PR fixes that break and hardens CI against the next one:

1. **Fix `torch_trt` (torch/torchvision/torch-tensorrt trio pin).** The
base install pulled torch 2.13 / torchvision 0.28, then `torch-tensorrt
2.12.1` downgraded torch back to 2.12 but left torchvision at 0.28
(which pins `torch==2.13`) — breaking `import`.
`examples/torch_trt/requirements.txt` now caps
`torch-tensorrt>=2.4.0,<2.13` + `torchvision<0.28` so the trio stays
consistent (also protects direct `pip install -r` users).
2. **Unit tests default to torch 2.13.** `noxfile.py` gains `torch_213`
(`torchvision~=0.28.0`); the required `linux`/`windows` jobs and the
multi-version Python spread (3.10/3.11/3.13/3.14) now run torch 2.13,
with torch 2.8–2.12 kept as back-compat legs on Python 3.12.
3. **Per-example allow-failure escape hatch.**
`_example_tests_runner.yml` gains an `allow_failure` input; when set, a
**test-run** failure is surfaced as a `::warning::` via
`continue-on-error` instead of blocking the PR. `example_tests.yml`
derives it per example from the repo variable
**`ALLOW_FAILURE_EXAMPLE_TESTS`** (comma-separated example names,
comma-wrapped so `onnx` ≠ `torch_onnx`). Future breakages can be
quarantined by updating the variable — no code change / PR required.

### Testing

- Ran the new default unit session locally in an isolated uv venv (torch
**2.13.0**+cu130, torchvision **0.28.0**+cu130, transformers
**5.12.1**):
  ```
  nox -s "unit-3.12(torch_213, tf_latest)"
  => 2813 passed, 15 skipped, 1786 warnings in 250.37s
  ```
- Verified the allow-failure hatch: with
`ALLOW_FAILURE_EXAMPLE_TESTS=torch_trt`, the (previously failing)
`torch_trt` job reports success with a warning and does not block the
required example check. `vars` is re-read on each job attempt, so
"Re-run failed jobs" picks up the variable without a fresh trigger.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (CI-only; older torch versions
still covered)
- If you copied code from any other sources or added a new PIP
dependency: N/A (no new dependency; only version caps)
- Did you write any new necessary tests?: N/A (CI configuration change)
- Did you update Changelog?: N/A (CI infra, no user-facing API change)
- Did you get Claude approval on this PR?: ❌ (pending — will run
`/claude review`)

### Additional Information

The `ALLOW_FAILURE_EXAMPLE_TESTS` repo variable can be cleared for
`torch_trt` now that the requirements pin lands the real fix; keep it as
the standing escape hatch for future example breakages.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Example test jobs can now be configured to “allow failure” without
failing the workflow; when enabled, a warning annotation is emitted.
* **Bug Fixes**
* CI unit-test and GPU-test configurations were refreshed (including a
reduced timeout for the `gpu_megatron` job).
* **Tests**
* Updated unit-test coverage to use the newest Torch 2.13-based setup by
default, with back-compat retained where applicable.
* **Documentation**
* Added inline guidance for how the allow-failure examples list is
specified.
* **Chores**
* Refreshed `torch_trt` example dependency constraints to improve
compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 19:01:37 +05:30
Ajinkya RasaneandClaude Opus 4.8 fbbc5989ce [6078291][OMNIML-3716] Add ViT FP8 + Torch-TRT example, wire softmax_quantizer in _QuantAttention (#1569)
### What does this PR do?

Type of change: new feature + bug fix

Adds a Torch-TensorRT deployment path for HuggingFace ViT and closes the
modelopt-side gap that prevented `*softmax_quantizer` from being applied
on the standard attention forward path.

* **New ViT PTQ recipes** under `modelopt_recipes/huggingface/vit/ptq/`:
* `fp8.yaml` — W8A8 per-tensor FP8 E4M3 on encoder Linear
weights/inputs;
attention Q/K/V BMMs + softmax output at FP8; per-block LayerNorm output
    at FP8 (one shared Q/DQ feeds Q/K/V + MLP); patch-embed `nn.Conv2d`,
    `classifier`, and the final `vit.layernorm` left FP16. Uses max
    calibration.
  * The recipe is self-contained (no `$import` of shared snippets) and
use the "specific-enable" style: narrow `parent_class` + path scoping
    on every enable rule, so no `enable: false` carve-outs are needed.

* **New example** under `examples/torch_trt/`:
  * `torch_tensorrt_ptq.py` — single-model pipeline (load HF model,
    calibrate from `zh-plus/tiny-imagenet`, `mtq.quantize`,
`torch_tensorrt.compile`, verify the compiled-model argmax matches the
    fake-quant argmax). Defaults to `google/vit-large-patch16-224`; pass
`--model_id` and `--recipe` to target any model + recipe combination.
`--no_pretrained` + `--model_kwargs` shrink the model for fast tests.
  * `README.md` documenting the flow, the shipped recipes, hardware
    requirements, and CLI usage.
  * `requirements.txt`.

* **Bug fix in `modelopt/torch/quantization/plugins/huggingface.py`** —
inside
  `_QuantAttention._quantized_attention`, the non-kitchen branch now
  temporarily replaces `torch.nn.functional.softmax` (via the existing
`replace_function` context manager) with a wrapper that pipes the
softmax
output through `self.softmax_quantizer`. Previously the slot was created
  on every registered attention class but only consumed by the optional
Kitchen MXFP8 flash-attention path, so FP8 / NVFP4 recipes that enabled
`*softmax_quantizer` saw it stay uncalibrated (`amax=None`) and emitted
  no Q/DQ around the softmax output during ONNX / Torch-TRT export. With
  this fix the `softmax_quantizer` is calibrated alongside the rest of
the model, and both the modelopt ONNX exporter and
`torch_tensorrt.compile`
  pick up the Q/DQ pair. The patch short-circuits to the unwrapped call
when the quantizer is disabled (zero-overhead) and has no effect on SDPA
  paths that fuse softmax inside a C++ kernel.

* **New e2e integration test** at
  `tests/examples/torch_trt/test_torch_tensorrt_ptq.py` — mirrors the
  `torch_onnx` test pattern: invokes the example through
  `run_example_command`, parametrizes over the two precision modes (fp8,
nvfp4), uses a 1-layer ViT config (`--no_pretrained` + `--model_kwargs`)
so each parametrized case completes in under a minute. `importorskip` on
  `torch_tensorrt` so the test is automatically skipped on hosts without
  the package.

### Usage

```bash
# FP8 (Hopper / Ada) — default model is google/vit-large-patch16-224
python examples/torch_trt/torch_tensorrt_ptq.py \
    --precision fp8 \
    --calib_samples 128 \
    --batch_size 1

# Custom model + custom recipe
python examples/torch_trt/torch_tensorrt_ptq.py \
    --model_id <huggingface/model-id> \
    --recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml>
```

### Testing

* Recipes load via `modelopt.recipe.load_recipe()` and pass
  `QuantizeConfig` schema validation.
* Run `pytest tests/examples/torch_trt/test_torch_tensorrt_ptq.py` →
  1 parametrized case passes on RTX 6000 Ada (fp8).
* End-to-end on `google/vit-base-patch16-224`: `mtq.quantize` with the
new
FP8 recipe followed by `torch_tensorrt.compile(ir="dynamo")` produces a
  TRT engine whose argmax matches the FP16 baseline.
* ONNX exported from the torch path now contains Q/DQ on **12 / 12**
  softmax outputs (was 0 / 12 before this PR's `_QuantAttention` fix),
  matching the ONNX-CLI output's quantization layout.

Both FP8 paths land within 0.13 pp Top-1 of the FP16 baseline; Top-5 is
  within 0.02 pp across all three.
* ImageNet-1k validation accuracy via the new
`torch_tensorrt_accuracy.py`
  (full 50000 samples, batch=1, **every model Torch-TensorRT-compiled —
including the baseline** — so the comparison is apples-to-apples) for
the
  example's default `google/vit-large-patch16-224`:

  | Model (Torch-TRT) | Top-1 | Top-5 | Δ Top-1 vs baseline |
  |---|---:|---:|---:|
  | Baseline (FP16) | 81.99% | 96.01% | — |
  | FP8 | 82.01% | 96.05% | +0.02 pp |

  FP8 is within noise of the FP16 TRT baseline and NVFP4 W4A4 costs only
−0.13 pp Top-1 / −0.05 pp Top-5. Absolute Top-1 sits below the model
card's
  ~85.5% because evaluation uses the HF `AutoImageProcessor` default
preprocessing (direct 224×224 resize, no resize-then-center-crop),
applied
identically to all three models — so the deltas are the comparison
signal.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — new e2e integration test
under `tests/examples/torch_trt/`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Torch‑TensorRT FP8/NVFP4 deployment examples and end‑to‑end
scripts for HuggingFace ViT, plus ViT-specific PTQ recipes and
ImageNet-1k vs FP16 accuracy reporting.

* **Bug Fixes**
* Fixed softmax quantization and export/compilation edge cases (softmax
calibration during export, IO casting for empty tensors, routed expert
weight syncing, importer key handling).

* **Documentation**
* Added comprehensive example README with setup, usage, recipes,
evaluation, and hardware guidance.

* **Requirements**
  * Pinned minimum versions for example dependencies.

* **Tests**
* Added tests validating the Torch‑TensorRT quantization examples for
fp8.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 20:28:48 +00:00