mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
449a39922b |
Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do? Type of change: Bug fix `nemo_automodel` 0.6.0 removed `nemo_automodel.recipes.diffusion.train.is_main_process` without a replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0 defines no equivalent anywhere in the package). `examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's import guard fires and **every** test in `tests/examples/diffusers/` errors at collection: ``` ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train' tests/examples/diffusers/fastgen/test_resume_dataloader.py E ImportError: The DMD2 fastgen example requires `nemo_automodel`. ... collected 42 items / 1 error ``` The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was published and the `onnx (diffusers)` job started failing on every PR (e.g. runs 33020467654, 33019418460, 33010815298, 33007613265, 33006944292 — all unrelated branches). Capping at `<0.6` restores the tested range. Every other `nemo_automodel` symbol the example imports still exists in 0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`, `recipes.diffusion.train.TrainDiffusionRecipe`, and the four `components.datasets.diffusion.*` helpers), so `is_main_process` is the only blocker; the alternative is defining that predicate locally and widening the cap again, which is worth doing separately if the example is meant to track 0.6. ### Usage ```bash pip install -r examples/diffusers/fastgen/requirements.txt ``` ### Testing Reproduced the break by diffing the published wheels: `is_main_process` is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0 and absent from 0.6.0 (`grep -rn "def is_main_process"` over the unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported symbols are all still present in 0.6.0. CI on this PR exercises the fix directly: the `onnx (diffusers)` job installs from this requirements file and is the job that has been failing. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing dependency, tightened bound. - Did you write any new necessary tests?: N/A — the existing `tests/examples/diffusers/` suite is what this unblocks. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — dependency-pin fix for a break introduced and fixed within the same unreleased cycle. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed dependency compatibility for the FastGen diffusion example. * Prevented installation of versions that could cause the example to fail at startup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
090b1c5114 |
fastgen DMD2: make the Qwen-Image example self-contained on stock nemo_automodel (#1688)
### What does this PR do?
**Type of change:** new example / refactor (example self-containment)
The published `examples/diffusers/fastgen` DMD2 Qwen-Image distillation
example previously relied on
**local, unpublished modifications to the sibling `nemo_automodel`
package** (data path, collate, a
partial-load checkpointer, and Qwen-Image preprocessing). An external
user running the published
Model-Optimizer example against **official stock `nemo_automodel`**
would hit import/attribute errors.
This PR makes the example **self-contained on stock
`nemo_automodel>=0.4.0`** — with the changes kept
**as small as possible**: the team's actual AutoModel delta is only ~430
lines, so wherever the
upstream module is importable, the delta is expressed as a thin subclass
/ small reimplementation rather than a
full-file copy.
- **`fastgen_data/`** — the DMD2 data path:
- `collate_fns.py` — **reuses the upstream `SequentialBucketSampler` and
reimplements the collate**:
it builds the DMD2 batch directly from the vendored dataset's per-item
output (`image_latents` /
`text_embeddings` / `text_embeddings_mask` + an optional broadcast
`negative_text_embeddings` for
CFG) and deliberately does **not** call upstream
`collate_fn_production`, which stacks
model-specific token keys (`clip_tokens` / `t5_tokens`) absent from the
Qwen-Image cache. The
builder loads an optional `negative_prompt_embedding_path`.
- `text_to_image_dataset.py` — a **faithful vendored copy** of the
upstream reader (its
`prompt_embeds_mask` emission is interleaved with cache loading, so
wrapping it would force a
redundant per-item `torch.load`; carried verbatim instead).
- **`fastgen_checkpoint.py`** — `PartialLoadCheckpointer(Checkpointer)`
that overrides only
`load_optimizer` (FSDP2 `DefaultLoadPlanner(allow_partial_load=True)`)
so optimizer resume works
without patching upstream. Injected via an in-place re-bless of
`self.checkpointer` in the
recipe's `load_checkpoint` (model-state load stays strict).
- **`preprocess/`** — Qwen-Image preprocessing
(`preprocessing_multiprocess.py` + `processors/`),
trimmed to the image path (drops the flux/wan/hunyuan processors and the
video base class).
It lives in AutoModel's top-level `tools/` tree, which is **not**
shipped in the pip package, so
it cannot be wrapped and is vendored; `MultiTierBucketCalculator` is
imported from stock upstream.
- **`make_negative_prompt_embedding.py`** — generates the optional CFG
negative-prompt embedding.
- All `configs/*.yaml` target `fastgen_data.build_*` (a test enumerates
every config).
- Licensing: the AutoModel-copied files are NVIDIA-authored Apache-2.0,
so they carry only the
standard NVIDIA SPDX header (managed by the `insert-license` hook) — no
per-file provenance note,
no duplicated license, no pre-commit exclusion, and no separate
`LICENSE` note.
`nemo_automodel[diffusion]` version bound in `requirements.txt`.
The DMD2 math in `modelopt/torch/fastgen/` is **unchanged** — only
example/training-time glue moved.
### Usage
```bash
# Install example deps (stock nemo_automodel) from a source checkout
pip install -r examples/diffusers/fastgen/requirements.txt
# Build the training cache from raw images (Qwen-Image VAE latents + text embeddings)
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
--image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
--caption_format meta_json
# Generate the CFG negative-prompt embedding once
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
--output <cache dir>/negative_prompt_embedding.pt
# Point the config's data.dataloader.cache_dir + negative_prompt_embedding_path at the cache, then train.
```
See `examples/diffusers/fastgen/README.md` → "Requirements &
self-contained data path".
### Testing
- New `tests/examples/diffusers/fastgen/test_vendored_migration.py`:
environment-independent
invariants (every config targets a vendored builder; no `tools.*`
imports; each former AutoModel
patch is vendored / wrapped / a documented exclusion; the
former-vendored files carry the standard NVIDIA SPDX header, no
provenance note or duplicate license) plus
dependency-guarded structural tests (the collate emits the batch
contract + broadcasts the
negative embedding; the builder accepts
`negative_prompt_embedding_path`; the checkpointer
overrides only `load_optimizer`; the Qwen-Image processor
self-registers).
- **Validated 9/9 against a pure stock `nemo_automodel` 0.4.0 worktree**
(none of the local patches
present) via SLURM — re-run after this slim-down.
- The migrated code path is exercised by a live multi-GPU DMD2 run that
resumed from a checkpoint
through the vendored `PartialLoadCheckpointer`.
- `ruff check` + `ruff format --check` clean on all changed files.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (additive; bundled configs
target the vendored builders)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
(NeMo-AutoModel @ `e42584e3`, Apache-2.0; per review these
NVIDIA-authored files carry the standard NVIDIA SPDX header, no separate
provenance / `LICENSE` note; `nemo_automodel` was already a dependency)
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌ (pending — opened as draft)
### Additional Information
Opened as a **draft** pending: OSRB review of the vendored
NeMo-AutoModel (Apache-2.0) code, CI green,
and (optional) a multi-GPU smoke-train + resume on stock upstream.
Vendored from
NVIDIA-NeMo/Automodel at commit `e42584e3`.
### Update (post-review)
- **Mid-run resume data-correctness fix** (`6ffbc52c9`): on resume the
`StatefulDataLoader`'s restored state did not advance past the resume
point, so each window re-served the same data slice and multi-window
(SLURM-windowed) runs under-covered the dataset. Fixed by rebuilding a
fresh loader and skipping the deterministic sampler to the position
implied by `global_step`; added a SLURM-free, GPU-free CPU regression
test (`tests/examples/diffusers/fastgen/test_resume_dataloader.py`).
- **Licensing review** (`be832ae95`): the AutoModel-copied files are
NVIDIA-authored Apache-2.0, so they now carry only the standard NVIDIA
SPDX header — dropped the per-file provenance note, the duplicated
original-license block, the `insert-license` pre-commit exclusion, and
the `LICENSE` note.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added vendored Qwen‑Image preprocessing (multi-process) with processor
registry support.
* Updated DMD2 data loading with a dedicated dataset/collation pipeline,
including negative-prompt embedding/mask handling.
* Improved training resume behavior by rebuilding dataloader state and
making checkpoint optimizer restore tolerant of partial FSDP2 optimizer
shards.
* **Documentation**
* Refreshed the fastgen README and config notes for real-data training;
removed the prior mock-data smoke workflow.
* **Tests**
* Added regression and migration tests covering vendored wiring,
collate/dataloader contracts, processor registration, and
resume/checkpoint behavior.
* **Chores**
* Updated licenses/attribution, vendoring/tooling guards, requirements
pinning, linting configuration, and repository ownership rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
a21197c277 |
Remove unsafe torch.load from examples/diffusers/fastgen (#1740)
Follow-up to #1326 - Remove unsafe `torch.load(..., weights_only=False)` in `examples/diffusers/fastgen` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated checkpoint loading mechanisms in FastGen examples to improve compatibility and reliability during model restoration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
aec72ffa68 |
Add DMD2 distillation for Qwen-Image (fastgen) (#1326)
### What does this PR do?
**Type of change:** New example + new `modelopt.torch.fastgen` library
module.
Adds **DMD2 (Distribution Matching Distillation) for Qwen-Image** —
distilling the base model into a few-step (1–4) generator. Includes the
framework-agnostic `modelopt.torch.fastgen` loss library (DMD pipeline,
EMA, optional GAN discriminator) and a NeMo AutoModel–based training
example with a mock-data smoke config, a real-data config, and inference
/ export scripts.
**Noted**: the example script will be migrated to AutoModel repo
### Usage
```bash
# Mock-data wiring smoke — runs end-to-end with no dataset to prepare
torchrun --nproc-per-node=8 \
examples/diffusers/fastgen/dmd2_finetune.py \
--config examples/diffusers/fastgen/configs/dmd2_qwen_image_smoke.yaml
```
See `examples/diffusers/fastgen/README.md` for real-data training and
inference.
### Testing
Unit tests under `tests/unit/torch/fastgen/`; `pre-commit` /
code-quality clean.
### Before your PR is "*Ready for review*"
- Backward compatible?: ✅ (new, additive module)
- Followed `CONTRIBUTING.md` for any copied code / new deps: ✅
- New tests added?: ✅
- Updated Changelog?: N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Adds a FastGen-based distillation framework (DMD2) with
student/fake-score training, EMA support, GAN discriminator branch,
inference pipeline, and export utilities.
* Qwen-Image integration with latent packing and feature-capture for
plugin-enabled pipelines.
* **Documentation**
* New README, example configs, and runnable example scripts for
Qwen-Image distillation and inference.
* **Tests**
* Comprehensive unit tests covering math parity, gradient routing,
plugins, hooks, EMA, and recipe setup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|