4 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Opus 5 449a39922b Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do?

Type of change: Bug fix

`nemo_automodel` 0.6.0 removed
`nemo_automodel.recipes.diffusion.train.is_main_process` without a
replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0
defines no equivalent anywhere in the package).
`examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's
import guard fires and **every** test in `tests/examples/diffusers/`
errors at collection:

```
ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train'
  tests/examples/diffusers/fastgen/test_resume_dataloader.py
E   ImportError: The DMD2 fastgen example requires `nemo_automodel`. ...
collected 42 items / 1 error
```

The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was
published and the `onnx (diffusers)` job started failing on every PR
(e.g. runs 33020467654, 33019418460, 33010815298, 33007613265,
33006944292 — all unrelated branches). Capping at `<0.6` restores the
tested range.

Every other `nemo_automodel` symbol the example imports still exists in
0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`,
`recipes.diffusion.train.TrainDiffusionRecipe`, and the four
`components.datasets.diffusion.*` helpers), so `is_main_process` is the
only blocker; the alternative is defining that predicate locally and
widening the cap again, which is worth doing separately if the example
is meant to track 0.6.

### Usage

```bash
pip install -r examples/diffusers/fastgen/requirements.txt
```

### Testing

Reproduced the break by diffing the published wheels: `is_main_process`
is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0
and absent from 0.6.0 (`grep -rn "def is_main_process"` over the
unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported
symbols are all still present in 0.6.0.

CI on this PR exercises the fix directly: the `onnx (diffusers)` job
installs from this requirements file and is the job that has been
failing.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing
dependency, tightened bound.
- Did you write any new necessary tests?: N/A — the existing
`tests/examples/diffusers/` suite is what this unblocks.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — dependency-pin fix for a break introduced and fixed within the
same unreleased cycle.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dependency compatibility for the FastGen diffusion example.
* Prevented installation of versions that could cause the example to
fail at startup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 05:59:02 +05:30
090b1c5114 fastgen DMD2: make the Qwen-Image example self-contained on stock nemo_automodel (#1688)
### What does this PR do?

**Type of change:** new example / refactor (example self-containment)

The published `examples/diffusers/fastgen` DMD2 Qwen-Image distillation
example previously relied on
**local, unpublished modifications to the sibling `nemo_automodel`
package** (data path, collate, a
partial-load checkpointer, and Qwen-Image preprocessing). An external
user running the published
Model-Optimizer example against **official stock `nemo_automodel`**
would hit import/attribute errors.

This PR makes the example **self-contained on stock
`nemo_automodel>=0.4.0`** — with the changes kept
**as small as possible**: the team's actual AutoModel delta is only ~430
lines, so wherever the
upstream module is importable, the delta is expressed as a thin subclass
/ small reimplementation rather than a
full-file copy.

- **`fastgen_data/`** — the DMD2 data path:
- `collate_fns.py` — **reuses the upstream `SequentialBucketSampler` and
reimplements the collate**:
it builds the DMD2 batch directly from the vendored dataset's per-item
output (`image_latents` /
`text_embeddings` / `text_embeddings_mask` + an optional broadcast
`negative_text_embeddings` for
CFG) and deliberately does **not** call upstream
`collate_fn_production`, which stacks
model-specific token keys (`clip_tokens` / `t5_tokens`) absent from the
Qwen-Image cache. The
    builder loads an optional `negative_prompt_embedding_path`.
- `text_to_image_dataset.py` — a **faithful vendored copy** of the
upstream reader (its
`prompt_embeds_mask` emission is interleaved with cache loading, so
wrapping it would force a
    redundant per-item `torch.load`; carried verbatim instead).
- **`fastgen_checkpoint.py`** — `PartialLoadCheckpointer(Checkpointer)`
that overrides only
`load_optimizer` (FSDP2 `DefaultLoadPlanner(allow_partial_load=True)`)
so optimizer resume works
without patching upstream. Injected via an in-place re-bless of
`self.checkpointer` in the
  recipe's `load_checkpoint` (model-state load stays strict).
- **`preprocess/`** — Qwen-Image preprocessing
(`preprocessing_multiprocess.py` + `processors/`),
trimmed to the image path (drops the flux/wan/hunyuan processors and the
video base class).
It lives in AutoModel's top-level `tools/` tree, which is **not**
shipped in the pip package, so
it cannot be wrapped and is vendored; `MultiTierBucketCalculator` is
imported from stock upstream.
- **`make_negative_prompt_embedding.py`** — generates the optional CFG
negative-prompt embedding.
- All `configs/*.yaml` target `fastgen_data.build_*` (a test enumerates
every config).
- Licensing: the AutoModel-copied files are NVIDIA-authored Apache-2.0,
so they carry only the
standard NVIDIA SPDX header (managed by the `insert-license` hook) — no
per-file provenance note,
no duplicated license, no pre-commit exclusion, and no separate
`LICENSE` note.
  `nemo_automodel[diffusion]` version bound in `requirements.txt`.

The DMD2 math in `modelopt/torch/fastgen/` is **unchanged** — only
example/training-time glue moved.

### Usage

```bash
# Install example deps (stock nemo_automodel) from a source checkout
pip install -r examples/diffusers/fastgen/requirements.txt

# Build the training cache from raw images (Qwen-Image VAE latents + text embeddings)
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
    --image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
    --caption_format meta_json

# Generate the CFG negative-prompt embedding once
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
    --output <cache dir>/negative_prompt_embedding.pt

# Point the config's data.dataloader.cache_dir + negative_prompt_embedding_path at the cache, then train.
```

See `examples/diffusers/fastgen/README.md` → "Requirements &
self-contained data path".

### Testing

- New `tests/examples/diffusers/fastgen/test_vendored_migration.py`:
environment-independent
invariants (every config targets a vendored builder; no `tools.*`
imports; each former AutoModel
patch is vendored / wrapped / a documented exclusion; the
former-vendored files carry the standard NVIDIA SPDX header, no
provenance note or duplicate license) plus
dependency-guarded structural tests (the collate emits the batch
contract + broadcasts the
negative embedding; the builder accepts
`negative_prompt_embedding_path`; the checkpointer
overrides only `load_optimizer`; the Qwen-Image processor
self-registers).
- **Validated 9/9 against a pure stock `nemo_automodel` 0.4.0 worktree**
(none of the local patches
  present) via SLURM — re-run after this slim-down.
- The migrated code path is exercised by a live multi-GPU DMD2 run that
resumed from a checkpoint
  through the vendored `PartialLoadCheckpointer`.
- `ruff check` + `ruff format --check` clean on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive; bundled configs
target the vendored builders)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
(NeMo-AutoModel @ `e42584e3`, Apache-2.0; per review these
NVIDIA-authored files carry the standard NVIDIA SPDX header, no separate
provenance / `LICENSE` note; `nemo_automodel` was already a dependency)
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌ (pending — opened as draft)

### Additional Information

Opened as a **draft** pending: OSRB review of the vendored
NeMo-AutoModel (Apache-2.0) code, CI green,
and (optional) a multi-GPU smoke-train + resume on stock upstream.
Vendored from
NVIDIA-NeMo/Automodel at commit `e42584e3`.

### Update (post-review)

- **Mid-run resume data-correctness fix** (`6ffbc52c9`): on resume the
`StatefulDataLoader`'s restored state did not advance past the resume
point, so each window re-served the same data slice and multi-window
(SLURM-windowed) runs under-covered the dataset. Fixed by rebuilding a
fresh loader and skipping the deterministic sampler to the position
implied by `global_step`; added a SLURM-free, GPU-free CPU regression
test (`tests/examples/diffusers/fastgen/test_resume_dataloader.py`).
- **Licensing review** (`be832ae95`): the AutoModel-copied files are
NVIDIA-authored Apache-2.0, so they now carry only the standard NVIDIA
SPDX header — dropped the per-file provenance note, the duplicated
original-license block, the `insert-license` pre-commit exclusion, and
the `LICENSE` note.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added vendored Qwen‑Image preprocessing (multi-process) with processor
registry support.
* Updated DMD2 data loading with a dedicated dataset/collation pipeline,
including negative-prompt embedding/mask handling.
* Improved training resume behavior by rebuilding dataloader state and
making checkpoint optimizer restore tolerant of partial FSDP2 optimizer
shards.
* **Documentation**
* Refreshed the fastgen README and config notes for real-data training;
removed the prior mock-data smoke workflow.
* **Tests**
* Added regression and migration tests covering vendored wiring,
collate/dataloader contracts, processor registration, and
resume/checkpoint behavior.
* **Chores**
* Updated licenses/attribution, vendoring/tooling guards, requirements
pinning, linting configuration, and repository ownership rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-22 16:22:38 -07:00
Keval Morabia a21197c277 Remove unsafe torch.load from examples/diffusers/fastgen (#1740)
Follow-up to #1326 - Remove unsafe `torch.load(..., weights_only=False)`
in `examples/diffusers/fastgen`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated checkpoint loading mechanisms in FastGen examples to improve
compatibility and reliability during model restoration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-16 08:44:13 +05:30
jingyu-mlandClaude Opus 4.8 aec72ffa68 Add DMD2 distillation for Qwen-Image (fastgen) (#1326)
### What does this PR do?

**Type of change:** New example + new `modelopt.torch.fastgen` library
module.

Adds **DMD2 (Distribution Matching Distillation) for Qwen-Image** —
distilling the base model into a few-step (1–4) generator. Includes the
framework-agnostic `modelopt.torch.fastgen` loss library (DMD pipeline,
EMA, optional GAN discriminator) and a NeMo AutoModel–based training
example with a mock-data smoke config, a real-data config, and inference
/ export scripts.

**Noted**: the example script will be migrated to AutoModel repo

### Usage

```bash
# Mock-data wiring smoke — runs end-to-end with no dataset to prepare
torchrun --nproc-per-node=8 \
    examples/diffusers/fastgen/dmd2_finetune.py \
    --config examples/diffusers/fastgen/configs/dmd2_qwen_image_smoke.yaml
```

See `examples/diffusers/fastgen/README.md` for real-data training and
inference.

### Testing

Unit tests under `tests/unit/torch/fastgen/`; `pre-commit` /
code-quality clean.

### Before your PR is "*Ready for review*"

- Backward compatible?: ✅ (new, additive module)
- Followed `CONTRIBUTING.md` for any copied code / new deps: ✅
- New tests added?: ✅
- Updated Changelog?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Adds a FastGen-based distillation framework (DMD2) with
student/fake-score training, EMA support, GAN discriminator branch,
inference pipeline, and export utilities.
* Qwen-Image integration with latent packing and feature-capture for
plugin-enabled pipelines.

* **Documentation**
* New README, example configs, and runnable example scripts for
Qwen-Image distillation and inference.

* **Tests**
* Comprehensive unit tests covering math parity, gradient routing,
plugins, hooks, EMA, and recipe setup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 16:03:35 -07:00