Files
090b1c5114 fastgen DMD2: make the Qwen-Image example self-contained on stock nemo_automodel (#1688)
### What does this PR do?

**Type of change:** new example / refactor (example self-containment)

The published `examples/diffusers/fastgen` DMD2 Qwen-Image distillation
example previously relied on
**local, unpublished modifications to the sibling `nemo_automodel`
package** (data path, collate, a
partial-load checkpointer, and Qwen-Image preprocessing). An external
user running the published
Model-Optimizer example against **official stock `nemo_automodel`**
would hit import/attribute errors.

This PR makes the example **self-contained on stock
`nemo_automodel>=0.4.0`** — with the changes kept
**as small as possible**: the team's actual AutoModel delta is only ~430
lines, so wherever the
upstream module is importable, the delta is expressed as a thin subclass
/ small reimplementation rather than a
full-file copy.

- **`fastgen_data/`** — the DMD2 data path:
- `collate_fns.py` — **reuses the upstream `SequentialBucketSampler` and
reimplements the collate**:
it builds the DMD2 batch directly from the vendored dataset's per-item
output (`image_latents` /
`text_embeddings` / `text_embeddings_mask` + an optional broadcast
`negative_text_embeddings` for
CFG) and deliberately does **not** call upstream
`collate_fn_production`, which stacks
model-specific token keys (`clip_tokens` / `t5_tokens`) absent from the
Qwen-Image cache. The
    builder loads an optional `negative_prompt_embedding_path`.
- `text_to_image_dataset.py` — a **faithful vendored copy** of the
upstream reader (its
`prompt_embeds_mask` emission is interleaved with cache loading, so
wrapping it would force a
    redundant per-item `torch.load`; carried verbatim instead).
- **`fastgen_checkpoint.py`** — `PartialLoadCheckpointer(Checkpointer)`
that overrides only
`load_optimizer` (FSDP2 `DefaultLoadPlanner(allow_partial_load=True)`)
so optimizer resume works
without patching upstream. Injected via an in-place re-bless of
`self.checkpointer` in the
  recipe's `load_checkpoint` (model-state load stays strict).
- **`preprocess/`** — Qwen-Image preprocessing
(`preprocessing_multiprocess.py` + `processors/`),
trimmed to the image path (drops the flux/wan/hunyuan processors and the
video base class).
It lives in AutoModel's top-level `tools/` tree, which is **not**
shipped in the pip package, so
it cannot be wrapped and is vendored; `MultiTierBucketCalculator` is
imported from stock upstream.
- **`make_negative_prompt_embedding.py`** — generates the optional CFG
negative-prompt embedding.
- All `configs/*.yaml` target `fastgen_data.build_*` (a test enumerates
every config).
- Licensing: the AutoModel-copied files are NVIDIA-authored Apache-2.0,
so they carry only the
standard NVIDIA SPDX header (managed by the `insert-license` hook) — no
per-file provenance note,
no duplicated license, no pre-commit exclusion, and no separate
`LICENSE` note.
  `nemo_automodel[diffusion]` version bound in `requirements.txt`.

The DMD2 math in `modelopt/torch/fastgen/` is **unchanged** — only
example/training-time glue moved.

### Usage

```bash
# Install example deps (stock nemo_automodel) from a source checkout
pip install -r examples/diffusers/fastgen/requirements.txt

# Build the training cache from raw images (Qwen-Image VAE latents + text embeddings)
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
    --image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
    --caption_format meta_json

# Generate the CFG negative-prompt embedding once
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
    --output <cache dir>/negative_prompt_embedding.pt

# Point the config's data.dataloader.cache_dir + negative_prompt_embedding_path at the cache, then train.
```

See `examples/diffusers/fastgen/README.md` → "Requirements &
self-contained data path".

### Testing

- New `tests/examples/diffusers/fastgen/test_vendored_migration.py`:
environment-independent
invariants (every config targets a vendored builder; no `tools.*`
imports; each former AutoModel
patch is vendored / wrapped / a documented exclusion; the
former-vendored files carry the standard NVIDIA SPDX header, no
provenance note or duplicate license) plus
dependency-guarded structural tests (the collate emits the batch
contract + broadcasts the
negative embedding; the builder accepts
`negative_prompt_embedding_path`; the checkpointer
overrides only `load_optimizer`; the Qwen-Image processor
self-registers).
- **Validated 9/9 against a pure stock `nemo_automodel` 0.4.0 worktree**
(none of the local patches
  present) via SLURM — re-run after this slim-down.
- The migrated code path is exercised by a live multi-GPU DMD2 run that
resumed from a checkpoint
  through the vendored `PartialLoadCheckpointer`.
- `ruff check` + `ruff format --check` clean on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive; bundled configs
target the vendored builders)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
(NeMo-AutoModel @ `e42584e3`, Apache-2.0; per review these
NVIDIA-authored files carry the standard NVIDIA SPDX header, no separate
provenance / `LICENSE` note; `nemo_automodel` was already a dependency)
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌ (pending — opened as draft)

### Additional Information

Opened as a **draft** pending: OSRB review of the vendored
NeMo-AutoModel (Apache-2.0) code, CI green,
and (optional) a multi-GPU smoke-train + resume on stock upstream.
Vendored from
NVIDIA-NeMo/Automodel at commit `e42584e3`.

### Update (post-review)

- **Mid-run resume data-correctness fix** (`6ffbc52c9`): on resume the
`StatefulDataLoader`'s restored state did not advance past the resume
point, so each window re-served the same data slice and multi-window
(SLURM-windowed) runs under-covered the dataset. Fixed by rebuilding a
fresh loader and skipping the deterministic sampler to the position
implied by `global_step`; added a SLURM-free, GPU-free CPU regression
test (`tests/examples/diffusers/fastgen/test_resume_dataloader.py`).
- **Licensing review** (`be832ae95`): the AutoModel-copied files are
NVIDIA-authored Apache-2.0, so they now carry only the standard NVIDIA
SPDX header — dropped the per-file provenance note, the duplicated
original-license block, the `insert-license` pre-commit exclusion, and
the `LICENSE` note.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added vendored Qwen‑Image preprocessing (multi-process) with processor
registry support.
* Updated DMD2 data loading with a dedicated dataset/collation pipeline,
including negative-prompt embedding/mask handling.
* Improved training resume behavior by rebuilding dataloader state and
making checkpoint optimizer restore tolerant of partial FSDP2 optimizer
shards.
* **Documentation**
* Refreshed the fastgen README and config notes for real-data training;
removed the prior mock-data smoke workflow.
* **Tests**
* Added regression and migration tests covering vendored wiring,
collate/dataloader contracts, processor registration, and
resume/checkpoint behavior.
* **Chores**
* Updated licenses/attribution, vendoring/tooling guards, requirements
pinning, linting configuration, and repository ownership rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-22 16:22:38 -07:00
..