Files
Keval MorabiaandClaude Opus 5 449a39922b Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do?

Type of change: Bug fix

`nemo_automodel` 0.6.0 removed
`nemo_automodel.recipes.diffusion.train.is_main_process` without a
replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0
defines no equivalent anywhere in the package).
`examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's
import guard fires and **every** test in `tests/examples/diffusers/`
errors at collection:

```
ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train'
  tests/examples/diffusers/fastgen/test_resume_dataloader.py
E   ImportError: The DMD2 fastgen example requires `nemo_automodel`. ...
collected 42 items / 1 error
```

The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was
published and the `onnx (diffusers)` job started failing on every PR
(e.g. runs 33020467654, 33019418460, 33010815298, 33007613265,
33006944292 — all unrelated branches). Capping at `<0.6` restores the
tested range.

Every other `nemo_automodel` symbol the example imports still exists in
0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`,
`recipes.diffusion.train.TrainDiffusionRecipe`, and the four
`components.datasets.diffusion.*` helpers), so `is_main_process` is the
only blocker; the alternative is defining that predicate locally and
widening the cap again, which is worth doing separately if the example
is meant to track 0.6.

### Usage

```bash
pip install -r examples/diffusers/fastgen/requirements.txt
```

### Testing

Reproduced the break by diffing the published wheels: `is_main_process`
is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0
and absent from 0.6.0 (`grep -rn "def is_main_process"` over the
unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported
symbols are all still present in 0.6.0.

CI on this PR exercises the fix directly: the `onnx (diffusers)` job
installs from this requirements file and is the job that has been
failing.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing
dependency, tightened bound.
- Did you write any new necessary tests?: N/A — the existing
`tests/examples/diffusers/` suite is what this unblocks.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — dependency-pin fix for a break introduced and fixed within the
same unreleased cycle.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dependency compatibility for the FastGen diffusion example.
* Prevented installation of versions that could cause the example to
fail at startup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 05:59:02 +05:30
..

DMD2 distillation for Qwen-Image

Distill Qwen/Qwen-Image into a few-step generator with DMD2 (Distribution Matching Distillation). The distilled student produces images in as few as 1–4 sampling steps while matching the base model's output distribution. Built on modelopt.torch.fastgen and NeMo AutoModel's TrainDiffusionRecipe.

Note

Qwen-Image is a third-party model with its own license terms. Review the Qwen-Image model card before downloading or redistributing weights or derivatives.

Requirements & self-contained data path

This example runs against stock upstream nemo_automodel (>=0.4.0,<1.0; see requirements.txt) from a source checkout of Model-Optimizer — the examples/ tree is not shipped in the nvidia-modelopt pip package. Install the example dependencies with:

pip install -r examples/diffusers/fastgen/requirements.txt

Tip

Prefer not to install nemo_automodel yourself? Use the NeMo AutoModel container, which bundles it (with the diffusion extras) — then you only need a source checkout of Model-Optimizer for the examples/ tree and can skip the pip install above:

docker run --gpus all -it --rm --shm-size=8g nvcr.io/nvidia/nemo-automodel:26.04

The DMD2 data loading (fastgen_data/) and raw-image preprocessing (preprocess/) are vendored into this example (from NeMo-AutoModel, Apache-2.0) so that no modifications to nemo_automodel are required. The entry points put this directory on sys.path, so the configs reference the vendored builders as _target_: fastgen_data.build_*. The DMD2 math in modelopt/torch/fastgen/ is unchanged.

Build the training cache from raw images (Qwen-Image VAE latents + text embeddings):

python examples/diffusers/fastgen/preprocess_qwen_image.py image \
    --image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
    --caption_format meta_json

The CFG negative-prompt embedding (the config's negative_prompt_embedding_path) is generated once from the same Qwen text encoder:

python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
    --output <cache dir>/negative_prompt_embedding.pt

Then point the config's data.dataloader.cache_dir at <cache dir> and its negative_prompt_embedding_path at <cache dir>/negative_prompt_embedding.pt, and train (below).

How DMD2 works

DMD2 trains three networks together:

Model Role
Student the few-step generator you keep
Fake-score a diffusion model that tracks the student's current output distribution
Teacher the frozen base Qwen-Image model (the target distribution)

The distribution-matching gradient pushes the student toward the teacher and away from the fake-score. Training alternates between two phases, controlled by student_update_freq:

each step:
  if step % student_update_freq == 0:   # student phase
      update the student (distribution-matching [+ optional GAN] loss)
      update the student EMA
  else:                                  # fake-score phase
      update the fake-score network to track the student

The canonical config additionally enables CFG (classifier-free guidance on the teacher) and a lightweight GAN branch (a discriminator head on a teacher feature block, plus an R1 gradient penalty) for sharper samples.

Install

From the repo root:

pip install -e ".[all]"                                      # ModelOpt + torch + diffusers
pip install -r examples/diffusers/fastgen/requirements.txt   # nemo_automodel

nemo_automodel[diffusion] pulls in diffusers, accelerate, and the TrainDiffusionRecipe this example subclasses.

Real-data training

configs/dmd2_qwen_image.yaml is the canonical config: 4-step student, CFG, and the GAN + R1 branch, trained on a preprocessed latent cache. Before launching, provide:

  • A preprocessed Qwen-Image latent cache — set data.dataloader.cache_dir.
  • A precomputed negative-prompt embedding (required for CFG) — set data.dataloader.negative_prompt_embedding_path.
  • An output directory — set checkpoint.checkpoint_dir.

The model path defaults to Qwen/Qwen-Image; point it at a local snapshot to avoid re-downloading on every job. Then:

torchrun --nproc-per-node=8 \
    examples/diffusers/fastgen/dmd2_finetune.py \
    --config examples/diffusers/fastgen/configs/dmd2_qwen_image.yaml \
    --step_scheduler.max_steps=5000

Any DMDConfig field can be overridden on the CLI (e.g. --dmd2.guidance_scale=3.5).

Checkpoints & resuming

Checkpoints land under checkpoint.checkpoint_dir. Alongside the student, the recipe saves the DMD2 sidecars needed to resume exactly: the fake-score model + optimizer, the student EMA (ema_shadow.pt), and the DMD iteration counter (dmd_state.pt). With restore_from: LATEST a re-launch auto-resumes from the newest checkpoint; pin a specific one with --checkpoint.restore_from=epoch_0_step_500.

Inference

After training, sample from the distilled student. The pipeline loads your consolidated student transformer plus the base Qwen-Image VAE / text encoder / tokenizer:

import torch
from inference_dmd2_qwen_image import QwenImageDMDInferencePipeline

pipe = QwenImageDMDInferencePipeline.from_pretrained(
    student_path="/path/to/checkpoint/epoch_0_step_500/model/consolidated",
    base_pipeline_path="Qwen/Qwen-Image",
    ema_path=None,                      # or ".../ema_shadow.pt" to sample the EMA weights
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt="a small red cube on a white table",
    num_inference_steps=4,              # match the student_sample_steps you trained with
    height=1024, width=1024,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("sample.png")

Or run the bundled CLI for a quick check:

python examples/diffusers/fastgen/inference_dmd2_qwen_image.py \
    --student_path /path/to/checkpoint/.../model/consolidated \
    --base_pipeline_path Qwen/Qwen-Image \
    --prompt "a small red cube on a white table" \
    --height 512 --width 512

Set num_inference_steps to the number of steps the student was trained for (dmd2.student_sample_steps — e.g. 4 for the canonical config, or 1 for a single-step student).

Config reference

Section Key Role
model pretrained_model_name_or_path Qwen-Image HF id or local snapshot.
model mode finetune — loads the pretrained weights.
step_scheduler global_batch_size, local_batch_size, max_steps, ckpt_every_steps, log_every Standard AutoModel scheduling knobs.
dmd2 recipe_path Built-in fastgen recipe to hydrate DMDConfig from (general/distillation/dmd2_qwen_image).
dmd2 pipeline_plugin qwen_image — selects QwenImageDMDPipeline (2×2 patch packing / img_shapes).
dmd2 student_sample_steps Number of student sampling steps (e.g. 4).
dmd2 guidance_scale CFG strength on the teacher (null disables CFG; requires a negative-prompt embedding when set).
dmd2 gan_loss_weight_gen, gan_r1_reg_weight, gan_feature_indices, … GAN branch (set gan_loss_weight_gen: 0 to disable).
dmd2 fake_score_lr, discriminator_lr Separate LRs for the fake-score / discriminator optimizers.
dmd2 sample_t_cfg, ema Timestep sampling + student EMA settings.
optim learning_rate, optimizer.* Student AdamW knobs.
fsdp dp_size, tp_size, activation_checkpointing, … FSDP2 parallelism (set dp_size to your GPU count).
data dataloader._target_, cache_dir, negative_prompt_embedding_path Latent cache dir + optional CFG negative-prompt embedding.
checkpoint checkpoint_dir, model_save_format, restore_from Output dir, save format, resume behavior.

Troubleshooting

CUDA out of memory. Training holds three Qwen-Image transformers (student + teacher

  • fake-score) plus optimizer state. Shard across more GPUs (raise --fsdp.dp_size), or enable --fsdp.activation_checkpointing=true.

Loss is NaN on step 0. Almost always an out-of-range timestep — confirm you haven't overridden dmd2.pred_type away from flow (Qwen-Image is a rectified-flow model) or changed the timestep schedule.

guidance_scale is set but negative_encoder_hidden_states was not provided. CFG needs a precomputed negative-prompt embedding. Set data.dataloader.negative_prompt_embedding_path, or set dmd2.guidance_scale: null to disable CFG.

Dataloader yields empty batches. Ensure your cache has at least local_batch_size * fsdp.dp_size items; the distributed sampler drops incomplete batches.

Reference