### What does this PR do? Type of change: Bug fix `nemo_automodel` 0.6.0 removed `nemo_automodel.recipes.diffusion.train.is_main_process` without a replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0 defines no equivalent anywhere in the package). `examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's import guard fires and **every** test in `tests/examples/diffusers/` errors at collection: ``` ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train' tests/examples/diffusers/fastgen/test_resume_dataloader.py E ImportError: The DMD2 fastgen example requires `nemo_automodel`. ... collected 42 items / 1 error ``` The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was published and the `onnx (diffusers)` job started failing on every PR (e.g. runs 33020467654, 33019418460, 33010815298, 33007613265, 33006944292 — all unrelated branches). Capping at `<0.6` restores the tested range. Every other `nemo_automodel` symbol the example imports still exists in 0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`, `recipes.diffusion.train.TrainDiffusionRecipe`, and the four `components.datasets.diffusion.*` helpers), so `is_main_process` is the only blocker; the alternative is defining that predicate locally and widening the cap again, which is worth doing separately if the example is meant to track 0.6. ### Usage ```bash pip install -r examples/diffusers/fastgen/requirements.txt ``` ### Testing Reproduced the break by diffing the published wheels: `is_main_process` is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0 and absent from 0.6.0 (`grep -rn "def is_main_process"` over the unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported symbols are all still present in 0.6.0. CI on this PR exercises the fix directly: the `onnx (diffusers)` job installs from this requirements file and is the job that has been failing. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing dependency, tightened bound. - Did you write any new necessary tests?: N/A — the existing `tests/examples/diffusers/` suite is what this unblocks. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — dependency-pin fix for a break introduced and fixed within the same unreleased cycle. - Did you get Claude approval on this PR?: ❌ 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed dependency compatibility for the FastGen diffusion example. * Prevented installation of versions that could cause the example to fail at startup. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
DMD2 distillation for Qwen-Image
Distill Qwen/Qwen-Image into a few-step
generator with DMD2 (Distribution Matching Distillation). The distilled student
produces images in as few as 1–4 sampling steps while matching the base model's
output distribution. Built on modelopt.torch.fastgen and NeMo AutoModel's
TrainDiffusionRecipe.
Note
Qwen-Image is a third-party model with its own license terms. Review the Qwen-Image model card before downloading or redistributing weights or derivatives.
Requirements & self-contained data path
This example runs against stock upstream nemo_automodel (>=0.4.0,<1.0; see
requirements.txt) from a source checkout of Model-Optimizer — the examples/ tree is not
shipped in the nvidia-modelopt pip package. Install the example dependencies with:
pip install -r examples/diffusers/fastgen/requirements.txt
Tip
Prefer not to install
nemo_automodelyourself? Use the NeMo AutoModel container, which bundles it (with the diffusion extras) — then you only need a source checkout of Model-Optimizer for theexamples/tree and can skip thepip installabove:docker run --gpus all -it --rm --shm-size=8g nvcr.io/nvidia/nemo-automodel:26.04
The DMD2 data loading (fastgen_data/) and raw-image preprocessing (preprocess/) are
vendored into this example (from NeMo-AutoModel, Apache-2.0) so that no modifications to
nemo_automodel are required. The entry points put this directory on sys.path, so the
configs reference the vendored builders as _target_: fastgen_data.build_*. The DMD2 math in
modelopt/torch/fastgen/ is unchanged.
Build the training cache from raw images (Qwen-Image VAE latents + text embeddings):
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
--image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
--caption_format meta_json
The CFG negative-prompt embedding (the config's negative_prompt_embedding_path) is generated
once from the same Qwen text encoder:
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
--output <cache dir>/negative_prompt_embedding.pt
Then point the config's data.dataloader.cache_dir at <cache dir> and its
negative_prompt_embedding_path at <cache dir>/negative_prompt_embedding.pt, and train (below).
How DMD2 works
DMD2 trains three networks together:
| Model | Role |
|---|---|
| Student | the few-step generator you keep |
| Fake-score | a diffusion model that tracks the student's current output distribution |
| Teacher | the frozen base Qwen-Image model (the target distribution) |
The distribution-matching gradient pushes the student toward the teacher and away from
the fake-score. Training alternates between two phases, controlled by student_update_freq:
each step:
if step % student_update_freq == 0: # student phase
update the student (distribution-matching [+ optional GAN] loss)
update the student EMA
else: # fake-score phase
update the fake-score network to track the student
The canonical config additionally enables CFG (classifier-free guidance on the teacher) and a lightweight GAN branch (a discriminator head on a teacher feature block, plus an R1 gradient penalty) for sharper samples.
Install
From the repo root:
pip install -e ".[all]" # ModelOpt + torch + diffusers
pip install -r examples/diffusers/fastgen/requirements.txt # nemo_automodel
nemo_automodel[diffusion] pulls in diffusers, accelerate, and the TrainDiffusionRecipe
this example subclasses.
Real-data training
configs/dmd2_qwen_image.yaml is the canonical config: 4-step student, CFG, and the
GAN + R1 branch, trained on a preprocessed latent cache. Before launching, provide:
- A preprocessed Qwen-Image latent cache — set
data.dataloader.cache_dir. - A precomputed negative-prompt embedding (required for CFG) — set
data.dataloader.negative_prompt_embedding_path. - An output directory — set
checkpoint.checkpoint_dir.
The model path defaults to Qwen/Qwen-Image; point it at a local snapshot to avoid
re-downloading on every job. Then:
torchrun --nproc-per-node=8 \
examples/diffusers/fastgen/dmd2_finetune.py \
--config examples/diffusers/fastgen/configs/dmd2_qwen_image.yaml \
--step_scheduler.max_steps=5000
Any DMDConfig field can be overridden on the CLI (e.g. --dmd2.guidance_scale=3.5).
Checkpoints & resuming
Checkpoints land under checkpoint.checkpoint_dir. Alongside the student, the recipe
saves the DMD2 sidecars needed to resume exactly: the fake-score model + optimizer, the
student EMA (ema_shadow.pt), and the DMD iteration counter (dmd_state.pt). With
restore_from: LATEST a re-launch auto-resumes from the newest checkpoint; pin a
specific one with --checkpoint.restore_from=epoch_0_step_500.
Inference
After training, sample from the distilled student. The pipeline loads your consolidated student transformer plus the base Qwen-Image VAE / text encoder / tokenizer:
import torch
from inference_dmd2_qwen_image import QwenImageDMDInferencePipeline
pipe = QwenImageDMDInferencePipeline.from_pretrained(
student_path="/path/to/checkpoint/epoch_0_step_500/model/consolidated",
base_pipeline_path="Qwen/Qwen-Image",
ema_path=None, # or ".../ema_shadow.pt" to sample the EMA weights
torch_dtype=torch.bfloat16,
).to("cuda")
image = pipe(
prompt="a small red cube on a white table",
num_inference_steps=4, # match the student_sample_steps you trained with
height=1024, width=1024,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("sample.png")
Or run the bundled CLI for a quick check:
python examples/diffusers/fastgen/inference_dmd2_qwen_image.py \
--student_path /path/to/checkpoint/.../model/consolidated \
--base_pipeline_path Qwen/Qwen-Image \
--prompt "a small red cube on a white table" \
--height 512 --width 512
Set num_inference_steps to the number of steps the student was trained for
(dmd2.student_sample_steps — e.g. 4 for the canonical config, or 1 for a single-step
student).
Config reference
| Section | Key | Role |
|---|---|---|
model |
pretrained_model_name_or_path |
Qwen-Image HF id or local snapshot. |
model |
mode |
finetune — loads the pretrained weights. |
step_scheduler |
global_batch_size, local_batch_size, max_steps, ckpt_every_steps, log_every |
Standard AutoModel scheduling knobs. |
dmd2 |
recipe_path |
Built-in fastgen recipe to hydrate DMDConfig from (general/distillation/dmd2_qwen_image). |
dmd2 |
pipeline_plugin |
qwen_image — selects QwenImageDMDPipeline (2×2 patch packing / img_shapes). |
dmd2 |
student_sample_steps |
Number of student sampling steps (e.g. 4). |
dmd2 |
guidance_scale |
CFG strength on the teacher (null disables CFG; requires a negative-prompt embedding when set). |
dmd2 |
gan_loss_weight_gen, gan_r1_reg_weight, gan_feature_indices, … |
GAN branch (set gan_loss_weight_gen: 0 to disable). |
dmd2 |
fake_score_lr, discriminator_lr |
Separate LRs for the fake-score / discriminator optimizers. |
dmd2 |
sample_t_cfg, ema |
Timestep sampling + student EMA settings. |
optim |
learning_rate, optimizer.* |
Student AdamW knobs. |
fsdp |
dp_size, tp_size, activation_checkpointing, … |
FSDP2 parallelism (set dp_size to your GPU count). |
data |
dataloader._target_, cache_dir, negative_prompt_embedding_path |
Latent cache dir + optional CFG negative-prompt embedding. |
checkpoint |
checkpoint_dir, model_save_format, restore_from |
Output dir, save format, resume behavior. |
Troubleshooting
CUDA out of memory. Training holds three Qwen-Image transformers (student + teacher
- fake-score) plus optimizer state. Shard across more GPUs (raise
--fsdp.dp_size), or enable--fsdp.activation_checkpointing=true.
Loss is NaN on step 0. Almost always an out-of-range timestep — confirm you haven't
overridden dmd2.pred_type away from flow (Qwen-Image is a rectified-flow model) or
changed the timestep schedule.
guidance_scale is set but negative_encoder_hidden_states was not provided. CFG needs
a precomputed negative-prompt embedding. Set data.dataloader.negative_prompt_embedding_path,
or set dmd2.guidance_scale: null to disable CFG.
Dataloader yields empty batches. Ensure your cache has at least
local_batch_size * fsdp.dp_size items; the distributed sampler drops incomplete batches.
Reference
- Fastgen library:
modelopt/torch/fastgen/ - Built-in recipe:
modelopt_recipes/general/distillation/dmd2_qwen_image.yaml - AutoModel recipe this example subclasses:
nemo_automodel/recipes/diffusion/train.py