Files
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
..
2025-06-05 13:24:07 -07:00
2026-08-13 03:59:47 +05:30
2025-06-05 13:24:07 -07:00
2026-09-10 07:20:45 +00:00
2025-10-27 11:16:14 -07:00