mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature **[2/2] of a split. Based on #2544 — merge that first; this PR's diff is only the Megatron-Bridge half.** #2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`. It was one of five scripts in that directory that write a checkpoint; the other four recorded nothing, so the provenance chain stopped at the PTQ checkpoint and a deployed model could not be traced back to the run that produced it. All five now take the same `--mlflow` / `--mlflow_experiment` / `--mlflow_run_name` flags, and **each declares what it records as a `Tool` beside its own flags** — the shared `mlflow_utils.py` knows none of them: | Script | Records | | --- | --- | | `prune_minitron.py` | command, arguments, log, `prune_score` metric, pointer | | `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | + resolved recipe, quantizer summary | | `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved config | | `export_quantized_megatron_to_hf.py` | command, arguments, log, pointer | | `export_distilled_megatron_to_hf.py` | same, one pointer per exported checkpoint | Each writes `.experiment.json` into the checkpoint it produced, and each tags what it consumed, so `prune → quantize → distill → export` is walkable both from disk and by tag query on the server. **`distill.py` opens the run and Megatron-Bridge joins it.** Its `LoggerConfig` records per-iteration metrics and the full resolved config — which a wrapper around `main()` cannot see — but nothing of `distill.py`'s own arguments and no invocation. Megatron-Bridge takes `mlflow.active_run()` when one exists, applies the tags and logs into it, so `distill_run()` opens the run on the rank Megatron-Bridge looks at (the **last** one) and the two share it. Its early exit is handled explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`, which a blanket handler would record as `FAILED`. **The library pieces that exist for that shared run land here with their first caller**, rather than in [1/2] where they would have none: `split_tracking_credentials`, so a URI handed to something which *records* it carries no credential; `log_active_run_experiment_json`, for pointing a checkpoint at a run this process did not open; and `MlflowRunLogger._reattach`, because a co-owner can end the run first — Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training. Two of Megatron-Bridge's defaults are deliberately not inherited: **checkpoint artifact upload stays off** unless `--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP after every save), and **an untracked run passes no `mlflow_*` fields at all**, since they landed in Megatron-Bridge 0.6 and sending them unconditionally would break an untracked run on an older one. ### Usage ```bash # Any of the five, same flags: torchrun --nproc_per_node 8 prune_minitron.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 quantize.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 distill.py ... --mlflow https://<server>/ torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/ # Each checkpoint names the run that wrote it: cat /output/qad/checkpoints/.experiment.json ``` Experiments default to `$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model basename>-<variant>`. ### Testing - Real runs on a toy Qwen3 in one MLflow experiment covering all five Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD distillation, quantized export, BF16 distillation, distilled export, HF PTQ — each closing `FINISHED` with the invocation, its arguments as params, its log, and a matching `.experiment.json` on disk. The chain tags line up: each stage's `source_checkpoint_path` is the previous stage's `checkpoint_path`. - `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the only lane that runs it: **76 passed**. Plus the three suites from #2544: **195 pass**. - `pre-commit run --files <changed>`: all hooks pass. - Each fix from the review rounds has a test that fails with the fix reverted: the resumed run, the foreign active run, the percent-decoded credential, the credential that cannot be moved, the rank-dependent `LoggerConfig`, the exit-callback guard, and the `iter_*` join. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: several rounds; re-requested on this head. ### Additional Information Split from a single ~1150-line PR at review's request; #2544 carries the library consolidation this builds on, and this branch is based on it. Earlier review threads here show as outdated after the rebases — they are all resolved and their fixes are in this branch. One known gap, stated in the README rather than implied: `distill.py --hf_export_path` writes a second HuggingFace checkpoint from rank 0, which is not the rank that owns the run, so it carries no pointer yet. For the same reason the uploaded `logs/distill.log` holds the last rank's output — `print_rank_0` keeps the script's own lines on rank 0 — which the README now says outright; carrying rank 0's log into a run owned by another rank needs cross-rank upload and is a follow-up. Two defects found on shared-run paths during review, both verified against the installed Megatron-Bridge 0.6 rather than its docs. Megatron-Bridge ends the run it shares with `distill.py` as `KILLED` from its SIGTERM handler (`train.py:1413`) and then leaves through `sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` — and MLflow's fluent calls resolve their target by *opening* a run when none is active, so a preempted distillation's log and metrics went to a second, empty run and its `KILLED` status was overwritten. Separately, an unreachable server disabled our logger but `logger_kwargs` still handed Megatron-Bridge the same URI, and `state.py` calls `set_experiment` unguarded from inside the training loop — so a best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of degrading to untracked. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>