6 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
Keval MorabiaandClaude Opus 5 4eb86524f0 [1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do?

Type of change: refactor (no functional change)

**[1/2] of a split. Merge this first; #2514 is [2/2] and is based on
this branch.**

Three example scripts had each reimplemented the same MLflow wiring: the
flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the
params/tags/artifacts a run uploads, and the open/close dance with its
status. The copies had already drifted — only `hf_ptq` wrote a
provenance pointer, only `vllm_serve` republished the resolved URI — and
every new tracked script meant another copy.

What a script records is now one declarative `Tool` record, **declared
in the script itself, beside the flags it reads**:

```python
# examples/megatron_bridge/quantize.py
QUANTIZE = Tool(
    name="megatron_bridge_quantize",
    tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...",
    variant_help="recipe name, or --quant_cfg if no --recipe",
    variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"),
    model=lambda args: args.hf_model_name_or_path,
    checkpoint=lambda args: args.export_megatron_path,
    texts=lambda args: resolved_recipe_texts(args.recipe),
    outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"},
)
```

`tracked_run` takes that record and runs the whole thing, so a script
adds tracking in three lines: `add_mlflow_args(parser, TOOL)`,
`resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args,
TOOL):`. The shared module knows no script's flags.

`examples/hf_ptq`, `examples/vllm_serve` and
`examples/megatron_bridge/quantize.py` move onto it. Three helpers fall
away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s
two flag pass-throughs).

### Usage

No user-facing change. The flags, their spellings and the experiment
naming are exactly as before; a script author now writes a `Tool`
instead of four functions.

### Testing

- `tests/unit/torch/utils/test_mlflow.py`,
`tests/examples/hf_ptq/test_hf_ptq_args.py`,
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the
only lane that runs it), which drives `quantize.py` for real: **34
passed**, locally and in this PR's `megatron` lane.
- `pre-commit run --files <changed>`: all hooks pass.
- The four suites shared four copies of a stand-in for the `mlflow`
module, which had drifted — one recorded artifacts as a list, another as
a dict, a third made `log_artifact` a no-op, so a test asserting on an
upload asserted nothing. They now share one
`tests/_test_utils/mlflow.py`, which also emulates the fluent API's
habit of opening a run when none is active.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `track_run` and
`checkpoint_run_tags` are removed, but neither shipped in a release
(0.47.0's `__all__` is `MlflowRunLogger`, `command_text`,
`current_user`, `default_experiment_name`, `validate_tracking_uri`, all
unchanged here). Three deliberate behaviour changes, each in shared code
and each tested:
- The `source_checkpoint_path` tag resolves to an absolute path where it
recorded the raw argument, which a chain of runs needs to join on the
pair. `run_tags` is shared, so this applies to every script that writes
the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local
`--hf_model_name_or_path`. A source that names no directory, such as a
Hub `org/name` id, is still recorded as given.
- `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a
block ending in `SystemExit(0)` as `FINISHED` where it recorded
`FAILED`, since a script that ends by calling `sys.exit()` rather than
returning has still finished.
- `.experiment.json`'s `tracking_uri` and the `run_url` built from it
drop a trailing `/` from the tracking URI, so the link is
`https://host/#/...` rather than `https://host//#/...`. Only reachable
by constructing `MlflowRunLogger` directly; every CLI path already
stripped the slash in `resolve_tracking_uri`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-visible change; the entry is in [2/2].
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split out of #2514. This half is the enabling refactor with no behaviour
change; #2514 is the feature it unlocks and is based on this branch. At
~605 changed lines of core logic it is over the ~500 guideline; the
owner accepted a two-PR split rather than three, and everything #2514
alone consumes — `split_tracking_credentials`,
`log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands
there rather than here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 19:57:03 +00:00
Keval MorabiaandClaude Opus 5 f2f0d6958e Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?

Type of change: new feature

`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.

Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:

- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.

Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.

One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.

The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).

### Usage

```bash
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --recipe general/ptq/nvfp4_default-kv_fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
    --mlflow https://<your-mlflow-server>/

# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```

The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.

### Testing

- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.

- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:25:30 -07:00
Chenjie LuoandClaude Opus 5 d69e93a72b Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?

Type of change: new feature

A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.

A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.

Two deliberate behaviours:

- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.

A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.

### Usage

```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
    --export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```

```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
  "tracking_uri": "https://<your-mlflow-server>",
  "experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
  "experiment_id": "36",
  "run_id": "7bec239a3a154970b062f3024a5ff20e",
  "run_name": "20260910-175422",
  "run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```

```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```

### Testing

**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.

**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.

**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:

- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 22:24:37 +00:00
Chenjie Luo b96841db3e Add optional MLflow tracking to the vLLM fake-quant server (#2120)
### What does this PR do?

Type of change: new feature

Wires `examples/vllm_serve/vllm_serve_fakequant.py` up to
`modelopt.torch.utils.mlflow` via `--mlflow <tracking-uri>`, the same
way #2023 did for `hf_ptq.py`, so a fake-quant serve records **what it
actually quantized** and an evaluation of that endpoint can be traced
back to a recipe. Without the flag, behavior is unchanged — every hook
is gated on it.

Three design points worth review:

1. **The run is recorded in the vLLM worker, not the launcher.**
`vllm_serve_fakequant.py` is the API-server frontend; the engine and its
workers are separate processes whose stdout it never sees, so a run
opened there would capture none of the calibration. The launcher instead
only settles the tracking configuration — validating the URI, naming the
experiment, recording the command the user actually typed — and
publishes it through the environment, which is how every other setting
in this example (`QUANT_CFG`, `RECIPE_PATH`, …) already reaches the
workers. Global rank 0 opens the run, so a TP-8 serve produces one run.

2. **The run covers load-through-warm-up, not the server's lifetime.**
It opens *before the weights load*, so an unreachable server or a
missing token fails in seconds rather than after a load and a full
calibration, and it closes `FINISHED` once the model is quantized and
warmed up. A run that stayed open for the serving lifetime would never
close cleanly on SIGTERM.

3. **`recipe/quant_cfg.yaml` is only written on the preset path.** With
`RECIPE_PATH`, `get_quant_config` returns the recipe's `quantize`
section unchanged and `resolved_recipe.yaml` already carries it. With
`QUANT_CFG`/`KV_QUANT_CFG` it is the *only* record of what ran: the
params carry the preset names, while the config reaching `mtq.quantize`
is those two deep-copied, merged, and — for an MLA model — extended at
runtime with `*kv_c_bmm_quantizer` / `*k_pe_bmm_quantizer` by inspecting
the loaded model.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The launcher's invocation, copy-pasteable, credentials
masked |
| `version.txt` | The ModelOpt version that ran |
| `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s
expanded |
| `recipe/quant_cfg.yaml` | Merged `QUANT_CFG`/`KV_QUANT_CFG` + MLA
fixup (preset path only) |
| `logs/<script>.log` | The rank-0 worker's stdout/stderr, including a
crash traceback |
| `summary/quant_summary.txt` | The per-quantizer summary |

Plus the quantization *and* serving settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` / `vllm_version`
tags. The `checkpoint_path` tag matches the one `hf_ptq.py` sets, so a
checkpoint's PTQ run and every serve of it join up.

Two small library additions, both consumed by the new example module:

- `command_text(argv=None)` — records another process's invocation,
since a spawned worker's own `sys.argv` is vLLM plumbing rather than
anything a user typed.
- `MlflowRunLogger.log_text()` — uploads a value settled midway through
a run, so a crash during calibration still keeps the config that caused
it.

The example `Dockerfile` installs the `mlflow` extra; the client remains
optional and is imported only once tracking is enabled.

### Usage

```bash
RECIPE_PATH=<recipe.yaml> python vllm_serve_fakequant.py <model_path> -tp 8 \
  --host 0.0.0.0 --port 8000 \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] tracking to https://<your-mlflow-server>, experiment $USER/vllm_serve_fakequant/<model>-<recipe>
(Worker_TP0) [mlflow] run: https://<your-mlflow-server>/#/experiments/19/runs/1c6679448f25...
```

`--mlflow-experiment` / `--mlflow-run-name` override the defaults.
`$MLFLOW_TRACKING_URI` enables tracking on its own and is best-effort;
an explicit `--mlflow` overrides it and fails loudly.

> This is the **quantization** tracking server. It is unrelated to any
server an evaluation harness exports its scores to — NeMo Evaluator
Launcher has its own `export.mlflow.tracking_uri`. The README calls this
out.

### Testing

**Unit — 87 passing**
(`tests/examples/vllm_serve/test_vllm_mlflow_utils.py`, 33 new;
`tests/unit/torch/utils/test_mlflow.py`, +5). `vllm_mlflow_utils`
deliberately imports no vLLM, so the whole launcher→worker handover is
covered without a GPU, a server, or the mlflow client.

**End to end on aws-cmh** (4× GB300, `simple_evals.gpqa_diamond`,
Nemotron-3.5-Lightning-30B-A3B-BF16 fake-quantized with
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`): run `FINISHED` in 261.5 s,
opened by `Worker_TP0` only, all artifacts present and verified by
content — `command.txt` held the launcher's invocation rather than the
worker's spawn argv, and `resolved_recipe.yaml` was 6797 B against 1845
B of source. 104 quantizers enabled (92 NVFP4 dynamic block-16 expert
weight/input with calibrated amax, 12 FP8 KV bmm). The eval then ran to
completion against the served endpoint, 22/22 requests HTTP 200.

Two bugs the hardware run caught, both fixed here with regression tests:

- `--mlflow_run_name` was rejected. vLLM's
`FlexibleArgumentParser.parse_args` rewrites **every** `--foo_bar` to
`--foo-bar` before matching, so a flag registered only under the
underscored spelling is unreachable from its CLI. Both spellings are now
registered. A unit test on a plain `ArgumentParser` could not have
caught this.
- `recipe/quant_cfg.yaml` uploaded a Python `repr` blob under a `.yaml`
name: a recipe's `quantize` is a `QuantizeConfig`, `yaml.safe_dump`
raises `RepresenterError` on it, and the old JSON fallback stringified
the object. `_dump_yaml` now unwraps pydantic via
`model_dump(mode="json")` and raises otherwise, with the caller
downgrading that to a warning so a bad config cannot take down a serve.

**Known coverage gap:** the preset (`QUANT_CFG`/`KV_QUANT_CFG`) path —
the only one that now writes `recipe/quant_cfg.yaml` — is covered by
unit test but has not been exercised on hardware; the canary used
`RECIPE_PATH`. Likewise the case where `$MLFLOW_TRACKING_URI` is present
*inside* the deployment container and `--mlflow` overrides it is
unit-tested only: NeMo Evaluator Launcher forwards only declared env
vars, so the eval server's URI never entered the container in the
canary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency. Uses the existing optional `nvidia-modelopt[mlflow]` extra
(`mlflow-skinny`, Apache-2.0) added in #2023; the example `Dockerfile`
now installs it. No code copied from other sources.
- Did you write any new necessary tests?: ✅ — 38 new tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 Misc.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run.

### Additional Information

Follows #2023, which added `MlflowRunLogger` and the `hf_ptq.py`
integration.

Note for anyone tracking from an OCI cluster:
`mlflow-modelopt.nvidia.com` is unreachable from oci-nrt and oci-hsg.
TCP 443 completes and the connection is then reset on the first
application byte, regardless of SNI or protocol, one RTT away — the PDX
PaaS ingress appears to apply a source-IP policy, and the OCI clusters
egress from Oracle-owned addresses (`155.248.190.0`, `168.110.199.1`)
rather than NVIDIA's. gcp-nrt, aws-cmh and cw-dfw all reach it. This is
an infrastructure matter, not a property of this change, but it
determines where the feature is usable today.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added optional MLflow tracking for vLLM fake-quantization serving
runs.
* Records serving, quantization, worker, and invocation metadata,
including configuration and summary artifacts.
* Supports tracking URI, credentials, environment, and command-line
configuration.
  * Added command and text artifact logging for active MLflow runs.
* **Documentation**
* Documented setup, configuration, recorded artifacts, lifecycle, and
fallback behavior.
  * Updated the example container to include MLflow support.
* **Tests**
* Added comprehensive coverage for tracking configuration, logging,
failures, and disabled tracking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-08-13 00:00:21 +00:00
Chenjie LuoandClaude Opus 5 77dbeb1872 Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do?

Type of change: new feature

Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper
for recording a script run on an MLflow tracking server, and wires
`examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a
PTQ run can be reproduced from its MLflow entry alone. Without the flag,
behavior is unchanged — every hook is gated on it.

The logger lives in the library rather than the example so other scripts
can record runs the same way: it takes a tracking URI, an experiment
name and an explicit `enabled` flag, with params, tags and artifacts
passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its
params, the resolved recipe, the quantization summaries). `mlflow` is an
optional dependency, imported only once tracking is enabled, so it is
not a new requirement for the library.

The run is opened **before the model loads**, so a bad URI or an
unreachable server fails in seconds rather than after hours of
calibration. The invocation and the recipe are uploaded at that point
too, which keeps a crashed run useful: it is still recorded, with status
`FAILED` and its log attached.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The full invocation, copy-pasteable |
| `version.txt` | The ModelOpt version that ran (also a searchable tag)
|
| `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s
expanded |
| `logs/hf_ptq.log` | Everything the run printed, including a crash
traceback |
| `summary/quant_summary.txt` | Per-quantizer summary (unless
`--no-verbose`) |
| `summary/moe.html` | Per-expert calibration token counts, when the run
produces them |

Plus model / format / calibration settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` tags.

Three design points worth review:

1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a
directory or use `$import`s, so the source file is not self-contained.
For
`huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe`
the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved —
the raw file records under 30% of what actually ran.
2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log
is produced by teeing stdout/stderr. Handlers that libraries bound to
`sys.stderr` at import time are re-pointed at the tee for the run's
duration and handed back afterwards; without that, `transformers` /
`huggingface_hub` warnings reach the console but never the log. Native
(C-level) output is still not captured — documented in the README.
3. **The recipe upload lives in the caller, not the library.** That
keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would
otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` →
`modelopt.torch.quantization` → `modelopt.torch.utils` import cycle.
4. **MLflow failures never fail the quantization.** Startup validation
is fatal by design (it is before any GPU work); the end-of-run upload is
best-effort.

Only the main rank uploads, so `--use_fsdp2` runs produce a single run.

### Usage

```bash
python hf_ptq.py \
  --pyt_ckpt_path <huggingface_model_card> \
  --recipe general/ptq/nvfp4_default-kv_fp8_cast \
  --export_path <quantized_ckpt_path> \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name>
[mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e...
```

`--mlflow_experiment` and `--mlflow_run_name` override the defaults
(`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start
time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`.
Authentication uses MLflow's own env vars.

### Testing

**Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the
library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the
hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven
against a stub module). Covers experiment-name derivation and
sanitization, URI accept/reject, tee pass-through, the pre-bound-handler
redirect, artifact renaming, skipping absent optional outputs, the
disabled path, and `version.txt`. 85 tests pass together with the
existing `test_hf_ptq_args.py` / `test_example_utils.py`.

**Hardware** — real PTQ runs against a live MLflow server:

| Run | Result |
| --- | --- |
| Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane
post-quant generations |
| Qwen3.6-35B-A3B MoE, AutoQuantize
`w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min,
search hit `effective bits: 6.00`; 106 KB log capturing every per-layer
decision, 4.4 MB quant summary |
| Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` |
| Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all
five artifacts including `version.txt` |
| Run **without** `--mlflow` after the review fixes | exactly 1
`[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked
path is untouched |
| Two runs sharing one `--export_path`, second crashed early | second
run uploads **no** summary — the first run's 124 KB file on disk is
correctly not attributed to it, and its traceback is in the log |
| Crash mid-run (gated HF dataset) | `FAILED` recorded with log +
traceback attached, summaries correctly absent |
| Malformed URI | Rejected by `argparse` with a `Did you mean
https://…?` hint |
| Unreachable host | Fails in 9.9 s total, before any model load |
| No `--mlflow` | Exit 0, no MLflow output, unchanged export |

**Coverage gap, stated plainly:** `summary/moe.html` is verified only
against a synthetic file (unit test + a real upload). It could not be
produced naturally — `expert_token_count` buffers live on
`_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused
`_QuantFusedExperts` path, so no such file is written for these models
regardless of `--moe_calib_experts_ratio`. The uploader's conditional is
correct; the branch simply had no natural input available here.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds
`mlflow` as an optional extra in `pyproject.toml`
(`nvidia-modelopt[mlflow]`, folded into `all`) and to
`examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported
lazily, so it is not required to install or import ModelOpt. No code
copied from other sources.
- Did you write any new necessary tests?: ✅ — 29 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 New Features.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run. A self-review was done first and its six findings are fixed in the
third commit (the notable one: gathering the MLflow inputs re-read the
recipe on *every* run, including without `--mlflow`).

### Additional Information

The one deliberate coverage gap is `summary/moe.html`, described under
Testing: no model available here takes the sparse-sequential MoE path
that writes it, so it is covered by unit test and a synthetic upload
rather than a natural one. The uploader treats it as an optional output
and skips it when absent, which is exercised by test.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 11:25:58 +05:30