Files
Model-Optimizer/tests
Chenjie LuoandClaude Opus 5 77dbeb1872 Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do?

Type of change: new feature

Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper
for recording a script run on an MLflow tracking server, and wires
`examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a
PTQ run can be reproduced from its MLflow entry alone. Without the flag,
behavior is unchanged — every hook is gated on it.

The logger lives in the library rather than the example so other scripts
can record runs the same way: it takes a tracking URI, an experiment
name and an explicit `enabled` flag, with params, tags and artifacts
passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its
params, the resolved recipe, the quantization summaries). `mlflow` is an
optional dependency, imported only once tracking is enabled, so it is
not a new requirement for the library.

The run is opened **before the model loads**, so a bad URI or an
unreachable server fails in seconds rather than after hours of
calibration. The invocation and the recipe are uploaded at that point
too, which keeps a crashed run useful: it is still recorded, with status
`FAILED` and its log attached.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The full invocation, copy-pasteable |
| `version.txt` | The ModelOpt version that ran (also a searchable tag)
|
| `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s
expanded |
| `logs/hf_ptq.log` | Everything the run printed, including a crash
traceback |
| `summary/quant_summary.txt` | Per-quantizer summary (unless
`--no-verbose`) |
| `summary/moe.html` | Per-expert calibration token counts, when the run
produces them |

Plus model / format / calibration settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` tags.

Three design points worth review:

1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a
directory or use `$import`s, so the source file is not self-contained.
For
`huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe`
the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved —
the raw file records under 30% of what actually ran.
2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log
is produced by teeing stdout/stderr. Handlers that libraries bound to
`sys.stderr` at import time are re-pointed at the tee for the run's
duration and handed back afterwards; without that, `transformers` /
`huggingface_hub` warnings reach the console but never the log. Native
(C-level) output is still not captured — documented in the README.
3. **The recipe upload lives in the caller, not the library.** That
keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would
otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` →
`modelopt.torch.quantization` → `modelopt.torch.utils` import cycle.
4. **MLflow failures never fail the quantization.** Startup validation
is fatal by design (it is before any GPU work); the end-of-run upload is
best-effort.

Only the main rank uploads, so `--use_fsdp2` runs produce a single run.

### Usage

```bash
python hf_ptq.py \
  --pyt_ckpt_path <huggingface_model_card> \
  --recipe general/ptq/nvfp4_default-kv_fp8_cast \
  --export_path <quantized_ckpt_path> \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name>
[mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e...
```

`--mlflow_experiment` and `--mlflow_run_name` override the defaults
(`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start
time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`.
Authentication uses MLflow's own env vars.

### Testing

**Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the
library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the
hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven
against a stub module). Covers experiment-name derivation and
sanitization, URI accept/reject, tee pass-through, the pre-bound-handler
redirect, artifact renaming, skipping absent optional outputs, the
disabled path, and `version.txt`. 85 tests pass together with the
existing `test_hf_ptq_args.py` / `test_example_utils.py`.

**Hardware** — real PTQ runs against a live MLflow server:

| Run | Result |
| --- | --- |
| Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane
post-quant generations |
| Qwen3.6-35B-A3B MoE, AutoQuantize
`w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min,
search hit `effective bits: 6.00`; 106 KB log capturing every per-layer
decision, 4.4 MB quant summary |
| Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` |
| Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all
five artifacts including `version.txt` |
| Run **without** `--mlflow` after the review fixes | exactly 1
`[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked
path is untouched |
| Two runs sharing one `--export_path`, second crashed early | second
run uploads **no** summary — the first run's 124 KB file on disk is
correctly not attributed to it, and its traceback is in the log |
| Crash mid-run (gated HF dataset) | `FAILED` recorded with log +
traceback attached, summaries correctly absent |
| Malformed URI | Rejected by `argparse` with a `Did you mean
https://…?` hint |
| Unreachable host | Fails in 9.9 s total, before any model load |
| No `--mlflow` | Exit 0, no MLflow output, unchanged export |

**Coverage gap, stated plainly:** `summary/moe.html` is verified only
against a synthetic file (unit test + a real upload). It could not be
produced naturally — `expert_token_count` buffers live on
`_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused
`_QuantFusedExperts` path, so no such file is written for these models
regardless of `--moe_calib_experts_ratio`. The uploader's conditional is
correct; the branch simply had no natural input available here.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds
`mlflow` as an optional extra in `pyproject.toml`
(`nvidia-modelopt[mlflow]`, folded into `all`) and to
`examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported
lazily, so it is not required to install or import ModelOpt. No code
copied from other sources.
- Did you write any new necessary tests?: ✅ — 29 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 New Features.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run. A self-review was done first and its six findings are fixed in the
third commit (the notable one: gathering the MLflow inputs re-read the
recipe on *every* run, including without `--mlflow`).

### Additional Information

The one deliberate coverage gap is `summary/moe.html`, described under
Testing: no model available here takes the sparse-sequential MoE path
that writes it, so it is covered by unit test and a synthetic upload
rather than a natural one. The uploader treats it as an optional output
and skips it when absent, which is exercised by test.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 11:25:58 +05:30
..