mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper for recording a script run on an MLflow tracking server, and wires `examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a PTQ run can be reproduced from its MLflow entry alone. Without the flag, behavior is unchanged — every hook is gated on it. The logger lives in the library rather than the example so other scripts can record runs the same way: it takes a tracking URI, an experiment name and an explicit `enabled` flag, with params, tags and artifacts passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its params, the resolved recipe, the quantization summaries). `mlflow` is an optional dependency, imported only once tracking is enabled, so it is not a new requirement for the library. The run is opened **before the model loads**, so a bad URI or an unreachable server fails in seconds rather than after hours of calibration. The invocation and the recipe are uploaded at that point too, which keeps a crashed run useful: it is still recorded, with status `FAILED` and its log attached. Uploaded artifacts: | Artifact | Contents | | --- | --- | | `command.txt` | The full invocation, copy-pasteable | | `version.txt` | The ModelOpt version that ran (also a searchable tag) | | `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s expanded | | `logs/hf_ptq.log` | Everything the run printed, including a crash traceback | | `summary/quant_summary.txt` | Per-quantizer summary (unless `--no-verbose`) | | `summary/moe.html` | Per-expert calibration token counts, when the run produces them | Plus model / format / calibration settings as searchable params, and `user` / `hostname` / `modelopt_version` / `git_sha` tags. Three design points worth review: 1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a directory or use `$import`s, so the source file is not self-contained. For `huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe` the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved — the raw file records under 30% of what actually ran. 2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log is produced by teeing stdout/stderr. Handlers that libraries bound to `sys.stderr` at import time are re-pointed at the tee for the run's duration and handed back afterwards; without that, `transformers` / `huggingface_hub` warnings reach the console but never the log. Native (C-level) output is still not captured — documented in the README. 3. **The recipe upload lives in the caller, not the library.** That keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` → `modelopt.torch.quantization` → `modelopt.torch.utils` import cycle. 4. **MLflow failures never fail the quantization.** Startup validation is fatal by design (it is before any GPU work); the end-of-run upload is best-effort. Only the main rank uploads, so `--use_fsdp2` runs produce a single run. ### Usage ```bash python hf_ptq.py \ --pyt_ckpt_path <huggingface_model_card> \ --recipe general/ptq/nvfp4_default-kv_fp8_cast \ --export_path <quantized_ckpt_path> \ --mlflow https://<your-mlflow-server>/ ``` ``` [mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name> [mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e... ``` `--mlflow_experiment` and `--mlflow_run_name` override the defaults (`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`. Authentication uses MLflow's own env vars. ### Testing **Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven against a stub module). Covers experiment-name derivation and sanitization, URI accept/reject, tee pass-through, the pre-bound-handler redirect, artifact renaming, skipping absent optional outputs, the disabled path, and `version.txt`. 85 tests pass together with the existing `test_hf_ptq_args.py` / `test_example_utils.py`. **Hardware** — real PTQ runs against a live MLflow server: | Run | Result | | --- | --- | | Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane post-quant generations | | Qwen3.6-35B-A3B MoE, AutoQuantize `w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min, search hit `effective bits: 6.00`; 106 KB log capturing every per-layer decision, 4.4 MB quant summary | | Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` | | Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all five artifacts including `version.txt` | | Run **without** `--mlflow` after the review fixes | exactly 1 `[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked path is untouched | | Two runs sharing one `--export_path`, second crashed early | second run uploads **no** summary — the first run's 124 KB file on disk is correctly not attributed to it, and its traceback is in the log | | Crash mid-run (gated HF dataset) | `FAILED` recorded with log + traceback attached, summaries correctly absent | | Malformed URI | Rejected by `argparse` with a `Did you mean https://…?` hint | | Unreachable host | Fails in 9.9 s total, before any model load | | No `--mlflow` | Exit 0, no MLflow output, unchanged export | **Coverage gap, stated plainly:** `summary/moe.html` is verified only against a synthetic file (unit test + a real upload). It could not be produced naturally — `expert_token_count` buffers live on `_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused `_QuantFusedExperts` path, so no such file is written for these models regardless of `--moe_calib_experts_ratio`. The uploader's conditional is correct; the branch simply had no natural input available here. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — new optional flags only; no `--mlflow` means no behavior change. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds `mlflow` as an optional extra in `pyproject.toml` (`nvidia-modelopt[mlflow]`, folded into `all`) and to `examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported lazily, so it is not required to install or import ModelOpt. No code copied from other sources. - Did you write any new necessary tests?: ✅ — 29 new unit tests. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — 0.47 New Features. - Did you get Claude approval on this PR?: ❌ — `/claude review` not yet run. A self-review was done first and its six findings are fixed in the third commit (the notable one: gathering the MLflow inputs re-read the recipe on *every* run, including without `--mlflow`). ### Additional Information The one deliberate coverage gap is `summary/moe.html`, described under Testing: no model available here takes the sparse-sequential MoE path that writes it, so it is covered by unit test and a synthetic upload rather than a natural one. The uploader treats it as an optional output and skips it when absent, which is exercised by test. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>