mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add day-0 verbosity gate + harden release skills from a live run (#2254)
### What does this PR do?
Type of change: Bug fix, new feature, documentation
Day-0 release treats verbosity as a hard gate, but nothing in the skill
measured it — `gate_compare.py` only does accuracy. A run could complete
Step 5 and report a publish recommendation with a gate silently
unmeasured. This adds the missing gate and folds in fixes for problems
that cost GPU-h on a live release run.
**New: `gate_verbosity.py` + Step 5b.** Pure `evaluate_verbosity()` plus
an artifact harvester, matching `gate_compare.py`'s conventions and
`failure_class` vocabulary.
Two details it encodes, both of which produced wrong verdicts before
they were understood:
- **Read `response_stats.avg_completion_tokens`.** The
`reasoning.*_tokens` fields are always `0`, which reads as "the harness
never captured tokens" and pushes you to word counts. Only the
reasoning/content *split* is missing, not the total. Words disagree with
the gate: one task read **+6.10% FAIL** in words and **+1.32% PASS** in
tokens.
- **Two run-hygiene filters.** Pooling a mismatched reasoning-effort run
reported **+43.06% FAIL** on a task that is **+1.24% PASS** matched; one
truncated run (n=200 vs 294) reported **+9.58% FAIL** on a task that is
**+1.71% PASS** without it. Tasks with no common sample count are
reported `not_comparable` rather than as a delta.
**Skill hardening**, each from a specific failure:
- **Step 2b canary** — poll ceiling must exceed load time (a 50 min poll
against a 51 min load failed a checkpoint that serves fine), print an
explicit `RESULT:` on every path (a fall-through exits 0 and reads as
PASS), log to shared storage, and canary the **as-exported** artifact
rather than a copy modified to make it work.
- **Step 4 config parity** — assert the candidate config differs from
the baseline's in nothing but checkpoint path and served-model name. A
mismatched `parallelism` was worth ~2 pp, enough to invert the sign of a
delta, and cost four re-runs.
- **Statistical power** — re-running does not guarantee fresh samples:
with a warm NEL response cache two runs came back bit-identical to 16
digits.
- **Step 6 closeout** — verify the published path against the evaluated
one by inode, and prefix rejected sibling exports.
- **Size gate** — growth is blocking by default and waived only when the
validation summary's
`source_precision` shows an already-sub-8-bit source (which cannot
shrink further under a
4-bit recipe) and the growth is within what that explains.
`source_precision` is now a
recorded field in the ptq validation table, so the waiver is reachable
from the normal
pipeline, and `SIZE_NOT_REDUCED` has a triage row pointing at declaring
it.
- **`ptq.py`** — `--calib_seq` matters more than `--calib_size`, and
`--mse_calibrate` is a no-op under `--cast_mxfp4_to_nvfp4` (it tunes
weight quantizers only, and the cast overwrites `weight_scale`).
- **`.gitignore workspaces/`** — the skills create scratch directories
inside the repo; nothing excluded them.
### Usage
```bash
python "$SKILL_DIR/scripts/gate_verbosity.py" \
--baseline <baseline_eval_root> --candidate <candidate_eval_root> \
--glob 'eval_*' --threshold 0.05
```
Exit codes match the sibling gates: `0` pass, `1` the gate ran and
failed, `2` the gate could not
read its input (wrong root, `--glob` matched nothing, everything
excluded). Prints per-task tokens,
delta, `within_threshold`, `sample_count`, run counts,
`dropped_mismatched_runs`, any
`truncated_comparison`, `not_comparable`, `harvest_diagnostics`, and a
`max_abs_delta` summary.
### Testing
- 8 new unit tests in `test_gates.py`, one per real failure mode
(two-sided threshold, partial-run filtering, unequal sample counts,
short-output warning, one-sided tasks, empty input). Full suite: **36
passed**, no GPU or network.
- `gate_verbosity.py` validated end-to-end against a real day-0 run's
artifacts: reproduces the hand-computed result (`max_abs_delta =
0.0171`, pass) and correctly marks the two unequal-sample tasks
`not_comparable`.
- `pre-commit run --files <changed>` clean, including ruff, mypy,
bandit, markdownlint, and the `.claude/skills` symlink sync.
- Verified `max_sample_length` is a real `get_dataset_dataloader`
parameter with default 512, matching `--calib_seq`'s default, so
existing `ptq()` callers are unaffected.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--calib_seq` defaults to the
pre-existing 512; `gate_ptq.py` reclassifies size growth from
`QUANT_COVERAGE_FAILURE` to `SIZE_NOT_REDUCED`, which is a more precise
class for an already-4-bit source and is covered by a new test asserting
a real coverage failure still outranks it.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies, stdlib only.
- Did you write any new necessary tests?: ✅ — 8 new tests for the new
gate.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The measured figures come from a completed day-0 NVFP4 release
qualification. Model-specific results were removed from the general
skills where the rule stands on its own; two references were kept
deliberately — a model card citation illustrating per-scenario sampling,
and a model-specific vLLM MoE kernel crash where the model name *is* the
evidence.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added configurable calibration sequence-length limits for DeepSeek-V4
quantization.
* Added automated release checks for output verbosity, evaluation
comparability, serving readiness, and configuration parity.
* **Bug Fixes**
* Improved quantization size-ratio reporting by distinguishing
explainable growth from blocking failures.
* Clarified handling of deployment memory-access errors and infeasible
evaluations.
* **Documentation**
* Expanded guidance for calibration, remote execution, workspace
management, evaluation setup, deployment troubleshooting, and
statistical reliability.
* Added task-specific guidance for SciCode and GDPVal feasibility
checks.
* **Chores**
* Excluded workspace session directories from version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d73278808b
commit
7ff81dd795
@@ -53,6 +53,25 @@ Usage (single node, 4 GPUs, MP=4):
|
||||
|
||||
For MP=8 across two nodes use torchrun's ``--nnodes=2 --node_rank=<i>
|
||||
--master_addr=<ip>`` flags.
|
||||
|
||||
Calibration settings
|
||||
--------------------
|
||||
|
||||
When the amax dumped here is later consumed by ``quantize_to_nvfp4.py
|
||||
--cast_mxfp4_to_nvfp4``, the expert weights are a lossless bit-cast, so the
|
||||
activation amax (``input_scale``) is the only calibrated quantity that survives:
|
||||
|
||||
* ``--calib_seq`` sets the tokenizer truncation cap, which decides how much of a
|
||||
long document survives. Raising it is **not** free on this path: the whole
|
||||
corpus is tokenized in one ``padding=True`` call, so every row is padded to
|
||||
the longest surviving sample, and ``calibrate_loop`` passes only ``input_ids``
|
||||
-- so those pad tokens participate in calibration. Choose it from the context
|
||||
length the activations must cover, and re-validate rather than assuming
|
||||
higher is better.
|
||||
|
||||
* ModelOpt's ``mse_calibrate`` has no effect on this path — it tunes weight
|
||||
quantizers only, and the MXFP4->NVFP4 cast overwrites ``weight_scale``
|
||||
afterwards.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -296,7 +315,14 @@ def _build_nvfp4_experts_cfg() -> dict:
|
||||
}
|
||||
|
||||
|
||||
def ptq(model, tokenizer, batch_size: int, calib_size: int, calib_datasets: list[str]):
|
||||
def ptq(
|
||||
model,
|
||||
tokenizer,
|
||||
batch_size: int,
|
||||
calib_size: int,
|
||||
calib_datasets: list[str],
|
||||
calib_seq: int = 512,
|
||||
):
|
||||
world_size = int(os.getenv("WORLD_SIZE", "1"))
|
||||
rank = int(os.getenv("RANK", "0"))
|
||||
|
||||
@@ -311,6 +337,7 @@ def ptq(model, tokenizer, batch_size: int, calib_size: int, calib_datasets: list
|
||||
tokenizer=tokenizer,
|
||||
batch_size=batch_size,
|
||||
num_samples=[calib_size] * len(calib_datasets),
|
||||
max_sample_length=calib_seq,
|
||||
device=device,
|
||||
)
|
||||
_trace("calib dataloader ready")
|
||||
@@ -436,6 +463,16 @@ def main():
|
||||
)
|
||||
p.add_argument("--batch_size", type=int, default=4)
|
||||
p.add_argument("--calib_size", type=int, default=64)
|
||||
p.add_argument(
|
||||
"--calib_seq",
|
||||
type=int,
|
||||
default=512,
|
||||
help=(
|
||||
"calibration sequence truncation cap (max_sample_length). Longer documents are "
|
||||
"truncated to this; because the corpus is tokenized in one padded call, shorter "
|
||||
"ones are padded up to it. Re-validate when changing it."
|
||||
),
|
||||
)
|
||||
p.add_argument(
|
||||
"--calib_dataset",
|
||||
dest="calib_datasets",
|
||||
@@ -471,7 +508,9 @@ def main():
|
||||
tokenizer = AutoTokenizer.from_pretrained(
|
||||
args.model_path, trust_remote_code=args.trust_remote_code
|
||||
)
|
||||
model = ptq(model, tokenizer, args.batch_size, args.calib_size, args.calib_datasets)
|
||||
model = ptq(
|
||||
model, tokenizer, args.batch_size, args.calib_size, args.calib_datasets, args.calib_seq
|
||||
)
|
||||
save_amax_and_quant_config(model, args.output_path)
|
||||
|
||||
if args.run_generate is not None:
|
||||
|
||||
Reference in New Issue
Block a user