Add day-0 verbosity gate + harden release skills from a live run (#2254)

### What does this PR do?

Type of change: Bug fix, new feature, documentation

Day-0 release treats verbosity as a hard gate, but nothing in the skill
measured it — `gate_compare.py` only does accuracy. A run could complete
Step 5 and report a publish recommendation with a gate silently
unmeasured. This adds the missing gate and folds in fixes for problems
that cost GPU-h on a live release run.

**New: `gate_verbosity.py` + Step 5b.** Pure `evaluate_verbosity()` plus
an artifact harvester, matching `gate_compare.py`'s conventions and
`failure_class` vocabulary.

Two details it encodes, both of which produced wrong verdicts before
they were understood:

- **Read `response_stats.avg_completion_tokens`.** The
`reasoning.*_tokens` fields are always `0`, which reads as "the harness
never captured tokens" and pushes you to word counts. Only the
reasoning/content *split* is missing, not the total. Words disagree with
the gate: one task read **+6.10% FAIL** in words and **+1.32% PASS** in
tokens.
- **Two run-hygiene filters.** Pooling a mismatched reasoning-effort run
reported **+43.06% FAIL** on a task that is **+1.24% PASS** matched; one
truncated run (n=200 vs 294) reported **+9.58% FAIL** on a task that is
**+1.71% PASS** without it. Tasks with no common sample count are
reported `not_comparable` rather than as a delta.

**Skill hardening**, each from a specific failure:

- **Step 2b canary** — poll ceiling must exceed load time (a 50 min poll
against a 51 min load failed a checkpoint that serves fine), print an
explicit `RESULT:` on every path (a fall-through exits 0 and reads as
PASS), log to shared storage, and canary the **as-exported** artifact
rather than a copy modified to make it work.
- **Step 4 config parity** — assert the candidate config differs from
the baseline's in nothing but checkpoint path and served-model name. A
mismatched `parallelism` was worth ~2 pp, enough to invert the sign of a
delta, and cost four re-runs.
- **Statistical power** — re-running does not guarantee fresh samples:
with a warm NEL response cache two runs came back bit-identical to 16
digits.
- **Step 6 closeout** — verify the published path against the evaluated
one by inode, and prefix rejected sibling exports.
- **Size gate** — growth is blocking by default and waived only when the
validation summary's
`source_precision` shows an already-sub-8-bit source (which cannot
shrink further under a
4-bit recipe) and the growth is within what that explains.
`source_precision` is now a
recorded field in the ptq validation table, so the waiver is reachable
from the normal
pipeline, and `SIZE_NOT_REDUCED` has a triage row pointing at declaring
it.
- **`ptq.py`** — `--calib_seq` matters more than `--calib_size`, and
`--mse_calibrate` is a no-op under `--cast_mxfp4_to_nvfp4` (it tunes
weight quantizers only, and the cast overwrites `weight_scale`).
- **`.gitignore workspaces/`** — the skills create scratch directories
inside the repo; nothing excluded them.

### Usage

```bash
python "$SKILL_DIR/scripts/gate_verbosity.py" \
    --baseline <baseline_eval_root> --candidate <candidate_eval_root> \
    --glob 'eval_*' --threshold 0.05
```

Exit codes match the sibling gates: `0` pass, `1` the gate ran and
failed, `2` the gate could not
read its input (wrong root, `--glob` matched nothing, everything
excluded). Prints per-task tokens,
delta, `within_threshold`, `sample_count`, run counts,
`dropped_mismatched_runs`, any
`truncated_comparison`, `not_comparable`, `harvest_diagnostics`, and a
`max_abs_delta` summary.

### Testing

- 8 new unit tests in `test_gates.py`, one per real failure mode
(two-sided threshold, partial-run filtering, unequal sample counts,
short-output warning, one-sided tasks, empty input). Full suite: **36
passed**, no GPU or network.
- `gate_verbosity.py` validated end-to-end against a real day-0 run's
artifacts: reproduces the hand-computed result (`max_abs_delta =
0.0171`, pass) and correctly marks the two unequal-sample tasks
`not_comparable`.
- `pre-commit run --files <changed>` clean, including ruff, mypy,
bandit, markdownlint, and the `.claude/skills` symlink sync.
- Verified `max_sample_length` is a real `get_dataset_dataloader`
parameter with default 512, matching `--calib_seq`'s default, so
existing `ptq()` callers are unaffected.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `--calib_seq` defaults to the
pre-existing 512; `gate_ptq.py` reclassifies size growth from
`QUANT_COVERAGE_FAILURE` to `SIZE_NOT_REDUCED`, which is a more precise
class for an already-4-bit source and is covered by a new test asserting
a real coverage failure still outranks it.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependencies, stdlib only.
- Did you write any new necessary tests?: ✅ — 8 new tests for the new
gate.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

The measured figures come from a completed day-0 NVFP4 release
qualification. Model-specific results were removed from the general
skills where the rule stands on its own; two references were kept
deliberately — a model card citation illustrating per-scenario sampling,
and a model-specific vLLM MoE kernel crash where the model name *is* the
evidence.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added configurable calibration sequence-length limits for DeepSeek-V4
quantization.
* Added automated release checks for output verbosity, evaluation
comparability, serving readiness, and configuration parity.

* **Bug Fixes**
* Improved quantization size-ratio reporting by distinguishing
explainable growth from blocking failures.
* Clarified handling of deployment memory-access errors and infeasible
evaluations.

* **Documentation**
* Expanded guidance for calibration, remote execution, workspace
management, evaluation setup, deployment troubleshooting, and
statistical reliability.
* Added task-specific guidance for SciCode and GDPVal feasibility
checks.

* **Chores**
  * Excluded workspace session directories from version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Keval Morabia
2026-08-27 12:52:09 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent d73278808b
commit 7ff81dd795
14 changed files with 1436 additions and 40 deletions
+41 -2
View File
@@ -53,6 +53,25 @@ Usage (single node, 4 GPUs, MP=4):
For MP=8 across two nodes use torchrun's ``--nnodes=2 --node_rank=<i>
--master_addr=<ip>`` flags.
Calibration settings
--------------------
When the amax dumped here is later consumed by ``quantize_to_nvfp4.py
--cast_mxfp4_to_nvfp4``, the expert weights are a lossless bit-cast, so the
activation amax (``input_scale``) is the only calibrated quantity that survives:
* ``--calib_seq`` sets the tokenizer truncation cap, which decides how much of a
long document survives. Raising it is **not** free on this path: the whole
corpus is tokenized in one ``padding=True`` call, so every row is padded to
the longest surviving sample, and ``calibrate_loop`` passes only ``input_ids``
-- so those pad tokens participate in calibration. Choose it from the context
length the activations must cover, and re-validate rather than assuming
higher is better.
* ModelOpt's ``mse_calibrate`` has no effect on this path — it tunes weight
quantizers only, and the MXFP4->NVFP4 cast overwrites ``weight_scale``
afterwards.
"""
from __future__ import annotations
@@ -296,7 +315,14 @@ def _build_nvfp4_experts_cfg() -> dict:
}
def ptq(model, tokenizer, batch_size: int, calib_size: int, calib_datasets: list[str]):
def ptq(
model,
tokenizer,
batch_size: int,
calib_size: int,
calib_datasets: list[str],
calib_seq: int = 512,
):
world_size = int(os.getenv("WORLD_SIZE", "1"))
rank = int(os.getenv("RANK", "0"))
@@ -311,6 +337,7 @@ def ptq(model, tokenizer, batch_size: int, calib_size: int, calib_datasets: list
tokenizer=tokenizer,
batch_size=batch_size,
num_samples=[calib_size] * len(calib_datasets),
max_sample_length=calib_seq,
device=device,
)
_trace("calib dataloader ready")
@@ -436,6 +463,16 @@ def main():
)
p.add_argument("--batch_size", type=int, default=4)
p.add_argument("--calib_size", type=int, default=64)
p.add_argument(
"--calib_seq",
type=int,
default=512,
help=(
"calibration sequence truncation cap (max_sample_length). Longer documents are "
"truncated to this; because the corpus is tokenized in one padded call, shorter "
"ones are padded up to it. Re-validate when changing it."
),
)
p.add_argument(
"--calib_dataset",
dest="calib_datasets",
@@ -471,7 +508,9 @@ def main():
tokenizer = AutoTokenizer.from_pretrained(
args.model_path, trust_remote_code=args.trust_remote_code
)
model = ptq(model, tokenizer, args.batch_size, args.calib_size, args.calib_datasets)
model = ptq(
model, tokenizer, args.batch_size, args.calib_size, args.calib_datasets, args.calib_seq
)
save_amax_and_quant_config(model, args.output_path)
if args.run_generate is not None: