Files
Model-Optimizer/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml
T
h-guo18 5db2682519 [Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do?

Type of change: new example

Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI
that quantizes an exported speculative-decoding drafter to FP8 or NVFP4
— weight-only or weight+activation — with no calibration data.

It needs no modeling code either. Exported drafters such as
[`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark)
have no importable model class, so each 2-D weight is wrapped in a
throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual
`quantizer_name` patterns select over those names. Works for any drafter
layout (DSpark / DFlash / EAGLE3 / Medusa).

**Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt
formats vLLM's backend can actually serve. AWQ is deliberately not
offered, since `awq_lite` silently degrades to plain RTN without a
`forward_loop`.

**Static activation scales without calibration.** `fp8` and `nvfp4`
normally need an activation amax *measured* on calibration data; a fixed
`input_scale` of 1.0 is applied instead. That works because acceptance
length is governed almost entirely by **clipping**, not resolution:

Sweeping the fixed scale over three decades (same setup as the Testing
section below; bf16 baseline 3.1423):

| `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 |
|---|---|---|---|---|---|
| 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% |
| 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% |
| 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% |
| 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% |
| 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% |
| 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% |
| 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% |
| **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** |
**-3.91%** |
| 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% |
| 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% |

Both formats fall off a cliff below ~0.03, where the declared range sits
far under the activations' true magnitude and most of the tensor is
clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no
drop-off at the top**, so the scale only has to be big enough. 1.0 is
the middle of that plateau, which is why it is hardcoded rather than
exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau
— that gap is the 4-bit resolution cost, and no choice of scale recovers
it.

Deriving the amax from the weights instead was tried and does not work:
`max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier
channels in the tens, so the range lands 1–2 orders of magnitude low and
clips, measuring -31% to -46% AL.

**Where calibration would go.** All of this sits behind
`resolve_activation_scales()`, the single place deciding where a static
amax comes from. Real calibration slots in ahead of the fixed fallback
with no change to the CLI or the call site, and composes because
`set_static_activation_amax()` skips quantizers that already have an
amax:

```python
if calib_forward_loop is not None:
    mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop)
set_static_activation_amax(root)   # fills in what calibration did not reach
```

**Serving a quantized drafter.** Four things had to be written into the
exported checkpoint before vLLM would load one:

- emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that
key, ModelOpt writes only `quant_algo`
- emit the exclusion list under `ignore` too — that is the key read from
the flat `quantization_config`; `exclude_modules` alone yields an empty
exclusion set
- add `*<name>` wildcards so exclusions match a runtime's nested module
prefix (`model.fc`) rather than the checkpoint key (`fc`)
- add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses,
whose names appear in no checkpoint key

Nothing is then needed on the caller side. **This closes the open
question left in the previous revision of this PR: vLLM does read
`quantization_config` off the draft checkpoint.**
`ModelConfig._verify_quantization` fills `quantization` in from
`quant_method` when it is unset, so once the export declares that key —
the first fix above — detection works on its own. Verified on
Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4
checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`,
AL 4.278 against 4.203 measured earlier.

`specdec_bench` also gains a `DSPARK` algorithm, which it did not have:
an exported `Qwen3DSparkModel` would otherwise have to go through
`DFLASH` and be built with vLLM `method="dflash"`. The branch sets
`method="dspark"` and leaves `draft_sample_method` on vLLM's own default
of `greedy`. A target whose fused-collective workspace (sized at
CUDA-graph capture) overflows at large speculative batches can disable
graphs with `--runtime_params '{"engine_args": {"enforce_eager":
true}}'`.

For DFlash-family drafters, `qwen3_dflash.py` builds its fused
context-KV projection by reading `qkv_proj.weight` raw and calling
`F.linear`, which cannot consume a packed weight. Keep those layers in
bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`;
`o_proj` and the MLP — the bulk of the drafter — still quantize. That
exclusion is mandatory, not a tuning choice.

`fc` (the projection from the target's captured layers into the draft)
is the one real knob, and it is a genuine trade rather than a free win —
see the Testing section for both models' numbers. The examples quantize
it; add `'*fc*'` to the exclude list to keep it in bf16.

`embed_tokens`, `markov_head` and `confidence_head` are excluded by
default: they are 2-D so the flat view treats them as GEMMs, but they
are embeddings or a single-output projection. `lm_head` is excluded by
the preset itself — unlike on a base model it is 37% of this drafter's
parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but
measure AL first. The flag re-enables both of `lm_head`'s quantizers;
re-enabling only the weight one would ship a W+A checkpoint whose
`lm_head` has no `input_scale` while the config still advertises it as
quantized.

### Usage

```bash
# weight+activation FP8, calibration-free, lossless on both models measured below
python scripts/quantize_drafter.py \
    --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \
    --qformat fp8 \
    --export_path ./dspark-qwen3-8b-fp8 \
    --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'

# smallest: weight-only NVFP4
python scripts/quantize_drafter.py \
    --drafter_path nvidia/MiniMax-M3-DSpark \
    --qformat w4a16_nvfp4 \
    --export_path ./MiniMax-M3-DSpark-W4A16
```

Or end to end on Slurm — quantize, then measure AL — via the launcher
examples added here, one per target:

```bash
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes
```

Serving one, if you are not going through `specdec_bench`:

```python
speculative_config = {
    "method": "dspark",
    "model": "./dspark-qwen3-8b-fp8",   # quantization is read from its config.json
    "num_speculative_tokens": 7,
}
```

### Testing

Two targets with different architectures, so the conclusions are not one
model's quirk:

* **Qwen3-8B** (dense transformer) +
[`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7),
`block_size` 7, TP1.
* **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) +
[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark),
`block_size` 8, TP8, with the mamba engine settings the model card pins
(`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic
SSM-cache rounding).

Both: MT-Bench 80 questions, greedy, one vLLM instance per point.

| recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs
bf16 |
|---|---|---|---|---|---|
| bf16 baseline | — | 3.1423 | — | 4.3296 | — |
| **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** |
**4.3289** | **-0.02%** |
| `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% |
| `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% |
4.2899 | -0.92% |
| `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% |
4.2334 | -2.22% |
| **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** |
**4.2030** | **-2.92%** |

**FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on
both.** +0.11% and -0.02% are both inside run-to-run noise — the
Nemotron baseline was measured twice under identical settings and the
two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of
that column. On the same reading, `fp8` and `fp8_pc_pt` are
indistinguishable on Nemotron; the dynamic variant only pulls ahead on
Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit
weight resolution is the real price and it is model-dependent but
bounded.

Whether to quantize `fc` is a per-model call rather than a general
recommendation — it buys a few percent of size for an AL cost that
differs by ~2x between these two drafters:

| `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 |
|---|---|---|
| checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB
(-4.4%) |
| AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) |

`fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter
parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by
default.

The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the
rest of that column; the `fc`-in-bf16 run reproduced the original number
to four decimals (3.0392), so the column is internally comparable.

Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s
on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip
within 0.0952 relative error; the 29 untouched tensors are bit-identical
to `bf16(source)`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (example-only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — validated manually as
above. Can add a `tests/examples/speculative_decoding/` test over a
small synthetic drafter if wanted before merge.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (example-only)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

The measurements above are one drafter on one target with one benchmark;
the plateau's location and the ~3.5% NVFP4 gap should be re-measured
before assuming they carry to a different drafter.

Note when reading an exported checkpoint: `input_scale` is `amax/448`
for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as
1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same
activation range.

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-26 22:04:09 +08:00

89 lines
3.8 KiB
YAML

# Calibration-free NVFP4 PTQ of the DSpark drafter for Nemotron-3.5-Lightning-30B-A3B.
#
# Same pipeline as examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml, against a hybrid
# Mamba-MoE target and the published DSpark drafter. Quantizes weight-only to NVFP4,
# then measures acceptance length so the cost is visible. No calibration data needed:
# every scale comes from the weights.
#
# 2-step pipeline:
# task_0: Quantize the drafter (CPU-only, well under a minute for this 0.97B draft)
# task_1: Benchmark acceptance length on MT-Bench via vLLM
#
# Usage:
# uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes
job_name: Nemotron-3.5-Lightning-30B-A3B_DSpark_PTQ_nvfp4
pipeline:
allow_to_fail: false
skip: false
note:
global_vars:
hf_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
draft_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark
# Step 1: Quantize. w4a16_nvfp4 keeps activations in bf16; use --qformat nvfp4
# for weight+activation (fixed input_scale 1.0).
#
# The q/k/v exclusions are mandatory, not a tuning choice: DFlash-family
# drafters build their fused context-KV projection by reading qkv_proj.weight
# raw, which cannot be a packed tensor. o_proj and the MLP still quantize.
# `fc` is optional -- quantizing it saves ~4% more size for ~1.3% AL on this
# model (4.2334 vs 4.2899); add '*fc*' to keep it in bf16.
#
# This drafter has has_lm_head=false (it shares the target's), so
# --quantize_lm_head does not apply. embed_tokens is 37% of the checkpoint but
# is excluded by default: it is an nn.Embedding the drafter inherits.
task_0:
script: common/specdec/quantize_drafter.sh
args:
- --qformat w4a16_nvfp4
- --export_path /scratchspace/export_quantized
- --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'
environment:
- DRAFTER_CKPT: <<global_vars.draft_model>>
slurm_config:
_factory_: "slurm_factory"
nodes: 1
ntasks_per_node: 1
gpus_per_node: 1
container: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc20
# Step 2: Acceptance length on MT-Bench. Compare against the same run with
# --draft_model_dir <<global_vars.draft_model>> to see what quantization cost.
#
# This target is a hybrid Mamba-MoE model, so it differs from the Qwen3 example
# in three ways. It needs TP8. Startup is dominated by Mamba2 kernel warmup
# (~6 min before the engine is ready), not weight loading. And it needs the
# mamba engine settings from the model card (runtime_params below): without
# them the first draft token is rejected ~88% of the time and acceptance length
# collapses from ~4.3 to ~1.5, while every later position stays normal -- so it
# looks like a bad drafter rather than a serving misconfiguration.
task_1:
script: common/specdec_bench/quick_check.sh
args:
- --draft_model_dir /scratchspace/export_quantized
# DSPARK reads --block_size, not --draft_length; must match the drafter's
# block_size, which is 8 for this checkpoint.
- --block_size 8
- --output_length 4096
- --engine VLLM
- --tp_size 8
- --ep_size 1
- --speculative_algorithm DSPARK
- --trust_remote_code
- --temperature 0
- --runtime_params modules/Model-Optimizer/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/engine_args.json
- --mtbench /hf-local/HuggingFaceH4/mt_bench_prompts/raw/question.jsonl
- --concurrency 8
environment:
- HF_MODEL_CKPT: <<global_vars.hf_model>>
# The model card sets this on every Nemotron-3.5 serve command.
- VLLM_ALLOW_LONG_MAX_MODEL_LEN: "1"
slurm_config:
_factory_: "slurm_factory"
nodes: 1
ntasks_per_node: 1
gpus_per_node: 8
container: vllm/vllm-openai:nightly