Files
h-guo18 5db2682519 [Example]: Calibration-free FP8/NVFP4 PTQ for speculative-decoding drafters (#2027)
### What does this PR do?

Type of change: new example

Adds `examples/speculative_decoding/scripts/quantize_drafter.py`, a CLI
that quantizes an exported speculative-decoding drafter to FP8 or NVFP4
— weight-only or weight+activation — with no calibration data.

It needs no modeling code either. Exported drafters such as
[`nvidia/MiniMax-M3-DSpark`](https://huggingface.co/nvidia/MiniMax-M3-DSpark)
have no importable model class, so each 2-D weight is wrapped in a
throwaway `nn.Linear` under its checkpoint key and ModelOpt's usual
`quantizer_name` patterns select over those names. Works for any drafter
layout (DSpark / DFlash / EAGLE3 / Medusa).

**Formats:** `w4a16_nvfp4`, `nvfp4`, `fp8`, `fp8_pc_pt` — the ModelOpt
formats vLLM's backend can actually serve. AWQ is deliberately not
offered, since `awq_lite` silently degrades to plain RTN without a
`forward_loop`.

**Static activation scales without calibration.** `fp8` and `nvfp4`
normally need an activation amax *measured* on calibration data; a fixed
`input_scale` of 1.0 is applied instead. That works because acceptance
length is governed almost entirely by **clipping**, not resolution:

Sweeping the fixed scale over three decades (same setup as the Testing
section below; bf16 baseline 3.1423):

| `input_scale` | amax | FP8 AL | vs bf16 | NVFP4 AL | vs bf16 |
|---|---|---|---|---|---|
| 0.003 | 1.3 | 2.2204 | -29.34% | 2.2076 | -29.75% |
| 0.01 | 4.5 | 2.6719 | -14.97% | 2.6641 | -15.22% |
| 0.03 | 13.4 | 2.9751 | -5.32% | 2.9259 | -6.89% |
| 0.1 | 44.8 | 3.1013 | -1.31% | 3.0206 | -3.88% |
| 0.2 | 89.6 | 3.1178 | -0.78% | 3.0015 | -4.48% |
| 0.3 | 134.4 | 3.1370 | -0.17% | 3.0222 | -3.82% |
| 0.5 | 224.0 | 3.1268 | -0.50% | 3.0360 | -3.38% |
| **1.0 (default)** | **448.0** | **3.1457** | **+0.11%** | **3.0193** |
**-3.91%** |
| 2.0 | 896.0 | 3.1354 | -0.22% | 3.0172 | -3.98% |
| 4.0 | 1792.0 | 3.1245 | -0.57% | 3.0034 | -4.42% |

Both formats fall off a cliff below ~0.03, where the declared range sits
far under the activations' true magnitude and most of the tensor is
clipped. Both then sit on a flat plateau from ~0.3 to 4.0 **with no
drop-off at the top**, so the scale only has to be big enough. 1.0 is
the middle of that plateau, which is why it is hardcoded rather than
exposed. NVFP4 trails FP8 by a roughly constant 3.5% across the plateau
— that gap is the 4-bit resolution cost, and no choice of scale recovers
it.

Deriving the amax from the weights instead was tried and does not work:
`max|W|` averages 0.79 while a RMSNorm'd activation is O(1) with outlier
channels in the tens, so the range lands 1–2 orders of magnitude low and
clips, measuring -31% to -46% AL.

**Where calibration would go.** All of this sits behind
`resolve_activation_scales()`, the single place deciding where a static
amax comes from. Real calibration slots in ahead of the fixed fallback
with no change to the CLI or the call site, and composes because
`set_static_activation_amax()` skips quantizers that already have an
amax:

```python
if calib_forward_loop is not None:
    mtq.calibrate(root, quant_cfg["algorithm"], forward_loop=calib_forward_loop)
set_static_activation_amax(root)   # fills in what calibration did not reach
```

**Serving a quantized drafter.** Four things had to be written into the
exported checkpoint before vLLM would load one:

- emit `quant_method` (`modelopt_fp4` / `modelopt`) — vLLM reads that
key, ModelOpt writes only `quant_algo`
- emit the exclusion list under `ignore` too — that is the key read from
the flat `quantization_config`; `exclude_modules` alone yields an empty
exclusion set
- add `*<name>` wildcards so exclusions match a runtime's nested module
prefix (`model.fc`) rather than the checkpoint key (`fc`)
- add `*qkv_proj` / `*gate_up_proj` aliases for layers a runtime fuses,
whose names appear in no checkpoint key

Nothing is then needed on the caller side. **This closes the open
question left in the previous revision of this PR: vLLM does read
`quantization_config` off the draft checkpoint.**
`ModelConfig._verify_quantization` fills `quantization` in from
`quant_method` when it is unset, so once the export declares that key —
the first fix above — detection works on its own. Verified on
Nemotron-3.5-Lightning passing nothing: `Detected ModelOpt NVFP4
checkpoint (quant_algo=NVFP4)` → `FlashInferCuteDslNvFp4LinearKernel`,
AL 4.278 against 4.203 measured earlier.

`specdec_bench` also gains a `DSPARK` algorithm, which it did not have:
an exported `Qwen3DSparkModel` would otherwise have to go through
`DFLASH` and be built with vLLM `method="dflash"`. The branch sets
`method="dspark"` and leaves `draft_sample_method` on vLLM's own default
of `greedy`. A target whose fused-collective workspace (sized at
CUDA-graph capture) overflows at large speculative batches can disable
graphs with `--runtime_params '{"engine_args": {"enforce_eager":
true}}'`.

For DFlash-family drafters, `qwen3_dflash.py` builds its fused
context-KV projection by reading `qkv_proj.weight` raw and calling
`F.linear`, which cannot consume a packed weight. Keep those layers in
bf16 with `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`;
`o_proj` and the MLP — the bulk of the drafter — still quantize. That
exclusion is mandatory, not a tuning choice.

`fc` (the projection from the target's captured layers into the draft)
is the one real knob, and it is a genuine trade rather than a free win —
see the Testing section for both models' numbers. The examples quantize
it; add `'*fc*'` to the exclude list to keep it in bf16.

`embed_tokens`, `markov_head` and `confidence_head` are excluded by
default: they are 2-D so the flat view treats them as GEMMs, but they
are embeddings or a single-output projection. `lm_head` is excluded by
the preset itself — unlike on a base model it is 37% of this drafter's
parameters, so `--quantize_lm_head` is a real lever (~1.9 GiB), but
measure AL first. The flag re-enables both of `lm_head`'s quantizers;
re-enabling only the weight one would ship a W+A checkpoint whose
`lm_head` has no `input_scale` while the config still advertises it as
quantized.

### Usage

```bash
# weight+activation FP8, calibration-free, lossless on both models measured below
python scripts/quantize_drafter.py \
    --drafter_path deepseek-ai/dspark_qwen3_8b_block7 \
    --qformat fp8 \
    --export_path ./dspark-qwen3-8b-fp8 \
    --exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'

# smallest: weight-only NVFP4
python scripts/quantize_drafter.py \
    --drafter_path nvidia/MiniMax-M3-DSpark \
    --qformat w4a16_nvfp4 \
    --export_path ./MiniMax-M3-DSpark-W4A16
```

Or end to end on Slurm — quantize, then measure AL — via the launcher
examples added here, one per target:

```bash
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_dspark_ptq_nvfp4.yaml --yes
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_dspark_ptq_nvfp4.yaml --yes
```

Serving one, if you are not going through `specdec_bench`:

```python
speculative_config = {
    "method": "dspark",
    "model": "./dspark-qwen3-8b-fp8",   # quantization is read from its config.json
    "num_speculative_tokens": 7,
}
```

### Testing

Two targets with different architectures, so the conclusions are not one
model's quirk:

* **Qwen3-8B** (dense transformer) +
[`deepseek-ai/dspark_qwen3_8b_block7`](https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7),
`block_size` 7, TP1.
* **Nemotron-3.5-Lightning-30B-A3B** (hybrid Mamba-MoE) +
[`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-DSpark),
`block_size` 8, TP8, with the mamba engine settings the model card pins
(`mamba_backend=flashinfer`, `mamba_ssm_cache_dtype=float16`, stochastic
SSM-cache rounding).

Both: MT-Bench 80 questions, greedy, one vLLM instance per point.

| recipe | activations | Qwen3-8B AL | vs bf16 | Nemotron-3.5 AL | vs
bf16 |
|---|---|---|---|---|---|
| bf16 baseline | — | 3.1423 | — | 4.3296 | — |
| **`fp8`** | static, `input_scale` 1.0 | **3.1457** | **+0.11%** |
**4.3289** | **-0.02%** |
| `fp8_pc_pt` | dynamic per-token | 3.1228 | -0.62% | 4.3411 | +0.26% |
| `w4a16_nvfp4`, `fc` in bf16 | bf16 (weight-only) | 3.0392 | -3.28% |
4.2899 | -0.92% |
| `w4a16_nvfp4`, `fc` quantized | bf16 (weight-only) | 3.0186 | -3.94% |
4.2334 | -2.22% |
| **`nvfp4`** | static, `input_scale` 1.0 | **3.0193** | **-3.91%** |
**4.2030** | **-2.92%** |

**FP8 weight+activation at the fixed `input_scale` of 1.0 is lossless on
both.** +0.11% and -0.02% are both inside run-to-run noise — the
Nemotron baseline was measured twice under identical settings and the
two runs differ by 0.94% (4.3093 / 4.3499), which sets the resolution of
that column. On the same reading, `fp8` and `fp8_pc_pt` are
indistinguishable on Nemotron; the dynamic variant only pulls ahead on
Qwen3. NVFP4 costs 3-4% on Qwen3 and 2-3% on Nemotron, i.e. the 4-bit
weight resolution is the real price and it is model-dependent but
bounded.

Whether to quantize `fc` is a per-model call rather than a general
recommendation — it buys a few percent of size for an AL cost that
differs by ~2x between these two drafters:

| `fc` bf16 → quantized | Qwen3-8B | Nemotron-3.5 |
|---|---|---|
| checkpoint size | 3.293 → 3.181 GiB (-3.4%) | 1.316 → 1.258 GiB
(-4.4%) |
| AL | 3.0392 → 3.0186 (-0.68%) | 4.2899 → 4.2334 (-1.32%) |

`fc` itself is only 3.5% (Qwen3) / 4.5% (Nemotron) of drafter
parameters; `embed_tokens` is the bulk (26% / 36%) and is excluded by
default.

The Qwen3 `w4a16_nvfp4` rows were measured in a later session than the
rest of that column; the `fc`-in-bf16 run reproduced the original number
to four decimals (3.0392), so the column is internally comparable.

Also validated on `nvidia/MiniMax-M3-DSpark`: `w4a16_nvfp4` runs in 67 s
on CPU, 9.98 GiB (fp32) -> 3.51 GiB; all 43 quantized tensors round-trip
within 0.0952 relative error; the 29 untouched tensors are bit-identical
to `bf16(source)`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (example-only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — validated manually as
above. Can add a `tests/examples/speculative_decoding/` test over a
small synthetic drafter if wanted before merge.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (example-only)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

The measurements above are one drafter on one target with one benchmark;
the plateau's location and the ~3.5% NVFP4 gap should be re-measured
before assuming they carry to a different drafter.

Note when reading an exported checkpoint: `input_scale` is `amax/448`
for FP8 but `amax/(6*448)` for NVFP4, so the one fixed amax records as
1.0 in an FP8 checkpoint and 0.1667 in an NVFP4 one. Both mean the same
activation range.

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-08-26 22:04:09 +08:00
..

Speculative Decoding (SpecDec) Bench

Installation

This benchmark is meant to be a lightweight layer ontop of an existing vLLM/SGLang/TRTLLM installation. For example, no install is required if one is running in the following dockers: vllm/vllm-openai:v0.11.0 (vLLM), lmsysorg/sglang:v0.5.4.post2 (SGLang), or nvcr.io/nvidia/tensorrt-llm/release:1.2.0 (TRT-LLM).

Next

cd examples/specdec_bench

Purpose

Collect relevant metrics on acceptance rate, timing, and outputs for Speculative Decoding methods. Acceptance rate refers to the number of tokens generated on every iteration. For a standard Autoregressive LLM, this number is just 1.

Getting Started

A basic example run script is provided which benchmarks MTBench (a standard 160 prompts spanning 8 categories). MTBench is available here

Running MTBench on GPT OSS + Eagle3

Download nvidia/gpt-oss-120b-Eagle3 to a local directory /path/to/eagle.

python3 run.py \
    --model_dir openai/gpt-oss-120b \
    --tokenizer openai/gpt-oss-120b \
    --draft_model_dir /path/to/eagle \
    --mtbench question.jsonl \
    --tp_size 1 \
    --ep_size 1 \
    --draft_length 3 \
    --output_length 4096 \
    --num_requests 80 \
    --engine TRTLLM \
    --concurrency 1 \
    --postprocess gptoss

Running Random ids on GPT OSS + Eagle3

Download nvidia/gpt-oss-120b-Eagle3 to a local directory /path/to/eagle.

python3 run.py \
    --model_dir openai/gpt-oss-120b \
    --tokenizer openai/gpt-oss-120b \
    --draft_model_dir /path/to/eagle \
    --random_isl 1024 \
    --tp_size 1 \
    --ep_size 1 \
    --draft_length 3 \
    --output_length 4096 \
    --num_requests 40 \
    --engine TRTLLM \
    --concurrency 1

Running SPEED-Bench on Llama 3.3 70B + Eagle 3

  1. Install the requirements file using pip install -r requirements.txt

  2. Prepare the data using the provided script:

python3 prepare_data.py --dataset speed --config all

The data will be saved to data/ directory, each config type (qualitative, throughput_1k, ...) to each own directory.

License

GOVERNING TERMS: This dataset is governed by the NVIDIA Evaluation Dataset License Agreement.

ADDITIONAL INFORMATION: MIT for bigcode/humanevalpack, RUCAIBox/MMATH, RUCAIBox/BAMBOO and EQ-Bench. Apache 2.0 for Writing Bench and Spec-Bench. CC BY 4.0 for FBK-MT/MCIF. MIT and Apache 2.0 for tianyang/repobench_python_v1.1, JetBrains-Research/lca-project-level-code-completion and tianyang/repobench_java_v1.1.

NOTICE: For each dataset a user elects to use, the user is responsible for checking if the dataset license is fit for the intended purpose. The prepare_data.py script automatically fetches data from all the source datasets.

Additional details are in HuggingFace dataset repository.

Qualitative split

python3 run.py \
    --model_dir meta-llama/Llama-3.3-70B-Instruct \
    --tokenizer meta-llama/Llama-3.3-70B-Instruct \
    --draft_model_dir yuhuili/EAGLE3-LLaMA3.3-Instruct-70B \
    --dataset speed \
    --dataset_path data/speed/qualitative \
    --tp_size 8 \
    --ep_size 1 \
    --draft_length 3 \
    --output_length 4096 \
    --engine TRTLLM \
    --concurrency 32 \
    --show_progress

Throughput split

python3 run.py \
    --model_dir meta-llama/Llama-3.3-70B-Instruct \
    --tokenizer meta-llama/Llama-3.3-70B-Instruct \
    --draft_model_dir yuhuili/EAGLE3-LLaMA3.3-Instruct-70B \
    --dataset speed \
    --dataset_path data/speed/throughput_1k \
    --tp_size 8 \
    --ep_size 1 \
    --draft_length 3 \
    --output_length 4096 \
    --engine TRTLLM \
    --concurrency 32 \
    --show_progress

For longer context (>8192 tokens), please use the following configuration when using TRTLLM:

engine_args:
  max_seq_len: 131072   # Model max context length (for Llama 3.3 70B)
  enable_chunked_prefill: true
python3 run.py \
    --model_dir meta-llama/Llama-3.3-70B-Instruct \
    --tokenizer meta-llama/Llama-3.3-70B-Instruct \
    --draft_model_dir yuhuili/EAGLE3-LLaMA3.3-Instruct-70B \
    --dataset speed \
    --dataset_path data/speed/throughput_16k \
    --tp_size 8 \
    --ep_size 1 \
    --draft_length 3 \
    --output_length 4096 \
    --engine TRTLLM \
    --concurrency 32 \
    --show_progress \
    --runtime_params runtime_args_long_context.yaml

Uploading results to S3

Each run.py invocation writes a result directory containing configuration.json, timing.json, acceptance_rate.json, and (when applicable) mtbench.json / specbench.json. upload_to_s3.py is a single-file, drop-in tool that uploads one run — or an entire sweep — to any S3-compatible bucket:

python upload_to_s3.py /path/to/run_or_sweep_dir s3://your-bucket/some/prefix \
    --endpoint https://your-s3-endpoint \
    --key-id YOUR_KEY_ID \
    --secret YOUR_SECRET

--endpoint, --key-id, and --secret default to the S3_ENDPOINT, S3_KEY_ID, and S3_SECRET environment variables. Omit --endpoint (or set S3_ENDPOINT="") to use AWS S3's default endpoint. Use --dry-run to preview the upload plan, and --skip-existing to skip runs already present at the destination instead of failing.

The tool handles two directory layouts and mirrors them into S3:

  • Flat — LOCAL_DIR/run_name/{configuration,timing,...}.json
  • Sweep — LOCAL_DIR/sweep_name/run_name/{configuration,timing,...}.json

LOCAL_DIR's basename is preserved in the destination prefix, so re-uploads from the same source land in the same place.

Notes

The goal of this benchmark is to provide an easy way to configure, run, and compare speculative implementations across frameworks in an apples-to-apples method. This benchmark sends request in a single-threaded fashion, so running large concurrency (>256) may result in python async scheduling delays and skew metrics. If larger concurrency is needed, it is recommended to fully deploy the model using vllm serve, python -m sglang.launch_server, or trtllm-serve (for vLLM, SGlang, or TRTLLM respectively) and use a more robust benchmarking client like NVIDIA AI Perf.