# Serve fakequant models with vLLM This is a simple example to demonstrate calibrating and serving ModelOpt fakequant models in vLLM. Compared with realquant, fakequant is 2-5x slower, but doesn't require dedicated kernel support and facilitates research. The general fakequant example is tested with vLLM 0.9.0, 0.19.1, 0.26.0, 0.28.0, 0.29.0 and 0.30.0. The compact NVFP4 attention worker documented below requires vLLM 0.15.0 or newer. ## Prepare environment Use the Dockerfile to build an environment with vLLM 0.30.0: ```bash docker build -f examples/vllm_serve/Dockerfile -t vllm-modelopt:v0.30.0 . ``` To build the same environment with another tested vLLM release, override `VLLM_VERSION`: ```bash docker build --build-arg VLLM_VERSION=0.26.0 \ -f examples/vllm_serve/Dockerfile -t vllm-modelopt:v0.26.0 . ``` For a direct installation from the ModelOpt repository root, install the tested vLLM release and the ModelOpt extras used by this example: ```bash python3 -m pip install "vllm==0.30.0" python3 -m pip install -e ".[all,mlflow]" ``` See the [ModelOpt installation guide](../../docs/source/getting_started/_installation_for_Linux.rst) for details about installing partial dependency sets. ## Calibrate and serve fake quant model in vLLM Step 1: Configure quantization settings. You can either edit the `quant_config` dictionary in `vllm_serve_fakequant.py`, or set the following environment variables to control quantization behavior: | Variable | Description | Default | |-----------------|--------------------------------------------------|---------------------| | QUANT_DATASET | Dataset name for calibration | cnn_dailymail | | QUANT_CALIB_SIZE| Number of samples used for calibration | 512 | | QUANT_CFG | Quantization config | None | | KV_QUANT_CFG | KV-cache quantization config | None | | QUANT_FILE_PATH | Optional path to exported quantizer state dict `quantizer_state.pth` | None | | MODELOPT_STATE_PATH | Optional path to exported `vllm_fq_modelopt_state.pth` (restores quantizer state and parameters) | None | | CALIB_BATCH_SIZE | Calibration batch size | 1 | | RECIPE_PATH | Optional path to a ModelOpt PTQ recipe YAML | None | Set these variables in your shell or Docker environment as needed to customize calibration. Step 2: Run the following command, with all supported flag as `vllm serve`: ```bash python vllm_serve_fakequant.py -tp 8 --host 0.0.0.0 --port 8000 ``` Hybrid attention/Mamba models such as Nemotron 3 Nano are supported on vLLM 0.26.0, 0.28.0, 0.29.0 and 0.30.0. For example, calibrate and serve with NVFP4 KV-cache fakequant as follows: ```bash KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \ python vllm_serve_fakequant.py -tp 8 \ --max-model-len 8192 --enforce-eager --host 0.0.0.0 --port 8000 ``` Calibration uses dedicated scratch KV-cache blocks, so reducing `--max-num-batched-tokens` is not required to avoid NaNs. For vLLM versions that expose `--moe-backend`, this launcher defaults to `--moe-backend triton`. ModelOpt expert fakequant needs a decomposed MoE backend so both expert GEMMs are visible during calibration. Step 3: test the API server with curl: ```bash curl -X POST "http://127.0.0.1:8000/v1/chat/completions" -H "Content-Type: application/json" -d '{ "model": "", "messages": [ {"role": "user", "content": "Hi, what is your name"} ], "max_tokens": 8 }' ``` Step 4 (Optional): using lm_eval to run evaluation ```bash lm_eval --model local-completions --tasks gsm8k --model_args model=,base_url=http://127.0.0.1:8000/v1/completions,num_concurrent=1,max_retries=3,tokenized_requests=False,batch_size=128,tokenizer_backend=None ``` ## Tracking a serve with MLflow Pass `--mlflow `, or set MLflow's own `MLFLOW_TRACKING_URI`, to record what this server actually quantized, so the numbers an evaluation produces can be traced back to a recipe: ```bash RECIPE_PATH= python vllm_serve_fakequant.py -tp 8 \ --host 0.0.0.0 --port 8000 \ --mlflow https:/// ``` This is the *quantization* tracking server. It is unrelated to any tracking server an evaluation harness exports its scores to — NeMo Evaluator Launcher, for instance, has its own `export.mlflow.tracking_uri`. Keep the two separate. Quantization runs in the vLLM **worker**, not in `vllm_serve_fakequant.py`, so that is where the run is recorded: the launcher validates the URI and hands the settings to the workers through the environment, and global rank 0 opens the run. It opens *before the weights load*, so a bad URI or a missing token fails within seconds rather than after a load and a full calibration, and it closes `FINISHED` once the model is quantized and warmed up — serving itself is not tracked.
Uploaded artifacts | Artifact | Contents | | --- | --- | | `command.txt` | The launcher's full invocation, copy-pasteable, with credentials masked | | `version.txt` | The ModelOpt version that ran | | `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s expanded, so it stands alone | | `recipe/quant_cfg.yaml` | `QUANT_CFG` and `KV_QUANT_CFG` merged, plus any MLA fixup — only when no recipe is used, since a recipe's config is already in `resolved_recipe.yaml` | | `logs/