mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: Bug fix
Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.
Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.
### Usage
```python
with disable_compilation(model):
mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```
No caller changes are required.
### Testing
- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
45 lines
1.2 KiB
Docker
45 lines
1.2 KiB
Docker
ARG VLLM_VERSION=0.30.0
|
|
FROM vllm/vllm-openai:v${VLLM_VERSION}
|
|
|
|
# Set environment variables
|
|
ENV PIP_NO_CACHE_DIR=off \
|
|
PIP_CONSTRAINT=
|
|
|
|
WORKDIR /workspace
|
|
|
|
# Install system dependencies needed for modelopt
|
|
RUN apt-get update && apt-get install -y \
|
|
git \
|
|
build-essential \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
# Copy the entire Model-Optimizer source code
|
|
COPY . Model-Optimizer
|
|
|
|
# Remove .git directory to reduce image size
|
|
RUN rm -rf Model-Optimizer/.git
|
|
|
|
# Install modelopt from local source with all dependencies. `mlflow` is the optional
|
|
# tracking client used by --mlflow; it is not part of `all`.
|
|
RUN cd Model-Optimizer && \
|
|
pip install -e ".[all,dev-test,mlflow]"
|
|
|
|
# Llama4 requires this
|
|
RUN pip install flash-attn==2.7.4.post1 --no-build-isolation
|
|
|
|
# Pre-compile CUDA extensions into a world-accessible directory so the vllm
|
|
# user can use the cache at runtime.
|
|
ENV TORCH_EXTENSIONS_DIR=/workspace/torch_extensions
|
|
RUN python3 -c "import modelopt.torch.quantization.extensions as ext; ext.precompile()" || true
|
|
|
|
# Allow the non-root vllm user to access the workspace
|
|
RUN chmod -R 777 /workspace
|
|
|
|
USER vllm
|
|
|
|
# Override the ENTRYPOINT from the base image to allow flexible usage
|
|
ENTRYPOINT []
|
|
|
|
# Set the default command
|
|
CMD ["/bin/bash"]
|