mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature Packages the existing ModelOpt agent skills as installable Codex and Claude plugins: - Adds a repo-scoped Codex marketplace and Claude-compatible marketplace. - Adds the canonical `plugins/modelopt/` plugin tree and manifests. - Moves the skill tree into the plugin and keeps `.agents/skills` as a compatibility symlink. - Adds a minimal `common` placeholder skill required by Codex validation. - Documents installation from this repository. ### Usage ```bash codex plugin marketplace add NVIDIA/Model-Optimizer ``` Then open `/plugins`, select the `modelopt` marketplace, and install `modelopt`. For Claude Code: ```bash claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git claude plugin install modelopt@modelopt ``` ### Testing - Codex plugin validator - `claude plugin validate . --strict` - `claude plugin validate plugins/modelopt --strict` 1. Install the marketplace plugin with Codex and Claude from an unrelated temporary workspace. 2. Exercise packaged evaluation helpers, a day-0 gate, and the shared remote helper from that workspace. 3. Run `uv run --frozen --extra dev python -m pytest -q plugins/modelopt/skills/day0-release/tests/test_gates.py plugins/modelopt/skills/benchmark-model-kernels/tests`. 4. Run pre-commit hooks for all changed files. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added a plugin-path validator; existing focused skill tests and installed-plugin smoke tests pass. - Did you update Changelog?: N/A — agent tooling and distribution only. - Did you get Claude approval on this PR?: N/A ### Additional Information Skills remain available through `.agents/skills`; bundled helpers are packaged under the plugin and resolved from `$SKILL_DIR` so installed workflows do not depend on the current workspace. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added installable ModelOpt plugins for Claude Code and Codex. * Added skills for PTQ, deployment, evaluation, monitoring, debugging, benchmarking, MLflow access, EAGLE3 workflows, and release management. * Added deployment helpers, evaluation recipes, checkpoint validation, and release-gating tools. * **Documentation** * Expanded setup, credential, SLURM, benchmarking, deployment, evaluation, troubleshooting, and workspace guidance. * Added installation instructions and updated agent-skill discovery guidance. * **Maintenance** * Updated skill references and compatibility links for reliable use across supported plugin environments. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
4.1 KiB
4.1 KiB
Model Card Research
Use WebSearch to find the model card (HuggingFace, build.nvidia.com). Read it carefully, the FULL text, the devil is in the details. Extract ALL relevant configurations:
- Sampling params (
temperature,top_p) - Context length (
deployment.extra_args: "--max-model-len <value>") - Output length (
max_new_tokens) — mandatory extraction. Scan the card for anymax_tokens/max_new_tokens/ "output length" recommendation. Cards often list two values (e.g., Qwen3.x:32768thinking-general +81920math/coding). Pick the highest value and apply at the top level (no per-task overrides). If the card is genuinely silent on output length, note that explicitly and fall back to the generic default (64K reasoning / 16K non-reasoning) — never write a config with "card not yet checked" + generic default. See SKILL.md Step 3 "max_new_tokens— pick a single top-level value" for the full rule. - TP/DP settings (to set them appropriately, AskUserQuestion on how many GPUs the model will be deployed)
- Reasoning config (if applicable):
- reasoning on/off: use either:
adapter_config.custom_system_prompt(like/think,/no_think) and noadapter_config.params_to_add(leaveparams_to_addunrelated to reasoning untouched)adapter_config.params_to_addfor payload modifier (like"chat_template_kwargs": {"enable_thinking": true/false}) and noadapter_config.custom_system_promptandadapter_config.use_system_prompt: false(leavecustom_system_promptanduse_system_promptunrelated to reasoning untouched).
- The
chat_template_kwargstoggle key drifts across model generations — read the card /chat_template.jinja, don't extrapolate, and set only the one key the model uses. Known:enable_thinking(Qwen3.5/3.6, GLM 5.1 — note GLM-4.x usedthinking+/nothink);thinking(Kimi K2.6 — renamed from K2.5'senable_thinking; DeepSeek V3.2/V4 — Python encoder, not Jinja, so an unused kwarg can error rather than be ignored). - reasoning effort/budget (if configurable, e.g. DeepSeek V4
reasoning_effort): default tomax(the highest effort the card documents), honoring any tied requirement (e.g. V4 Think Max needs--max-model-len >= 393216). AskUserQuestion only if the user signals a cost/latency preference. - etc.
- reasoning on/off: use either:
- Deployment-specific
extra_argsfor vLLM/SGLang (look for the vLLM/SGLang deployment command) - Deployment-specific vLLM/SGLang versions (by default we use latest docker images, but you can control it with
deployment.imagee.g. vLLM abovevllm/vllm-openai:v0.11.0stopped supportingrope-scalingarg used by Qwen models) - ARM64 / non-standard GPU compatibility: The default
vllm/vllm-openaiimage only supports common GPU architectures. For ARM64 platforms or GPUs with non-standard compute capabilities (e.g., NVIDIA GB10 with sm_121), use NGC vLLM images instead:- Example:
deployment.image: nvcr.io/nvidia/vllm:26.01-py3 - AskUserQuestion about their GPU architecture if the model card doesn't specify deployment constraints
- Example:
- Any preparation requirements (e.g., downloading reasoning parsers, custom plugins):
- If the model card mentions downloading files (like reasoning parsers, custom plugins) before deployment, add
deployment.pre_cmdwith the download command - Use
curlinstead ofwgetas it's more widely available in Docker containers - Example:
pre_cmd: curl -L -o reasoning_parser.py https://huggingface.co/.../reasoning_parser.py - When using
pip installinpre_cmd, always use--no-cache-dirto avoid cross-device link errors in Docker containers (the pip cache and temp directories may be on different filesystems) - Example:
pre_cmd: pip3 install --no-cache-dir flash-attn --no-build-isolation
- If the model card mentions downloading files (like reasoning parsers, custom plugins) before deployment, add
- Any other model-specific requirements
Remember to check evaluation.nemo_evaluator_config and evaluation.tasks.*.nemo_evaluator_config overrides too for parameters to adjust (e.g. disabling reasoning)!
Present findings, explain each setting, ask user to confirm or adjust. If no model card found, ask user directly for the above configurations.