mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
pull-request/2597
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e27f76fbfe |
Refine agent contribution guidance (#1488)
### What does this PR do? Type of change: documentation. This PR centralizes repository agent guidance around `AGENTS.md` as the shared entrypoint. `CLAUDE.md` now points to `AGENTS.md`, while detailed coding principles and tool-specific setup notes live under `.agents/`. Key changes: - Add `AGENTS.md` as the shared repository agent instructions file. - Point `CLAUDE.md` at `AGENTS.md` so Claude Code reads the same entrypoint. - Add `.agents/developer-guidelines.md` for production code and review principles, including minimal changes, extension points, testing, performance, and compatibility expectations. - Add `.agents/TOOLING.md` for human-maintained notes about local agent overrides and shared instruction maintenance. - Update the Claude review workflow to read `AGENTS.md` and `.agents/developer-guidelines.md`. - Update contributor and README guidance with focused local validation examples and an AI-agent pointer. - Ignore local agent override files. ### Usage N/A. Documentation-only change. ### Testing - `git diff --check` - Commit pre-commit hooks passed, including `markdownlint-cli2`. - GitHub PR checks passed, including code quality, docs, unit Linux, required gate checks, DCO, CodeRabbit, and Codecov. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information Docs-only agent guidance update. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive AI-agent instructions, developer guidelines, and tooling notes for AI-assisted workflows. * Updated contribution docs with clearer test-running guidance and linked agent resources from the main README. * Adjusted the code-review workflow to reference the new guidance materials. * **Chores** * Updated ignore rules to exclude local agent override files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
ed1d6a9061 |
[CI] Update PR Template with claude review details (#1404)
### What does this PR do? Type of change: documentation Adds a `Did you get Claude approval on this PR?` bullet to the pre-review checklist in `.github/PULL_REQUEST_TEMPLATE.md`. Also adds two CRITICAL rules to `CLAUDE.md` so Claude Code (and contributors using it) follow the project's PR conventions: - Use `.github/PULL_REQUEST_TEMPLATE.md` verbatim instead of the harness's default `## Summary` / `## Test plan` format - Run `/claude review` on non-trivial PRs before merging Builds on #1400 (workflows added) and #1401 (proxy compat fixes). ### Usage N/A — documentation only. ### Testing - Verified the new checklist bullet renders under "Before your PR is *Ready for review*" in this PR's description - Plan to run `/claude review` on this PR as a smoke-test of the integration on a docs-only diff ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ (will run after PR is open) ### Additional Information Related: #1400, #1401 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added a new checklist item asking for Claude approval on pull requests. * Added guidance requiring the PR template be used verbatim when creating PRs via the CLI. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
c7966119eb |
Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do? Type of change: re-org code <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ We changed the import path <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ❌ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Calibration support for skip-softmax multi-threshold measurement in sparse attention. * N:M sparse softmax masking and helpers for sparsity-aware attention. * **Chores** * Reorganized and consolidated kernel/backends for quantization and sparsity to a unified kernels layout, updating tests and examples to match. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
3d0f0db49e |
[CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do? Type of change: New feature / infrastructure improvement Follow-up to #1285 for correct CI test environment for megatron based tests Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs, and wheel build sessions. The primary motivation was that `tox-current-env` is incompatible with uv venvs in NGC containers (e.g. NeMo's `/opt/venv`) — it picks the system Python via `sys._base_executable` instead of the container's venv Python which has megatron packages pre-installed. Key changes: - **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit, partial-install, pre-commit, docs, and wheel sessions - **GPU sessions** use `venv_backend="none"` (run directly in container env) and `python -m pip/pytest` to avoid PATH mismatches - **uv** is set as the default venv backend (if available) for CPU sessions (faster installs) Also includes CI workflow simplifications: - **`_pr_gate.yml`** new reusable workflow centralizing file-change detection + linux-check wait logic (was duplicated across 3 workflow files) - **Collapsed pr/non-pr job pairs** into single jobs with conditional `runs-on` in `gpu_tests.yml`, `example_tests.yml`, `regression_tests.yml` - **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a single `multi-version` matrix job in `unit_tests.yml` - **PR path filtering** for unit test secondary jobs (multi-version, launcher, partial-install) — skipped if no relevant files changed - **Fixed schedule/workflow_dispatch skipping** — jobs with `needs: [pr-gate]` were incorrectly skipped when all pr-gate internal jobs were skipped; fixed by making the gate job always run - **multi-version, launcher, partial-install** now also run on `schedule` / `workflow_dispatch` ### Usage ```bash python -m pip install nox uv # install nox and uv (once) nox -l # list all sessions nox -s gpu_megatron # run a GPU session (inside container) nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes) COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage ``` ### Testing - Ran `nox -l` to verify all session names - Ran `gpu_megatron` session locally inside NeMo container — confirmed it uses `/opt/venv/bin/python` correctly - Manually triggered nightly-runs: - Unit: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657 - GPU: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763 - Examples: https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A — CI infrastructure only - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox` and `uv` to `dev-test`, both Apache-2.0) - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no user-facing changes ### Additional Information Supersedes the tox-current-env workaround in the parent branch. --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3249d0b041 |
[1/N] Polish PTQ skills (#1198)
## What does this PR do?
Polish PTQ skills based on learnings from
Devstral-Small-2-24B-Instruct-2512 NVFP4 PTQ, plus reviewer feedback.
### Changes to `.claude/skills/ptq/SKILL.md`
- Compacted FP8 rule to core principle (prefer `_QuantFP8Linear` lazy
dequant), defers details to `unsupported-models.md`
- Updated VLM rule: `hf_ptq.py` auto-handles VLMs via
`extract_and_prepare_language_model_from_vl()`
- Common Pitfalls: transformers version upgrade order, gated dataset
HF_TOKEN, NFS root_squash (references common docs)
### Changes to `.claude/skills/ptq/references/unsupported-models.md`
- Step A: Guidance on transformers install order (ModelOpt first, then
upgrade with deps)
- Step B: Rewritten — `_QuantFP8Linear` plugin handles standard FP8
automatically; manual dequant only for non-standard params
- Pattern 4 (VLM): Note that `hf_ptq.py` already handles VLM extraction
automatically
- Pattern 5 (FP8): Rewritten — plugin is primary path; added Triton
dependency note
### Changes to `.claude/skills/ptq/references/slurm-setup-ptq.md`
- Replaced verbose offline-nodes and NFS sections with compact
"PTQ-Specific Notes" section (gated datasets + cross-reference to common
docs)
### Changes to `.claude/skills/common/slurm-setup.md`
- Docker (non-pyxis) variant: full template for clusters using `docker
run` instead of pyxis/enroot
- Removed redundant `--runtime=nvidia` from Docker template (`--gpus
all` suffices)
- New section 5: NFS root_squash and Docker permissions (moved from
PTQ-specific docs). Promotes `docker run --user` as preferred fix;
tightened fallback chmod from `a+rwX` to `g+rwX`
### Changes to `CLAUDE.md`
- Made function signatures vague (`apply_mode(model, mode, ...)`) to
avoid documentation drift
### Changes to `.claude/skills/ptq/tests.json`
- Added test 5: MiniMax-M2.5 FP8 pre-quantized model — covers
`_QuantFP8Linear` plugin path and manual MoE expert dequantization
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
174f3a228d |
[OMNIML-3779] Add recipe into CLAUDE.md (#1106)
### What does this PR do? Add recipe into CLAUDE.md <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated internal codebase documentation to reflect the addition of a new recipe namespace for optimization specifications. Expanded architectural overview to include recipe loading and validation capabilities, along with reference materials for recipe-based model optimization workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
a007820afa |
Add CLAUDE.md (#956)
### What does this PR do? Add CLAUDE.md file with repo overview for AI agents <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a comprehensive CLAUDE.md documenting the Model Optimizer: concepts, architecture, design patterns and anti-patterns, security and contribution guidelines, common commands, architecture layout, core abstractions (modes), key components overview, CI/testing and export guidance, setup and workflow tips, and links to further documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com> |