Commit Graph
7 Commits
Author SHA1 Message Date
realAsma e27f76fbfe Refine agent contribution guidance (#1488)
### What does this PR do?

Type of change: documentation.

This PR centralizes repository agent guidance around `AGENTS.md` as the
shared
entrypoint. `CLAUDE.md` now points to `AGENTS.md`, while detailed coding
principles and tool-specific setup notes live under `.agents/`.

Key changes:

- Add `AGENTS.md` as the shared repository agent instructions file.
- Point `CLAUDE.md` at `AGENTS.md` so Claude Code reads the same
entrypoint.
- Add `.agents/developer-guidelines.md` for production code and review
principles, including minimal changes, extension points, testing,
performance,
  and compatibility expectations.
- Add `.agents/TOOLING.md` for human-maintained notes about local agent
  overrides and shared instruction maintenance.
- Update the Claude review workflow to read `AGENTS.md` and
  `.agents/developer-guidelines.md`.
- Update contributor and README guidance with focused local validation
examples
  and an AI-agent pointer.
- Ignore local agent override files.

### Usage

N/A. Documentation-only change.

### Testing

- `git diff --check`
- Commit pre-commit hooks passed, including `markdownlint-cli2`.
- GitHub PR checks passed, including code quality, docs, unit Linux,
required
  gate checks, DCO, CodeRabbit, and Codecov.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Docs-only agent guidance update.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive AI-agent instructions, developer guidelines, and
tooling notes for AI-assisted workflows.
* Updated contribution docs with clearer test-running guidance and
linked agent resources from the main README.
* Adjusted the code-review workflow to reference the new guidance
materials.
* **Chores**
  * Updated ignore rules to exclude local agent override files.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-05-14 11:55:15 -07:00
Keval MorabiaandClaude Opus 4.7 ed1d6a9061 [CI] Update PR Template with claude review details (#1404)
### What does this PR do?

Type of change: documentation

Adds a `Did you get Claude approval on this PR?` bullet to the
pre-review checklist in `.github/PULL_REQUEST_TEMPLATE.md`. Also adds
two CRITICAL rules to `CLAUDE.md` so Claude Code (and contributors using
it) follow the project's PR conventions:

- Use `.github/PULL_REQUEST_TEMPLATE.md` verbatim instead of the
harness's default `## Summary` / `## Test plan` format
- Run `/claude review` on non-trivial PRs before merging

Builds on #1400 (workflows added) and #1401 (proxy compat fixes).

### Usage

N/A — documentation only.

### Testing

- Verified the new checklist bullet renders under "Before your PR is
*Ready for review*" in this PR's description
- Plan to run `/claude review` on this PR as a smoke-test of the
integration on a docs-only diff

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ (will run after PR is open)

### Additional Information

Related: #1400, #1401

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added a new checklist item asking for Claude approval on pull
requests.
* Added guidance requiring the PR template be used verbatim when
creating PRs via the CLI.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 01:15:51 +05:30
jingyu-ml c7966119eb Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do?

Type of change: re-org code <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ We changed the import path
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅
<!--- Only for new features, API changes, critical bug fixes or backward
incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Calibration support for skip-softmax multi-threshold measurement in
sparse attention.
  * N:M sparse softmax masking and helpers for sparsity-aware attention.

* **Chores**
* Reorganized and consolidated kernel/backends for quantization and
sparsity to a unified kernels layout, updating tests and examples to
match.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-22 23:34:56 +00:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00
ZhiyuandClaude Opus 4.6 3249d0b041 [1/N] Polish PTQ skills (#1198)
## What does this PR do?
Polish PTQ skills based on learnings from
Devstral-Small-2-24B-Instruct-2512 NVFP4 PTQ, plus reviewer feedback.
                  
### Changes to `.claude/skills/ptq/SKILL.md`
                  
- Compacted FP8 rule to core principle (prefer `_QuantFP8Linear` lazy
dequant), defers details to `unsupported-models.md`
- Updated VLM rule: `hf_ptq.py` auto-handles VLMs via
`extract_and_prepare_language_model_from_vl()`
- Common Pitfalls: transformers version upgrade order, gated dataset
HF_TOKEN, NFS root_squash (references common docs)
  ### Changes to `.claude/skills/ptq/references/unsupported-models.md`
- Step A: Guidance on transformers install order (ModelOpt first, then
upgrade with deps)
- Step B: Rewritten — `_QuantFP8Linear` plugin handles standard FP8
automatically; manual dequant only for non-standard params
- Pattern 4 (VLM): Note that `hf_ptq.py` already handles VLM extraction
automatically
- Pattern 5 (FP8): Rewritten — plugin is primary path; added Triton
dependency note
   
### Changes to `.claude/skills/ptq/references/slurm-setup-ptq.md`
                  
- Replaced verbose offline-nodes and NFS sections with compact
"PTQ-Specific Notes" section (gated datasets + cross-reference to common
docs)
   
### Changes to `.claude/skills/common/slurm-setup.md`
                  
- Docker (non-pyxis) variant: full template for clusters using `docker
run` instead of pyxis/enroot
- Removed redundant `--runtime=nvidia` from Docker template (`--gpus
all` suffices)
- New section 5: NFS root_squash and Docker permissions (moved from
PTQ-specific docs). Promotes `docker run --user` as preferred fix;
tightened fallback chmod from `a+rwX` to `g+rwX`
  ### Changes to `CLAUDE.md`
- Made function signatures vague (`apply_mode(model, mode, ...)`) to
avoid documentation drift

### Changes to `.claude/skills/ptq/tests.json`
   
- Added test 5: MiniMax-M2.5 FP8 pre-quantized model — covers
`_QuantFP8Linear` plugin path and manual MoE expert dequantization

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 22:49:52 -07:00
Shengliang Xu 174f3a228d [OMNIML-3779] Add recipe into CLAUDE.md (#1106)
### What does this PR do?

Add recipe into CLAUDE.md

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated internal codebase documentation to reflect the addition of a
new recipe namespace for optimization specifications. Expanded
architectural overview to include recipe loading and validation
capabilities, along with reference materials for recipe-based model
optimization workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-03-24 17:45:45 -07:00
Rohan Joshi a007820afa Add CLAUDE.md (#956)
### What does this PR do?

Add CLAUDE.md file with repo overview for AI agents


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive CLAUDE.md documenting the Model Optimizer:
concepts, architecture, design patterns and anti-patterns, security and
contribution guidelines, common commands, architecture layout, core
abstractions (modes), key components overview, CI/testing and export
guidance, setup and workflow tips, and links to further documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
2026-03-06 19:30:42 +00:00