### What does this PR do?
Type of change: documentation.
Replaces the legacy Sphinx RTD theme with Shibuya and gives the ModelOpt
documentation a modern responsive light/dark presentation.
The configuration uses the Shibuya green palette with NVIDIA green
(#76b900) as the primary accent, follows the reader system color
preference by default, enables dark code blocks, and expands the first
level of global navigation. RTD-specific CSS is removed, while the
announcement page now uses Shibuya semantic color tokens in both light
and dark modes.
Shibuya 2026.7.12 is licensed under BSD-3-Clause.
### Usage
```python
html_theme = "shibuya"
html_theme_options = {
"accent_color": "green",
"color_mode": "auto",
"dark_code": True,
"globaltoc_expand_depth": 1,
}
```
Doc preview:
https://nvidia.github.io/Model-Optimizer/pr-preview/pr-2242/
### Testing
- `nox -N --envdir /tmp/modelopt-shibuya-nox -s docs`
- Sphinx 9.1 built all 411 pages successfully with `--fail-on-warning`.
- `pre-commit run --files docs/source/conf.py
docs/source/_static/custom.css docs/source/_static/announcements.css
pyproject.toml uv.lock --show-diff-on-failure`
- All applicable hooks passed.
- Verified generated HTML loads Shibuya assets, declares the green
accent, includes automatic light/dark mode logic, and includes the
custom theme-token CSS.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A — documentation theme
change covered by the full warning-as-error build.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation presentation change.
- Did you get Claude approval on this PR?: N/A — focused documentation
theme migration.
### Additional Information
No source code or public API behavior changes.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated the documentation site with the Shibuya theme.
* Added green theme accents, automatic light/dark mode, and dark code
blocks.
* Improved table-of-contents behavior and refreshed announcement
styling.
* Removed outdated layout and table customization overrides.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: realAsma <akuriparambi@nvidia.com>
Fix broken doc building CI and pin docs/test dependencies to avoid such
issues in future
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Improved several guide cross-references so links resolve more reliably
in the rendered docs.
* Updated references in the quantization guides for better navigation
between related topics.
* **Chores**
* Tightened versions for documentation and test tooling to improve
consistency in local and CI environments.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
sphinx 9.x requires Python >=3.12, but requires-python is >=3.10, so
`uv lock` failed resolving dev-docs for 3.10/3.11. Mark the dev-docs
group with `python_version >= '3.12'` so uv skips it for older Pythons
(we don't build docs there). Includes the resulting uv.lock upgrade.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
### What does this PR do?
Type of change: New feature / infrastructure improvement
Follow-up to #1285 for correct CI test environment for megatron based
tests
Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.
Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)
Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`
### Usage
```bash
python -m pip install nox uv # install nox and uv (once)
nox -l # list all sessions
nox -s gpu_megatron # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)" # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)" # with coverage
```
### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes
### Additional Information
Supersedes the tox-current-env workaround in the parent branch.
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
## Summary
- **hf_online_dflash.yaml**: Add 100K-sample training config with
regression baselines (B200 loss curve),
`MAX_FINAL_LOSS`/`MIN_FINAL_ACC`/`MIN_ACCEPTANCE_LENGTH` thresholds,
vLLM nightly container for DFlash support
- **vllm_smoke_test.sh**: Parse acceptance length from vLLM server log
for regression check; `pip install pandas` workaround for broken nightly
container; capture server output to temp file
- **query.sh**: Detect vLLM server death during startup (PID liveness
check) + 600s timeout to prevent infinite polling that wastes GPU hours;
`pip install pandas` workaround
- Fix empty `environment:` key in DFlash YAML causing nemo_run
`ListParseError`
## Test plan
- [x] E2E pipeline passed on 8x B200 (training + vLLM smoke test + AR
eval)
- [x] Training regression: final loss 3.82 < 5.0, acc 0.20 > 0.15
- [x] vLLM acceptance length: 1.79 >= 1.4 threshold
- [x] AR evaluation: 2.02 overall on MT-Bench (8 categories)
- [x] Server liveness check prevents GPU waste on vLLM crash
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added optional regression validation for vLLM acceptance metrics
* Introduced configurable vLLM server startup timeout (default 600
seconds)
* **Improvements**
* Enhanced logging for vLLM server startup with progress tracking and
waited time reporting
* Faster detection of vLLM server process failures during initialization
* **Configuration Updates**
* Increased training dataset size and logging granularity
* Scaled tensor parallelism from 4 to 8 across multiple pipelines
* Expanded PTQ quantization to multi-step pipeline
* Added configurable training metric thresholds
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
## What does this PR do?
**Type of change:** Project config update <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
Remove unnecessary setup.py as the same can be done with pyproject.toml.
No change for users in terms of installation command.
## Testing
<!-- Mention how have you tested your change if applicable. -->
Per PR tests passing is sufficient validation
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Migrated project configuration from setup.py to pyproject.toml,
adopting modern Python packaging standards.
* Updated CI/CD workflows and development tools to reference the new
configuration location.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>