mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: new feature
Adds a general Slurm-only QAD skill based on the supported Megatron
Bridge
workflow. The skill:
- starts from a measured BF16-to-PTQ benchmark gap and preserves the
preceding
PTQ configuration or recipe;
- gates QAD on exact Megatron Bridge model support and successful
Megatron PTQ,
using its master-rank quantizer summary as a scoped `amax` sanity check;
- requires model- and hardware-derived TP/PP/CP/EP/ETP topology
selection;
- streams and randomly samples only the required
`nvidia/Nemotron-Cascade-2-SFT-Data` token budget and uses Megatron
sequence
packing;
- defaults to 32K sequences, LR `1e-5` with cosine decay, a 1000-step
cap, and
GBS 512;
- requires explicit user authorization because QAD is costly, validates
two
batches every 25 steps, saves every 50 steps, and monitors a decreasing
smoothed loss trend;
- evaluates an early checkpoint around step 150 and continues only when
benchmark recovery and the loss trend justify more training;
- follows the established common Slurm and remote-execution guidance
instead of
duplicating mutable commands from the Megatron Bridge README.
Also exposes Megatron Bridge `save_interval`, `exit_interval`, and
`exit_duration_in_mins` through `examples/megatron_bridge/distill.py`,
with example-test coverage for checkpoint
and ModelOpt-state preservation at an early exit.
### Usage
```text
Use the QAD skill to recover the measured BF16-to-PTQ benchmark gap for
<model> on <Slurm cluster>, preserving the validated PTQ recipe.
```
### Testing
- `PYTHONPATH=$PWD pre-commit run --all-files`
- Passed every hook on the rebased branch, including Ruff, Ruff format,
mypy,
YAML/recipe validation, launcher reference validation, Bandit, generated
arguments, symlink synchronization, and Markdown lint.
- `python
~/.codex/skills/.system/skill-creator/scripts/quick_validate.py
.agents/skills/qad`
- `Skill is valid!`
Qwen3-0.6B result-bearing validation:
- Resources: one exclusive node, 8 H100 GPUs
- Container: `nvcr.io/nvidia/nemo:26.06`
- Quantization: NVFP4, group size 16, embedding excluded
- QAD topology: TP=1, PP=1, CP=4, EP=1, DP=2
- Training validation configuration: sequence length 32768, MBS=1,
GBS=8,
`train_iters=1000`, LR `1e-5` / minimum LR `1e-6`, 50 warmup iterations,
cosine decay, `eval_interval=150`, `exit_interval=150`,
`exit_duration_in_mins=220`
- This result-bearing run used the then-current coupled eval/save
cadence. The
final skill now validates two batches every 25 steps and saves every 50;
the
example test covers the independent checkpoint cadence.
- The reduced GBS 8 is intentionally validation-only; the skill retains
GBS 512
as the production default.
- Data: exactly 10,000,000 sampled tokens from four
`nvidia/Nemotron-Cascade-2-SFT-Data` configs:
- math: 2,306,011 tokens / 364 documents
- science: 1,191,257 tokens / 285 documents
- chat: 6,142,077 tokens / 1,800 documents
- instruction following: 360,655 tokens / 411 documents
- Megatron built packed 32K GPT samples from the materialized prefixes;
the full
dataset was not downloaded.
- QAD loss was finite and decreased from `0.2640341` at iteration 10 to
`0.1060580` at iteration 150. Final gradient norm was `0.747`, with zero
skipped and zero NaN iterations. Validation distillation loss was
`0.09715855`.
- The iteration-150 checkpoint saved successfully with `modelopt_state`,
and
both PTQ and QAD-150 exported to unified Hugging Face format.
- Identical full MMLU 0-shot comparison through the Megatron evaluator:
| Model | Accuracy |
| --- | ---: |
| BF16 | 0.39517164 |
| PTQ | 0.32851446 |
| QAD-150 | 0.38740921 |
QAD-150 recovered `0.05889475 / 0.06665718 = 88.35%` of the measured PTQ
gap,
so validation stopped at the early evidence gate rather than continuing
blindly toward 1000 iterations.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did
you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — this adds an agent skill and
example-only
lifecycle flags.
- Did you get Claude approval on this PR?: N/A
### Additional Information
All seven branch commits are cryptographically signed and include a
`Signed-off-by` trailer.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Updated the QAD skill documentation with a clear “Execute in this
order” workflow, including a revised default recovery training policy.
* Added a new `nemotron-cascade-2` dataset blend configuration with an
increased token budget.
* Enhanced the MeGatron Bridge distillation CLI with stricter interval
argument validation and support for configurable save-and-exit controls.
* **Documentation**
* Expanded Megatron Bridge README guidance for dataset preparation,
token-budget recalculation, and resume expectations.
* **Tests**
* Improved distillation and QAD tests to validate early-exit behavior
and checkpoint expectations.
* Added unit tests covering distillation CLI interval validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Meng Xin <mxin@nvidia.com>
46 lines
1.3 KiB
YAML
46 lines
1.3 KiB
YAML
# Set to the target model's Hugging Face ID or local tokenizer path.
|
|
tokenizer: <target-model-tokenizer>
|
|
output_dir: <session-model-workspace>/data/nemotron-cascade-2-through-1000
|
|
target_tokens: 17_300_000_000
|
|
sources:
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: math
|
|
split: train
|
|
content_field: messages
|
|
weight: 21.1
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: science
|
|
split: train
|
|
content_field: messages
|
|
weight: 10.9
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: chat
|
|
split: train
|
|
content_field: messages
|
|
weight: 56.2
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: instruction_following
|
|
split: train
|
|
content_field: messages
|
|
weight: 3.3
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: safety
|
|
split: train
|
|
content_field: messages
|
|
weight: 0.02
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: conversational_agent
|
|
split: train
|
|
content_field: messages
|
|
weight: 3.3
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: swe
|
|
split: train
|
|
content_field: messages
|
|
weight: 1.8
|
|
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
|
|
config: terminal_agent
|
|
split: train
|
|
content_field: messages
|
|
weight: 3.3
|