mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix + test robustness Fixes the puzzletron GPU tests and makes them robust to GPU vendor / transformers version drift. The previous exact-equality checks failed in CI because bf16 attention paths produce different operation orderings on different GPUs (Ada vs Blackwell), which completely reshuffles close-in-magnitude FFN importance scores on the small (hidden=256) test models. Earlier set-membership relaxations weren't enough — the orderings between GPUs were near-orthogonal. **Changes:** *Numerical determinism (root cause):* - Switch test forward to **fp32 autocast** (`autocast_dtype: torch.float32` in `validate_model_defaults.yaml`). Combined with TF32 disabled (below), this gives bit-identical results across GPU vendors. - Reduce `block_size` from 8192 → **2048**: SDPA's flash attention only supports bf16/fp16, so at fp32 it materializes the full attention matrix; seq=8192 OOMs (32 GB), seq=2048 fits. - Strengthen `set_seed`: also disable TF32 on cuda matmul + cudnn so fp32 matmuls produce bit-identical results across GPU vendors. *Test structure (orthogonal improvements):* - Refactor `_assert_*` helpers to `_check_*` that return error lists; the multiprocess job collects all failures and fails once at the end (instead of fail-fast). Single CI run now surfaces every mismatch. - Drop the brittle exact-equality `score[0]` (channel-rank) check. Use **set-membership at both ends of the importance ranking**: - Expected least-important channel must land in top-K least-important of actual. - Expected most-important channel must land in top-K most-important of actual. - `K = max(8, num_channels // 16)` gives slack for any residual drift without defeating the check. *Tolerances:* - `_check_lm_loss` default tolerance: **0.001** for non-MoE (now essentially bit-deterministic). - MoE caller passes **0.01** to absorb expert-routing run-to-run noise (~0.003). - Nemotron-3-Nano-30B-A3B (hybrid mamba+MoE) remains excluded from `EXPECTED_LM_LOSS` — its ~0.06 run-to-run noise would require a tolerance that defeats the check. *Values:* - Update `EXPECTED_FFN_PRUNING_VALUES` (with `least_important` + `most_important` per layer) and `EXPECTED_LM_LOSS` for the new fp32 + shorter-seq baseline. ### Testing Verified locally on RTX 6000 Ada across transformers `4.57.0`, `5.6.0`, `5.9.0`, in both 1-GPU and 2-GPU modes: | | 4.57 | 5.6 | 5.9 | |-------|-----|-----|-----| | 1-GPU | 8/9 † | 9/9 | 9/9 | | 2-GPU | — | — | 9/9 | † Qwen3-VL-30B-A3B fails on 4.57: its lm_loss is 4.92 on 4.57 but 5.03125 on 5.6+. This is a pre-existing cross-major-version drift documented in a code comment; not introduced by this PR. CI runs transformers 5.9 (container 26.04). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (test-only change) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>