Files
Model-Optimizer/modelopt
Keval Morabia 7ae18651a5 Fix puzzletron tests and make FFN pruning check robust (#1556)
### What does this PR do?

Type of change: Bug fix + test robustness

Fixes the puzzletron GPU tests and makes them robust to GPU vendor /
transformers version drift. The previous exact-equality checks failed in
CI because bf16 attention paths produce different operation orderings on
different GPUs (Ada vs Blackwell), which completely reshuffles
close-in-magnitude FFN importance scores on the small (hidden=256) test
models. Earlier set-membership relaxations weren't enough — the
orderings between GPUs were near-orthogonal.

**Changes:**

*Numerical determinism (root cause):*
- Switch test forward to **fp32 autocast** (`autocast_dtype:
torch.float32` in `validate_model_defaults.yaml`). Combined with TF32
disabled (below), this gives bit-identical results across GPU vendors.
- Reduce `block_size` from 8192 → **2048**: SDPA's flash attention only
supports bf16/fp16, so at fp32 it materializes the full attention
matrix; seq=8192 OOMs (32 GB), seq=2048 fits.
- Strengthen `set_seed`: also disable TF32 on cuda matmul + cudnn so
fp32 matmuls produce bit-identical results across GPU vendors.

*Test structure (orthogonal improvements):*
- Refactor `_assert_*` helpers to `_check_*` that return error lists;
the multiprocess job collects all failures and fails once at the end
(instead of fail-fast). Single CI run now surfaces every mismatch.
- Drop the brittle exact-equality `score[0]` (channel-rank) check. Use
**set-membership at both ends of the importance ranking**:
- Expected least-important channel must land in top-K least-important of
actual.
- Expected most-important channel must land in top-K most-important of
actual.
- `K = max(8, num_channels // 16)` gives slack for any residual drift
without defeating the check.

*Tolerances:*
- `_check_lm_loss` default tolerance: **0.001** for non-MoE (now
essentially bit-deterministic).
- MoE caller passes **0.01** to absorb expert-routing run-to-run noise
(~0.003).
- Nemotron-3-Nano-30B-A3B (hybrid mamba+MoE) remains excluded from
`EXPECTED_LM_LOSS` — its ~0.06 run-to-run noise would require a
tolerance that defeats the check.

*Values:*
- Update `EXPECTED_FFN_PRUNING_VALUES` (with `least_important` +
`most_important` per layer) and `EXPECTED_LM_LOSS` for the new fp32 +
shorter-seq baseline.

### Testing

Verified locally on RTX 6000 Ada across transformers `4.57.0`, `5.6.0`,
`5.9.0`, in both 1-GPU and 2-GPU modes:

|       | 4.57 | 5.6 | 5.9 |
|-------|-----|-----|-----|
| 1-GPU | 8/9 † | 9/9 | 9/9 |
| 2-GPU | — | — | 9/9 |

† Qwen3-VL-30B-A3B fails on 4.57: its lm_loss is 4.92 on 4.57 but
5.03125 on 5.6+. This is a pre-existing cross-major-version drift
documented in a code comment; not introduced by this PR. CI runs
transformers 5.9 (container 26.04).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (test-only change)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-29 07:33:11 +00:00
..