mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
check_scaled_pack_inputs checked that the block scales were on CUDA before check_pack_inputs looked at the input and grid, so a call with several bad arguments named the scales first. The pybind wrappers it replaced checked the input, grid and scales devices in that order, then everything else. check_pack_inputs now takes the scales optionally and checks their device right after the grid's, which restores that order. Over 5 formats and 8 combinations of bad arguments, every first error now matches main's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>