mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Every format had a CUDA encoder but decoded with PyTorch ops, which was fine while decoding looked like a one-off. Fake quant decodes every weight on every forward, though, so once packing was fast and cached, decoding became the cost: in the TinyLlama hf_ptq example it was 32 to 46 seconds of a 55 to 70 second run, 3 ms per weight per forward, for the 100 preview tokens. On a 27B model it would be about 10 seconds per forward. Each format's Format struct now has a decode() beside its store(), and decode_blocks in common.cuh launches one thread per 8-value vector: it reads that vector's index, sign or delta bits and local scale from the block, and writes the eight values in the requested dtype. common.cuh also holds the bit readers and the two value forms, sign-flipped for IQ2 and delta-shifted for IQ1. Every float operation is explicitly rounded and follows the PyTorch decoder's order, so the compiler cannot fuse a multiply into an add, and the output is bit-identical: across random payloads, including block scales that decode to inf or NaN, and real encodings, in float32, bfloat16, float16 and float64. The CUDA path also reproduces llama.cpp's own values on the conformance blocks. dequantize_<format> uses it for CUDA payloads and keeps the PyTorch path otherwise. Decoding a 5632x2048 weight drops from 3.3 to 5.2 ms to 0.06 to 0.14 ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>