Files
Chenjie LuoandClaude Opus 5.5 5bf43143c9 Decode the IQ formats on CUDA
Every format had a CUDA encoder but decoded with PyTorch ops, which was
fine while decoding looked like a one-off. Fake quant decodes every weight
on every forward, though, so once packing was fast and cached, decoding
became the cost: in the TinyLlama hf_ptq example it was 32 to 46 seconds of
a 55 to 70 second run, 3 ms per weight per forward, for the 100 preview
tokens. On a 27B model it would be about 10 seconds per forward.

Each format's Format struct now has a decode() beside its store(), and
decode_blocks in common.cuh launches one thread per 8-value vector: it
reads that vector's index, sign or delta bits and local scale from the
block, and writes the eight values in the requested dtype. common.cuh also
holds the bit readers and the two value forms, sign-flipped for IQ2 and
delta-shifted for IQ1. Every float operation is explicitly rounded and
follows the PyTorch decoder's order, so the compiler cannot fuse a multiply
into an add, and the output is bit-identical: across random payloads,
including block scales that decode to inf or NaN, and real encodings, in
float32, bfloat16, float16 and float64. The CUDA path also reproduces
llama.cpp's own values on the conformance blocks. dequantize_<format> uses
it for CUDA payloads and keeps the PyTorch path otherwise.

Decoding a 5632x2048 weight drops from 3.3 to 5.2 ms to 0.06 to 0.14 ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-10-01 06:36:40 +00:00
..
2026-10-01 06:36:40 +00:00