Files
Lei Wang 09385e7deb [Refactor] Support auto swizzling for tma store and phaseout related layout annotations (#1509)
* Remove unnecessary swizzled layout annotations from various attention sink examples and kernel analysis script for improved clarity and performance.

* Remove unnecessary swizzled layout annotations from various examples to enhance code clarity and performance.

* lint fix

* Uncomment main test execution for vectorized cast

* Add swizzled layout annotation for B_shared in dequantize GEMM example

This change introduces a layout annotation for the B_shared tensor in the example_dequant_gemm_bf16_mxfp4_hopper.py file, enhancing the memory layout optimization for better performance during matrix multiplication.

* Remove unnecessary cache disabling in GQA sink example for improved clarity and performance.

* lint fix

* Refactor layout functions to support row-major linear layout for any dimension

- Renamed `makeGemmLayoutLinear` to `makeLinearLayout` and updated its implementation to handle arbitrary dimensions.
- Updated related function calls in `gemm_layouts.cc` and `layout.cc` to use the new layout function.
- Enhanced layout inference in `CumSumOpNode` to enforce linear layout for shared buffers in strict mode.

* Refactor `make_linear_layout` to accept a single argument and support arbitrary dimensions

- Updated the function signature to take a `Buffer`, `BufferLoad`, or `BufferRegion` directly.
- Simplified the implementation by removing argument checks and directly obtaining the shape from the buffer info.
- Enhanced the documentation to clarify the function's purpose and usage.

* skip callback test

* Remove example_dequant_gemm_bf16_mxfp4_hopper_tma.py file, eliminating unused code related to dequantization GEMM example.

* typo fix
2025-12-24 14:22:35 +08:00
..

FP8 Matmul Benchmark (8192×8192)

This document records the throughput achieved by benchmark_matmul.py when multiplying FP8 matrices sized M = N = 8192 across different K dimensions. Each measurement relies on the default autotuning search space bundled with the benchmark.

Environment

  • Repository commit: 6b1faf71faf18c564f5f77e0f5c1671cd91dfbc3
  • GPUs: NVIDIA H800 SXM on driver 560.35.05

How to Reproduce

cd benchmark/matmul_fp8
python - <<'PY'
from benchmark_matmul import matmul

M = 8192
N = 8192
for K in [256, 512, 1024, 2048, 4096, 8192, 16384]:
    res = matmul(M, N, K, False)
    tflops = 2 * M * N * K / res.latency * 1e-12
    print(f"K={K:5d}  latency={res.latency:.6f}s  TFlops={tflops:.3f}")
PY

Results

K Latency (s) Throughput (TFLOPs)
256 0.060352 569
512 0.080096 858
1024 0.121696 1129
2048 0.204672 1343
4096 0.374816 1467
8192 0.729664 1507
16384 1.427264 1541