mirror of
https://github.com/tile-ai/tilelang.git
synced 2026-10-02 06:34:36 +08:00
* Remove unnecessary swizzled layout annotations from various attention sink examples and kernel analysis script for improved clarity and performance. * Remove unnecessary swizzled layout annotations from various examples to enhance code clarity and performance. * lint fix * Uncomment main test execution for vectorized cast * Add swizzled layout annotation for B_shared in dequantize GEMM example This change introduces a layout annotation for the B_shared tensor in the example_dequant_gemm_bf16_mxfp4_hopper.py file, enhancing the memory layout optimization for better performance during matrix multiplication. * Remove unnecessary cache disabling in GQA sink example for improved clarity and performance. * lint fix * Refactor layout functions to support row-major linear layout for any dimension - Renamed `makeGemmLayoutLinear` to `makeLinearLayout` and updated its implementation to handle arbitrary dimensions. - Updated related function calls in `gemm_layouts.cc` and `layout.cc` to use the new layout function. - Enhanced layout inference in `CumSumOpNode` to enforce linear layout for shared buffers in strict mode. * Refactor `make_linear_layout` to accept a single argument and support arbitrary dimensions - Updated the function signature to take a `Buffer`, `BufferLoad`, or `BufferRegion` directly. - Simplified the implementation by removing argument checks and directly obtaining the shape from the buffer info. - Enhanced the documentation to clarify the function's purpose and usage. * skip callback test * Remove example_dequant_gemm_bf16_mxfp4_hopper_tma.py file, eliminating unused code related to dequantization GEMM example. * typo fix
FP8 Matmul Benchmark (8192×8192)
This document records the throughput achieved by benchmark_matmul.py when multiplying FP8 matrices sized M = N = 8192 across different K dimensions. Each measurement relies on the default autotuning search space bundled with the benchmark.
Environment
- Repository commit:
6b1faf71faf18c564f5f77e0f5c1671cd91dfbc3 - GPUs:
NVIDIA H800 SXMon driver560.35.05
How to Reproduce
cd benchmark/matmul_fp8
python - <<'PY'
from benchmark_matmul import matmul
M = 8192
N = 8192
for K in [256, 512, 1024, 2048, 4096, 8192, 16384]:
res = matmul(M, N, K, False)
tflops = 2 * M * N * K / res.latency * 1e-12
print(f"K={K:5d} latency={res.latency:.6f}s TFlops={tflops:.3f}")
PY
Results
| K | Latency (s) | Throughput (TFLOPs) |
|---|---|---|
| 256 | 0.060352 | 569 |
| 512 | 0.080096 | 858 |
| 1024 | 0.121696 | 1129 |
| 2048 | 0.204672 | 1343 |
| 4096 | 0.374816 | 1467 |
| 8192 | 0.729664 | 1507 |
| 16384 | 1.427264 | 1541 |