Files
Yichen YanandSiriusNEO 1545f0065d [Metal] M5 Cooperative Tensor T.gemm (#2252)
* [Metal] Add cooperative tensor intrinsics

Expose TileLang-owned cooperative tensor builtins so Metal MPP lowering does not depend on extra TVM fork APIs.

* [Metal] Select cooperative tensor GEMM lowering

Add a shape-aware MPP instruction choice for shared-output Metal GEMM while preserving simdgroup fallback for fragments and unsupported tiles.

* [Metal] Emit MPP cooperative tensor shaders

Generate Metal 4 MPP matmul2d code for cooperative tensor intrinsics and keep source-only codegen separate from runtime compilation.

* [Metal] Lower GEMM through MPP tensor ops

Split Metal GEMM lowering into simdgroup and cooperative tensor emitters so M5 tiles use MPP while fragment accumulators keep the existing path.

* [Metal] Guard generic passes for cooperative tensors

Keep generic allocation and storage rewrites away from opaque Metal cooperative tensor scopes to avoid invalid scope analysis.

* [Metal] Test cooperative tensor GEMM coverage

Add runtime and source-only coverage for non-square MPP GEMM so the new cooperative tensor path is reproducible in CI and on M5.

* [Metal] Update TVM Metal 4 runtime guard

Point the submodule at the macOS SDK guarded Metal 4 runtime update used by cooperative tensor shaders.

* [Metal] Document cooperative tensor GEMM

Add a reference page covering the two Metal GEMM paths, selection rules, current limitations, and planned follow-up work.

* lint

* Improve Metal cooperative tensor GEMM

* Optimize Metal cooperative tensor GEMM

* Update TVM Metal 4 support

* Add Metal backend internals documentation

* Harden Metal cooperative tensor lowering

* test older impl with new framework

* Clean up Metal transform exports

* Use macos-latest for Metal CI

* resolve comments and lint

---------

Co-authored-by: SiriusNEO <chaofan@deepseek.com>
2026-07-28 16:17:15 +08:00
..