mirror of
https://github.com/tile-ai/tilelang.git
synced 2026-10-02 06:34:36 +08:00
* [Metal] Add cooperative tensor intrinsics Expose TileLang-owned cooperative tensor builtins so Metal MPP lowering does not depend on extra TVM fork APIs. * [Metal] Select cooperative tensor GEMM lowering Add a shape-aware MPP instruction choice for shared-output Metal GEMM while preserving simdgroup fallback for fragments and unsupported tiles. * [Metal] Emit MPP cooperative tensor shaders Generate Metal 4 MPP matmul2d code for cooperative tensor intrinsics and keep source-only codegen separate from runtime compilation. * [Metal] Lower GEMM through MPP tensor ops Split Metal GEMM lowering into simdgroup and cooperative tensor emitters so M5 tiles use MPP while fragment accumulators keep the existing path. * [Metal] Guard generic passes for cooperative tensors Keep generic allocation and storage rewrites away from opaque Metal cooperative tensor scopes to avoid invalid scope analysis. * [Metal] Test cooperative tensor GEMM coverage Add runtime and source-only coverage for non-square MPP GEMM so the new cooperative tensor path is reproducible in CI and on M5. * [Metal] Update TVM Metal 4 runtime guard Point the submodule at the macOS SDK guarded Metal 4 runtime update used by cooperative tensor shaders. * [Metal] Document cooperative tensor GEMM Add a reference page covering the two Metal GEMM paths, selection rules, current limitations, and planned follow-up work. * lint * Improve Metal cooperative tensor GEMM * Optimize Metal cooperative tensor GEMM * Update TVM Metal 4 support * Add Metal backend internals documentation * Harden Metal cooperative tensor lowering * test older impl with new framework * Clean up Metal transform exports * Use macos-latest for Metal CI * resolve comments and lint --------- Co-authored-by: SiriusNEO <chaofan@deepseek.com>