A visual explanation of how large matrix products are split into cache-sized tiles, accumulated across K, and reorganized for efficient memory access.
← Home
Writing tagged GEMM
Published writing tagged GEMM.
How owner-computes scheduling splits a large result matrix into disjoint C tiles so threads can use outer-product updates without atomics or competing writes.