Memory movement, made visible

Matrix Tiling, Step by Step: From Cache Lines to Race-Free GEMM

Step through C = A × B as cache-sized tiles are fetched, reused, accumulated, and written.

Progress 0 / 0 Ready

What this visualizer makes visible

Matrix multiplication looks like one compact equation, but its speed depends on the order in which data reaches the processor. This lab slows the computation down so you can watch that order—not just the final answer.

1 · Follow one multiply-add. Yellow cells are read from A and B; blue marks the C cell being updated. The loop walker shows exactly where that operation sits inside the tiled nest.

2 · Change the traversal. Compare row-column dot products with rank-1 outer products, then change tile size and row/column-major layouts to see which cache lines each choice touches.

3 · Add threads without write races. The companion view assigns every C tile to one worker. Reads may overlap, but writes stay disjoint—so the output needs no atomics or reduction buffer.

Explore race-free threading
cache line: 64B
fetched now active tile written now partial C completed C not used now

A matrix

M x K

B matrix

K x N

C matrix

M x N

Outer-product tile micro-step

T x T tile update

A sub-vector

B sub-vector

rank-1 product grid

C tile after this k

Memory layout heatmap

linearized arrays
read write/update same cache line not touched

A memory

B memory

C memory

This step

Next step

Tiled loop order

for ii in 0..M step T:
  for jj in 0..N step T:
    for kk in 0..K step T:
      for i in ii..ii+T-1:
        for j in jj..jj+T-1:
          for k in kk..kk+T-1:
            C[i,j] += A[i,k] * B[k,j]

From the model to the machine

The same ownership rule scales to ARM SME

A matching C++ implementation uses the same owner-computes schedule: each worker accumulates complete C tiles while sharing read-only A and B panels. On ARM SME, a row-major A column panel is packed into contiguous storage, B rows are already contiguous, and ZA accumulator tiles hold the partial output. Its portable fallback keeps the same race-free schedule on machines without SME.

This page explains execution order and memory access. It is not a claim that one tile size or traversal wins on every processor; real performance still depends on cache geometry, vector width, matrix shape, packing cost, and the implementation around the kernel.