Memory movement, made visible
Matrix Tiling, Step by Step: From Cache Lines to Race-Free GEMM
Step through C = A × B as cache-sized tiles are fetched, reused, accumulated, and written.
What this visualizer makes visible
Matrix multiplication looks like one compact equation, but its speed depends on the order in which data reaches the processor. This lab slows the computation down so you can watch that order—not just the final answer.
1 · Follow one multiply-add. Yellow cells are read from A and B; blue marks the C cell being updated. The loop walker shows exactly where that operation sits inside the tiled nest.
2 · Change the traversal. Compare row-column dot products with rank-1 outer products, then change tile size and row/column-major layouts to see which cache lines each choice touches.
3 · Add threads without write races. The companion view assigns every C tile to one worker. Reads may overlap, but writes stay disjoint—so the output needs no atomics or reduction buffer.
A matrix
M x KB matrix
K x NC matrix
M x NMemory layout heatmap
linearized arraysA memory
B memory
C memory
This step
Next step
Tiled loop order
for ii in 0..M step T:
for jj in 0..N step T:
for kk in 0..K step T:
for i in ii..ii+T-1:
for j in jj..jj+T-1:
for k in kk..kk+T-1:
C[i,j] += A[i,k] * B[k,j]
From the model to the machine
The same ownership rule scales to ARM SME
A matching C++ implementation uses the same owner-computes schedule: each worker accumulates complete C tiles while sharing read-only A and B panels. On ARM SME, a row-major A column panel is packed into contiguous storage, B rows are already contiguous, and ZA accumulator tiles hold the partial output. Its portable fallback keeps the same race-free schedule on machines without SME.
This page explains execution order and memory access. It is not a claim that one tile size or traversal wins on every processor; real performance still depends on cache geometry, vector width, matrix shape, packing cost, and the implementation around the kernel.