The owner-computes invariant
Thread-Owned C Tiles
Parallel column-row multiplication where each thread owns disjoint C tiles and never writes another thread's result cells.
Parallel wave
0 / 0
Ready
A matrix
M x KB matrix
K x NC matrix ownership
M x NThis wave
Race-free threaded loop order
parallel_for each thread t:
for C_tile in tiles_owned_by(t):
for k in 0..K-1:
# thread t is the only writer for this C tile
C_tile += A[tile_rows, k] outer B[k, tile_cols]