The owner-computes invariant

Thread-Owned C Tiles

Parallel column-row multiplication where each thread owns disjoint C tiles and never writes another thread's result cells.

Parallel wave 0 / 0 Ready
Return to tile visualizer
A/B fetched now C written now partial C completed C thread 0 owns thread 1 owns thread 2 owns thread 3 owns

A matrix

M x K

B matrix

K x N

C matrix ownership

M x N

Parallel outer-product wave

owner-computes

This wave

Race-free threaded loop order

parallel_for each thread t:
  for C_tile in tiles_owned_by(t):
    for k in 0..K-1:
      # thread t is the only writer for this C tile
      C_tile += A[tile_rows, k] outer B[k, tile_cols]