Increasing Arithmetic Intensity: Threads load into Registers to reuse data and slash Shared Memory access.
This visualizer uses a scaled-down teaching configuration (BM=8, BN=4, BK=4, TM=4) so the dataflow fits on screen. The CUDA kernel in this repo uses the same algorithmic structure with larger production constants.
Block tiling stays fixed at BM=8, BN=4, BK=4, TM=4.
BM = number of output rows each block computesBN = number of output columns each block computesBK = how much of the K dimension is loaded at a timeTM = number of output rows each thread computesFollow one thread's 4-register output strip across all phases.
As or Bs cell the selected thread is using right now.
(dotIdx, resultIndex) step, including other threads.
Btmp.
C: the 4-output column owned by the selected thread.
C.
| TID | Action | Btmp | threadResults[0] | threadResults[1] | threadResults[2] | threadResults[3] |
|---|
Is this 1D Tiling? Yes! Each thread computes a 1D column vector (in our viz, a \(4 \times 1\) tile) of the output block.
The Core Idea: We load a single value from Matrix B in shared memory into a fast, private hardware Register (Btmp).
Increasing Arithmetic Intensity: Because the thread calculates multiple rows (TM), it can reuse Btmp to multiply against multiple elements from Matrix A.
The Impact on the "Memory Wall":
Standard Shared Memory tiling fixes the Global Memory bottleneck, but replaces it with a Shared Memory bottleneck (Shared Memory is fast, but Registers are vastly faster).
Btmp once. Then we read 4 values from A to perform 4 FMA instructions. That's 5 Shared Reads for 4 FMAs. (Ratio 1.25:1)Why is the Shared Tile for A (8×4) larger than B (4×4)?
It directly relates to the work each thread is doing! Our block has 8 threads arranged in a 2×4 grid. Since each thread computes a 4×1 vertical strip (TM = 4), the block collectively computes an 8×4 block of the output Matrix C.
To compute an 8×4 block of C, the threads need 8 rows from Matrix A but only 4 columns from Matrix B. Thus, the tile fetched from A (As) must be twice as tall as the tile fetched from B (Bs).