Watch threads cooperatively load data to fast On-Chip SRAM before computing.
Use small sizes so the grids stay readable.
Pick a thread to follow across the full tile-local dot product.
| Thread (tx,ty) | Global Target | Action | Local partialSum |
|---|
Without Shared Memory: Each thread reads an entire row of A and column of B from slow global DRAM.
With Shared Memory: Threads cooperatively load a tile into fast SRAM, then reuse it \(\text{TILE_SIZE}\) times.
Current example: \(2 \times 4 \times 4 \times 8 / 2 = 128\) ideal DRAM reads.
Why exactly \(2 \times M \times N \times K\) for the naive approach?
Why Tiling Fixes This: Neighboring threads need the same rows of A or columns of B. The naive approach fetches the exact same data from slow DRAM repeatedly. Tiling brings a chunk of that data into fast shared memory once, then threads read from the shared memory, drastically cutting down that massive number!
threadIdx.x. At a fixed k, the warp reads tileA[ty][k] (one address for everyone: a broadcast) and tileB[k][tx] (32 consecutive words of one row: 32 different banks). Each thread does walk down a column of tileB over the k loop, but those reads are separate instructions, so they cannot conflict. Padding matters when one warp reads down a column at once, as in the matrix transpose example.