The Outer Product approach: Heavy math intensity with dual register arrays (regM and regN).
| TID | Action | regM[0] | regM[1] | regN[0] | regN[1] | Res[0,0] | Res[0,1] | Res[1,0] | Res[1,1] |
|---|
Visualizing the Outer Product Mapping:
During the FMA Math step, watch the Shared Memory grids (As and Bs) carefully. The lightly colored column and row indicate the vectors loaded into private registers. The brightly highlighted pink and indigo cells show exactly which single scalar value is currently being used by the thread to compute the intersection.
Notice how as the math progresses, these highlighted operands "slide" across the loaded vectors to systematically calculate every intersection in the 2x2 output grid!
The Shift to Outer Products: Instead of loading 1 element to reuse along a 1D strip, each thread loads a column vector (regM) from A and a row vector (regN) from B.
The Math: When you multiply a column vector by a row vector, you get a full 2D matrix! The thread performs \(\text{TM} \times \text{TN}\) math operations using only those few registers.
Shattering the Memory Wall: Arithmetic Intensity skyrockets. The ratio of math operations to Shared Memory reads goes through the roof.
The Problem with 1D Tiling (4×1 strip)
The Magic of 2D Tiling (4×4 square)
Why this matters on real GPUs:
In production kernels, you usually see 8×8 or 16×16 2D register tiling. If a thread uses an 8×8 tile, it reads 8 values from A and 8 values from B (16 reads total), and uses them to perform 64 math operations. That means for every single time the GPU goes to Shared Memory, it executes 4 math operations. By the time it needs to read from memory again, the math cores have been kept so busy that the memory latency is completely hidden. This is how highly optimized matrix multipliers (like cuBLAS) reach 90%+ of the GPU's maximum theoretical speed!