SGEMM: Ultimate Memory Visualizer

Explore SIMT Parallelism and DRAM transaction efficiency.

DRAM Transactions
0
Effective Bandwidth
0%

Execution Controls

Init
Animation Speed 500ms
// Mapping: Coalesced
const uint j = blockIdx.x * bDim.x + tIdx.x;
const uint i = blockIdx.y * bDim.y + tIdx.y;
for (int k = 0; k < K; ++k) {
tmp += A[i*K+k] * B[k*N+j];
}
C[i*N+j] = tmp + C[i*N+j];
Matrix A (6 × 6)
Broadcast
Matrix B (6 × 8)
Coalesced
Matrix C (6 × 8)
Thread Computing Result

Global Memory (DRAM) Layout & Bandwidth

A-DRAM (Row-Major)
B-DRAM (Row-Major)
C-DRAM (Row-Major)
Bus Utilization
0%
Useful: 0B
Moved: 0B

Memory Analysis Deep Dive

Why is utilization 50% in Coalesced Mode?

In the current simulation logic, for every "Load" step, a 4-thread Warp performs 8 useful float loads (4 threads \(\times\) 2 matrices). At 4 bytes per float, that is 32 bytes of useful data. To get this data, the hardware performs 2 transactions (1 for Matrix A, 1 for Matrix B). Since each transaction moves a 32-byte segment, the total data moved is 64 bytes.

\[\text{Utilization} = \frac{32 \text{ useful bytes}}{64 \text{ total bytes moved}} = 50\%\]

Note: Matrix B is perfectly coalesced (4 items used), but Matrix A is broadcast (1 item used by all), meaning the DRAM segment for A is highly underutilized.

Are we assuming each transaction reads four items?

No, we are assuming each transaction moves a 32-byte segment. Since a float is 4 bytes, each transaction actually moves 8 floats. Because our "Warp" is only 4 threads wide, even a "perfectly coalesced" access only uses 4 out of the 8 floats in that segment (50% of that specific transaction's capacity).

Summary
Bus Efficiency: 100%

Execution Monitor

Parallel SIMT
Warp Lane Global TID Coord (i, j) Read A (Type) Read B (Type) Private tmp