Explore SIMT Parallelism and DRAM transaction efficiency.
Why is utilization 50% in Coalesced Mode?
In the current simulation logic, for every "Load" step, a 4-thread Warp performs 8 useful float loads (4 threads \(\times\) 2 matrices). At 4 bytes per float, that is 32 bytes of useful data. To get this data, the hardware performs 2 transactions (1 for Matrix A, 1 for Matrix B). Since each transaction moves a 32-byte segment, the total data moved is 64 bytes.
Note: Matrix B is perfectly coalesced (4 items used), but Matrix A is broadcast (1 item used by all), meaning the DRAM segment for A is highly underutilized.
Are we assuming each transaction reads four items?
No, we are assuming each transaction moves a 32-byte segment. Since a float is 4 bytes, each transaction actually moves 8 floats. Because our "Warp" is only 4 threads wide, even a "perfectly coalesced" access only uses 4 out of the 8 floats in that segment (50% of that specific transaction's capacity).
| Warp Lane | Global TID | Coord (i, j) | Read A (Type) | Read B (Type) | Private tmp |
|---|