CUDA Shared Memory Bank Conflicts

Matrix Transpose Optimization (Scaled to 8x8 Tile, 8 Banks)

Step 1 of 7

Global Input

↓

Shared Memory [8][8]

↓

Global Output

Shared Memory Banks

0
1
2
3
4
5
6
7
Standard Kernel
(Conflicts)
Padded Kernel
(Conflict-Free)

Phase 1: Global Read

Description goes here.

Active Kernel Code


                
Execution Step: 0 / 16