CUDA Algorithm Lab
Interactive, step-by-step visualizations of how GPU algorithms map threads to data. Each one goes with a runnable, self-checking CUDA example in the repository.
Matrix multiplication (GEMM)
- Naive kernel memory patternsStep 00: one thread per output, uncoalesced loads
- Thread-to-memory mappingStep 01: how a warp's addresses become coalesced
- Shared memory tilingStep 02: staging tiles of A and B in shared memory
- 1D register tilingStep 03: each thread computes a column of outputs
- Register tiling for a single threadStep 03: one thread's loads and FMAs, up close
- Matrix multiplication modesInner product versus outer product formulations
- 2D register tilingStep 04: each thread computes an 8 x 8 patch
Reduction
- Naive reduction treeInterleaved addressing and its divergent warps
- Shared memory reductionSequential addressing without bank conflicts