CUDA Parallel Reduction
N=16 Elements | 1 Block | Block_Size=8 Threads
Mapping:
Naive (Growing/Scattered)
Interleaved (Growing/Packed)
Global Mem (Shrinking/Packed)
N:
8
16
32
Show Indices
Initial Array Data
Partial Sums
Final Block Sum