CUDA Shared Memory Reduction

1 Block | 8 Threads | 4 Segments/Thread (N=32)

Global Memory
Shared Memory
Initialization