Occupancy = Active Warps / Max Warps
Image: Coolcaesar, CC BY 4.0, via Wikimedia Commons
Occupancy = Active Warps / Max Warps
CPU cache
L1/L2 cache hierarchy reduces global memory latency
warp divergence kills performance
Warp divergence causes threads to execute non-uniformly, leading to idle cycles and reduced throughput
tensor cores are
Why can computers crunch numbers faster than humans?
CUDA
CUDA enables parallel computation on GPUs
Nvidia
Ever wondered how video games run so smoothly on your computer?
Dynamic random-access memory
DRAM requires periodic refreshing to maintain data integrity
Swipe through 100 ML concepts daily
Open Pocket Polymath