AbstractChapter Outline4.1 Architecture of a modern GPU4.2 Block scheduling4.3 Synchronization and transparent scalability4.4 Warps and SIMD hardware4.5 Control divergence4.6 Warp scheduling and latency tolerance4.7 Resource partitioning and occupancy4.8 Querying device properties4.9 SummaryExercisesReferences