NVTX Markers and Profiling ToolsProfiling PyTorch to Identify BottlenecksUsing PyTorch ProfilerSystem Profiling with Nsight Systems and NVTX TimelinesKernel Roofline Analysis for General Matrix Multiply (GEMM)CPU and GPU Profiling with Linux perfPyTorch Compiler (torch.compile)Using the PyTorch CompilerCompiling Versus Writing Custom KernelsCompilation Modes and Trade-Offs in Speed, Memory, and Compile TimeRegional CompilationProfiling and Debugging Compiler Performance IssuesPyTorch Optimized Attention MechanismsPyTorch Architecture Optimization (torchao), Quantization, Sparsity, and PruningConcurrency with CUDA StreamsOverlapping Communication and ComputationStream Synchronization with EventsUsing CUDA Streams with MoE ModelsReducing Kernel Launch Overhead with CUDA GraphsCapturing a CUDA Graph and Preallocating MemoryReplaying the GraphBest Practices for CUDA GraphsCUDA Graph Trees (PyTorch Compiler Internal)Profiling and Tuning Memory in PyTorchTuning the CUDA Memory AllocatorActivation Checkpointing for Memory SavingsOffloading Parameters to CPU and NVMeSuperOffload: Optimized CPU-GPU Superchip OffloadFSDP Automatic Checkpointing and OffloadingCombining FSDP with Tensor Parallel and Pipeline ParallelPluggable Memory Allocators and Cross-GPU Data TransfersEnabling Peer-to-Peer DMA and UCXPyTorch Symmetric MemoryOptimizing the Data Input PipelineScaling with PyTorch DistributedDDP with torch.compileFSDP with torch.compileTensor and Pipeline Parallelism with torch.compileTorchTitan, AsyncTP, AutoParallel, and SimpleFSDPMulti-GPU Profiling with HTAContinuous Integration and Performance BenchmarkingPyTorch HUD Performance DashboardPerformance Benchmarks and MLPerf LoggingKey TakeawaysConclusion