An Adaptive Framework for Oversubscription
Description: An Adaptive Framework for Oversubscription Management in CPU-GPU Unified Memory DATE 2021 Debashis Ganguly Rami Melhem, Jun Yang De-facto for General Purpose Computing 2 Integral Part of Computing Platforms 3 ORNLs Summit Supercomputer
Related Topics
Download Presentation
"An Adaptive Framework for Oversubscription" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
slide1. An Adaptive Framework for Oversubscription Management in CPU-GPU Unified Memory DATE 2021 Debashis GangulyRami Melhem, Jun Yang<br>
slide2. De-facto for General Purpose Computing 2<br>
slide3. Integral Part of Computing Platforms 3 ORNL’s Summit Supercomputer
Power Systems AC922, formerly Witherspoon Google Cloud’s HGX A100
NVIDIA HGX A100<br>
slide4. CPU-GPU Heterogeneous Memory Systems 4 Near-memory
Device-memory e.g. HBM
Bandwidth-optimized, but Capacity-constrained
Far-memory
Host-memory e.g. DDR4
Capacity-optimized, but Bandwidth-constrained
Connected by cache-coherent slow interconnect
Non-uniform Memory Access between host and device memories<br>
slide5. Unified Virtual Address Space 5 Micro-architectural support
Hardware page faulting
Page migration engine
Unified Memory Runtime
Responsible for paging in and out of device memory<br>
slide6. Slow CPU-GPU Interconnect 6<br>
slide7. Device Memory Oversubscription 7 Unnecessary page thrashing in application with oversubscribed memory footprint<br>
slide8. Performance Degradation Under Oversubscription 8 Memory management decision relies on ninja programming skill
Depends on intrusive profiling of memory access pattern
Enabled by detailed programmer annotation<br>
slide9. High-level Design Components 9<br>
slide10. Detecting Access Pattern Reuse Irregular Regular No Reuse Random Streaming Sparse Sequential Data Access Order 10<br>
slide11. Defining Access Patterns with Examples 11 Regular FDTD-2D Backpropagation Streaming<br>
slide12. Defining Access Patterns with Examples 12 Random RA SSSP Irregular<br>
slide13. Detecting Patterns in Interconnect Traffic 13 Kernel i Kernel i+1 Alloc. 1 Alloc. 2 0xc0000000 0xc0010000 0xc0020000 0xc0030000 0xc0430000 0xc0420000 0xc0450000 0xc0410000 Alloc. 1 Alloc. 2 0xc0040000 0xc0050000 0xc0060000 0xc0070000 0xc0450000 0xc0470000 0xc0460000 0xc0440000 Linear Linear Random Random No Reuse Reuse Address
---------- 0xc0000000
0xc0010000
0xc0020000
0xc0030000 222960
222961
222962
222963
222964
222965
222965
222970
222971
222972
222974
222975
222977
222977
222978
222979 Timestamp
-------------- 0xc0040000
0xc0050000
0xc0060000
0xc0070000 0xc0430000
0xc0420000
0xc0450000
0xc0410000 0xc0450000
0xc0470000
0xc0460000
0xc0440000<br>
slide14. State-transitions for Managed Allocations 14 U L R M LR RR MR Random Random Linear Linear Linear,
Random Random Linear Random,
Reuse Linear,
Reuse Random Linear Linear,
Random,
Reuse Reuse Reuse Reuse U: Undecided;
L: Linear/Streaming; R: Random; M: Mixed/Irregular;
LR: Linear Reuse/Regular; RR: Random Reuse; MR: Mixed Reuse;<br>
slide15. Workload-specific Memory Management Strategy 15<br>
slide16. Simulation Framework and Benchmarks 16 GPGPU-Sim UVM Smart
Support for Unified Memory extending GPGPU-Sim 3.x
Micro-architecture Modeling
Fault-driven migration, host-pinned access, delayed migration, and
Timing Modeling
Page table walk and TLB, Far-fault handling, PCIe latency and queuing delays
Runtime Modeling
Software prefetchers, page eviction routines
Unified Memory API
cudaMallocManaged, cudaDeviceSynchronize, etc.
Benchmarks
Rodinia, Lonestar, PolyBench, HPC Challenge
https://github.com/DebashisGanguly/gpgpu-sim_UVMSmart<br>
slide17. Performance Improvement by Smart Adaptive Runtime 17 An average 28% and 30% performance improvement under 125% and 150% memory oversubscription<br>
slide18. Reduction in Page Thrashing by Smart Adaptive Runtime 18<br>
slide19. Concluding Remarks 19 Design Principles Application Transparency
No need for intrusive and exhaustive profiling of execution
Reduce the burden of ninja programming
Minimal Implementation Overhead
No need for new hardware extensions
Minimal changes to the runtime
Leverage existing hardware and runtime support Achievements Accommodate large working sets of data-intensive workloads
Comprehensive memory management to reduce performance overhead under device-memory oversubscription<br>
slide20. Debashis GangulyPh.D.
debashis@cs.pitt.edu
https://people.cs.pitt.edu/~debashis/ An Adaptive Framework for Oversubscription Management in CPU-GPU Unified Memory<br>
slide2. De-facto for General Purpose Computing 2<br>
slide3. Integral Part of Computing Platforms 3 ORNL’s Summit Supercomputer
Power Systems AC922, formerly Witherspoon Google Cloud’s HGX A100
NVIDIA HGX A100<br>
slide4. CPU-GPU Heterogeneous Memory Systems 4 Near-memory
Device-memory e.g. HBM
Bandwidth-optimized, but Capacity-constrained
Far-memory
Host-memory e.g. DDR4
Capacity-optimized, but Bandwidth-constrained
Connected by cache-coherent slow interconnect
Non-uniform Memory Access between host and device memories<br>
slide5. Unified Virtual Address Space 5 Micro-architectural support
Hardware page faulting
Page migration engine
Unified Memory Runtime
Responsible for paging in and out of device memory<br>
slide6. Slow CPU-GPU Interconnect 6<br>
slide7. Device Memory Oversubscription 7 Unnecessary page thrashing in application with oversubscribed memory footprint<br>
slide8. Performance Degradation Under Oversubscription 8 Memory management decision relies on ninja programming skill
Depends on intrusive profiling of memory access pattern
Enabled by detailed programmer annotation<br>
slide9. High-level Design Components 9<br>
slide10. Detecting Access Pattern Reuse Irregular Regular No Reuse Random Streaming Sparse Sequential Data Access Order 10<br>
slide11. Defining Access Patterns with Examples 11 Regular FDTD-2D Backpropagation Streaming<br>
slide12. Defining Access Patterns with Examples 12 Random RA SSSP Irregular<br>
slide13. Detecting Patterns in Interconnect Traffic 13 Kernel i Kernel i+1 Alloc. 1 Alloc. 2 0xc0000000 0xc0010000 0xc0020000 0xc0030000 0xc0430000 0xc0420000 0xc0450000 0xc0410000 Alloc. 1 Alloc. 2 0xc0040000 0xc0050000 0xc0060000 0xc0070000 0xc0450000 0xc0470000 0xc0460000 0xc0440000 Linear Linear Random Random No Reuse Reuse Address
---------- 0xc0000000
0xc0010000
0xc0020000
0xc0030000 222960
222961
222962
222963
222964
222965
222965
222970
222971
222972
222974
222975
222977
222977
222978
222979 Timestamp
-------------- 0xc0040000
0xc0050000
0xc0060000
0xc0070000 0xc0430000
0xc0420000
0xc0450000
0xc0410000 0xc0450000
0xc0470000
0xc0460000
0xc0440000<br>
slide14. State-transitions for Managed Allocations 14 U L R M LR RR MR Random Random Linear Linear Linear,
Random Random Linear Random,
Reuse Linear,
Reuse Random Linear Linear,
Random,
Reuse Reuse Reuse Reuse U: Undecided;
L: Linear/Streaming; R: Random; M: Mixed/Irregular;
LR: Linear Reuse/Regular; RR: Random Reuse; MR: Mixed Reuse;<br>
slide15. Workload-specific Memory Management Strategy 15<br>
slide16. Simulation Framework and Benchmarks 16 GPGPU-Sim UVM Smart
Support for Unified Memory extending GPGPU-Sim 3.x
Micro-architecture Modeling
Fault-driven migration, host-pinned access, delayed migration, and
Timing Modeling
Page table walk and TLB, Far-fault handling, PCIe latency and queuing delays
Runtime Modeling
Software prefetchers, page eviction routines
Unified Memory API
cudaMallocManaged, cudaDeviceSynchronize, etc.
Benchmarks
Rodinia, Lonestar, PolyBench, HPC Challenge
https://github.com/DebashisGanguly/gpgpu-sim_UVMSmart<br>
slide17. Performance Improvement by Smart Adaptive Runtime 17 An average 28% and 30% performance improvement under 125% and 150% memory oversubscription<br>
slide18. Reduction in Page Thrashing by Smart Adaptive Runtime 18<br>
slide19. Concluding Remarks 19 Design Principles Application Transparency
No need for intrusive and exhaustive profiling of execution
Reduce the burden of ninja programming
Minimal Implementation Overhead
No need for new hardware extensions
Minimal changes to the runtime
Leverage existing hardware and runtime support Achievements Accommodate large working sets of data-intensive workloads
Comprehensive memory management to reduce performance overhead under device-memory oversubscription<br>
slide20. Debashis GangulyPh.D.
debashis@cs.pitt.edu
https://people.cs.pitt.edu/~debashis/ An Adaptive Framework for Oversubscription Management in CPU-GPU Unified Memory<br>