Multi-Resource Packing for Cluster Schedulers
DJ
Published · 38 slides · 0 views
1 / 1
Description
Multi-Resource Packing for Cluster Schedulers Robert Grandl, Ganesh Ananthanarayanan, Srikanth Kandula, Sriram Rao, Aditya Akella Tetris Performance of cluster schedulers We find that: 1Time to finish a set of jobs Resources are fragmented
Related Topics
Share
Embed code
Download this presentation From Below
"Multi-Resource Packing for Cluster Schedulers" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
01
Multi-Resource Packing for Cluster Schedulers Robert Grandl, Ganesh Ananthanarayanan,
Srikanth Kandula, Sriram Rao, Aditya Akella Tetris<br>
Srikanth Kandula, Sriram Rao, Aditya Akella Tetris<br>
02
Performance of cluster schedulers We find that: 1Time to finish a set of jobs Resources are fragmented i.e. machines run below capacity Even at 100% usage, goodput is smaller due to over-allocation Pareto-efficient multi-resource fair schemes do not lead to good avg. performance Tetris
Up to 40% improvement in makespan1 and job completion time with near-perfect fairness<br>
Up to 40% improvement in makespan1 and job completion time with near-perfect fairness<br>
03
Findings from Bing and Facebook traces analysis Tasks need varying amounts of each resource Demands for resources are weakly correlated Applications have (very) diverse resource needs Multiple resources become tight This matters, because no single bottleneck resource in the cluster: E.g., enough cross-rack network bandwidth to use all cores 3 Upper bound on potential gains Makespan reduces by ≈ 49% Avg. job completion time reduces by ≈ 46%<br>
04
4 Why so bad #1 Production schedulers neither pack tasks nor consider all their relevant resource demands #1 Resource Fragmentation #2 Over-allocation<br>
05
Current Schedulers “Packer” Scheduler Machine A
4 GB Memory Machine B
4 GB Memory T1: 2 GB T3: 4 GB T2: 2 GB Time Resource Fragmentation (RF) Machine A
4 GB Memory Machine B
4 GB Memory T1: 2 GB T3: 4 GB T2: 2 GB Time Avg. task compl. time = 1 t 5 Current Schedulers RF increase with the number of resources being allocated ! Avg. task compl.time = 1.33 t Allocate resources per slots, fairness. Are not explicit about packing.<br>
4 GB Memory Machine B
4 GB Memory T1: 2 GB T3: 4 GB T2: 2 GB Time Resource Fragmentation (RF) Machine A
4 GB Memory Machine B
4 GB Memory T1: 2 GB T3: 4 GB T2: 2 GB Time Avg. task compl. time = 1 t 5 Current Schedulers RF increase with the number of resources being allocated ! Avg. task compl.time = 1.33 t Allocate resources per slots, fairness. Are not explicit about packing.<br>
06
Current Schedulers “Packer” Scheduler Machine A
4 GB Memory; 20 MB/s Nw. Time T1: 2 GB
Memory 20 MB/s Nw. T2: 2 GB
Memory 20 MB/s Nw. T3: 2 GB
Memory Machine A
4 GB Memory; 20 MB/s Nw. Time T1: 2 GB
Memory 20 MB/s Nw. T2: 2 GB
Memory 20 MB/s Nw. T3: 2 GB
Memory 20 MB/s Nw. 20 MB/s Nw. 6 Over-Allocation Not all of the resources are explicitly allocated E.g.,disk and network can be over-allocated Avg. task compl.time= 2.33 t Avg. task compl. time = 1.33 t Current Schedulers<br>
4 GB Memory; 20 MB/s Nw. Time T1: 2 GB
Memory 20 MB/s Nw. T2: 2 GB
Memory 20 MB/s Nw. T3: 2 GB
Memory Machine A
4 GB Memory; 20 MB/s Nw. Time T1: 2 GB
Memory 20 MB/s Nw. T2: 2 GB
Memory 20 MB/s Nw. T3: 2 GB
Memory 20 MB/s Nw. 20 MB/s Nw. 6 Over-Allocation Not all of the resources are explicitly allocated E.g.,disk and network can be over-allocated Avg. task compl.time= 2.33 t Avg. task compl. time = 1.33 t Current Schedulers<br>
07
Work Conserving != no fragmentation, over-allocation Treat cluster as a big bag of resources Hides the impact of resource fragmentation Assume job has a fixed resource profile Different tasks in the same job have different demands Multi-resource Fairness Schemes do not solve the problem Why so bad #2 How the job is scheduled impacts jobs’ current resource profiles Can schedule to create complementarity Example in paper
Packer vs. DRF: makespan and avg. completion time improve by over 30% Pareto1 efficient != performant 1no job can increase its share without decreasing the share of another 7<br>
Packer vs. DRF: makespan and avg. completion time improve by over 30% Pareto1 efficient != performant 1no job can increase its share without decreasing the share of another 7<br>
08
Competing objectives Job completion time Fairness vs. Cluster efficiency vs. 8<br>
09
Tetris 9<br>
10
Theory Practice Multi-Resource Packing of Tasks
similar to
Multi-Dimensional Bin Packing Balls could be tasks
Bin could be machine, time 1APX-Hard is a strict subset of NP-hard APX-Hard1 Existing heuristics do not directly apply: Assume balls of a fixed size Assume balls are known apriori 10 vary with time / machine placed elastic cope with online arrival of jobs, dependencies, cluster activity Avoiding fragmentation looks like: Tight bin packing Reduce # of bins reduce makespan<br>
similar to
Multi-Dimensional Bin Packing Balls could be tasks
Bin could be machine, time 1APX-Hard is a strict subset of NP-hard APX-Hard1 Existing heuristics do not directly apply: Assume balls of a fixed size Assume balls are known apriori 10 vary with time / machine placed elastic cope with online arrival of jobs, dependencies, cluster activity Avoiding fragmentation looks like: Tight bin packing Reduce # of bins reduce makespan<br>
11
# 1 Packing heuristic Packing tasks to machines = Multi-Dimensional Bin Packing
Ball = Task resource demands vector
Bin = Machine available resource vector 1. Check for fit to ensure no over-allocation Alignment score (A) 11 A packing heuristic Tasks resources demand vector Machine resource vector < Fit “A” works because: 2. Bigger balls get bigger scores 3. Abundant resources used first<br>
Ball = Task resource demands vector
Bin = Machine available resource vector 1. Check for fit to ensure no over-allocation Alignment score (A) 11 A packing heuristic Tasks resources demand vector Machine resource vector < Fit “A” works because: 2. Bigger balls get bigger scores 3. Abundant resources used first<br>
12
Tetris 12<br>
13
13 CHALLENGE # 2 Shortest Remaining Time First1 (SRTF) 1SRTF – M. Harchol-Balter et al. Connection Scheduling in Web Servers [USITS’99] schedules jobs in ascending order of their remaining time Job Completion Time Heuristic Q: What is the shortest “remaining time” ? “remaining work” remaining # tasks tasks’ durations tasks’ resource demands & & = A job completion time heuristic Gives a score P to every job Extended SRTF to incorporate multiple resources<br>
14
14 CHALLENGE # 2 Job Completion Time Heuristic Combine A and P scores ! Packing Efficiency Completion Time ? A: delays job completion time P: loss in packing efficiency<br>
15
Tetris 15<br>
16
# 3 16 Packer says: “task T should go next to improve packing efficiency” Possible to satisfy all three
In fact, happens often in practice SRTF says: “schedule job J to improve avg. completion time” Fairness says: “this set of jobs should be scheduled next” Fairness Heuristic Performance and fairness do not mix well in general But …. We can get “perfect fairness” and much better performance<br>
In fact, happens often in practice SRTF says: “schedule job J to improve avg. completion time” Fairness says: “this set of jobs should be scheduled next” Fairness Heuristic Performance and fairness do not mix well in general But …. We can get “perfect fairness” and much better performance<br>
17
# 3 17 Fairness Knob, F [0, 1) Pick the best-for-perf. task from among
1-F fraction of jobs furthest from fair share Fairness Heuristic Fairness is not a tight constraint Long term fairness not short term fairness Lose a bit of fairness for a lot of gains in performance Heuristic F = 0 F → 1<br>
1-F fraction of jobs furthest from fair share Fairness Heuristic Fairness is not a tight constraint Long term fairness not short term fairness Lose a bit of fairness for a lot of gains in performance Heuristic F = 0 F → 1<br>
18
18 Putting it all together We saw: Other things in the paper: Packing efficiency Prefer small remaining work Fairness knob Estimate task demands Deal with inaccuracies, barriers Other cluster activities Job Manager1 Node Manager1 Cluster-wide Resource Manager Multi-resource asks; barrier hint Track resource usage; enforce allocations New logic to match tasks to machines (+packing, +SRTF, +fairness) Allocations Asks Offers Resource
availability reports Yarn architecture Changes to add Tetris(shown in orange)<br>
availability reports Yarn architecture Changes to add Tetris(shown in orange)<br>
19
Evaluation Implemented in Yarn 2.4
250 machine cluster deployment
Bing and Facebook workload 19<br>
250 machine cluster deployment
Bing and Facebook workload 19<br>
20
20 Efficiency Makespan Multi-resource Scheduler 28 % Avg. Job Compl. Time 35% Tetris Gains from avoiding fragmentation avoiding over-allocation Tetris vs. Single Resource Scheduler 29 % 30 % Single Resource Scheduler<br>
21
21 Fairness Fairness Knob quantifies the extent to which Tetris adheres to fair allocation No Fairness
F = 0 Makespan 50 % 10 % 25 % Job Compl.
Time 40 % 23 % 35 % Avg. Slowdown
[over impacted jobs] 25 % 2 % 5 % Full Fairness
F → 1 F = 0.25<br>
F = 0 Makespan 50 % 10 % 25 % Job Compl.
Time 40 % 23 % 35 % Avg. Slowdown
[over impacted jobs] 25 % 2 % 5 % Full Fairness
F → 1 F = 0.25<br>
22
Tetris Pack efficiently along multiple resources Prefer jobs with less “remaining work” Incorporate Fairness Combine heuristics that improve packing efficiency with those that lower average job completion time Achieving desired amounts of fairness can coexist with improving cluster performance Implemented inside YARN; deployment and trace-driven simulations show encouraging initial results We are working towards a Yarn check-in
http://research.microsoft.com/en-us/UM/redmond/projects/tetris/ 22<br>
http://research.microsoft.com/en-us/UM/redmond/projects/tetris/ 22<br>
23
23 Backup slides<br>
24
Estimating resource requirements Estimating Resource Demands Observe usage and backfill finished tasks in the same phase We estimate peak usage from reports resource usages of tasks and other cluster activity e.g., evacuation Resource Tracker collecting statistics from recurring jobs inputs size/location of tasks 24 Placement impacts usages of network and disk.<br>
25
25 Incorporating task placement 1. Disk and network demands depend on task placement 2. Remote resources cannot directly be included in dot-product Resource vectors increase with the number of machines in the cluster Can prefer remote placement! Compute packing score on local resources; use a fractional penalty to reduce use of remote resources Sensitivity Analysis Makespan & completion time change little for remote penalty [6%, 15%]<br>
26
26 Alternative Packing Heuristics di – task demand along dimens. i ai – avail. res. along dimens. i<br>
27
27 Virtual Machine Packing != Tetris Virtual Machine Packing But focus on different challenges and not task packing: balance load across servers ensure VM availability inspite of failures allow for quick software and hardware updates NO corresponding entity to a job and hence job completion time is inexpressible Explicit resource requirements (e.g. small VM) makes VM packing simpler Consolidating VMs, with multi-dimensional resource requirements, on to the fewest number of servers<br>
28
28 Barrier knob, b [0, 1) Tetris gives preference for last tasks in a stage Offer resources to tasks in a stage preceding a barrier, where b fraction of tasks have finished b = 1 no tasks preferentially treated<br>
29
29 Weighting Alignment vs. SRTF While the best choice of depends on the workload, we found that: gains from packing efficiency are only moderate sensitive to improving avg. job compl. time requires > 0, though gains stabilize quickly Sensitivity Analysis<br>
30
30 Cluster load vs. Tetris performance<br>
31
31 Starvation Prevention It could take a long time to accommodate large tasks
Working on a more principled solution But … most tasks have demands within one order of magnitude Free resources become available in “large clumps”
periodic availability reports
scheduler learns about resources freed up by all tasks that finish in the preceding period in one shot<br>
Working on a more principled solution But … most tasks have demands within one order of magnitude Free resources become available in “large clumps”
periodic availability reports
scheduler learns about resources freed up by all tasks that finish in the preceding period in one shot<br>
32
32 Workload analysis<br>
33
33 Ingestion / evacuation ingestion = storing incoming data for later analytics evacuation = data evacuated and re-replicated before
maintenance operations E.g., some clusters reports volumes of up to 10 TB per hour Other cluster activities which produce background traffic E.g., rack decommission for machines re-imaging Resource Tracker reports, used by Tetris to avoid contention between its tasks and these activities<br>
maintenance operations E.g., some clusters reports volumes of up to 10 TB per hour Other cluster activities which produce background traffic E.g., rack decommission for machines re-imaging Resource Tracker reports, used by Tetris to avoid contention between its tasks and these activities<br>
34
34 Fairness vs. Efficiency<br>
35
35 Fairness vs. Efficiency<br>
36
Packer Scheduler vs. DRF DRF Scheduler Packer Schedulers 2 tasks Job Schedule Resources used 2 tasks 2 tasks 2 tasks 2 tasks 2 tasks 6 tasks 6 tasks 6 tasks A B C 18 cores 16 GB 18 cores 16 GB 18 cores 16 GB t 2t 3t 0 tasks Job Schedule Resources used 0 tasks 6 tasks 0 tasks 6 tasks 18 tasks A B C 18 cores 18 cores 6 GB 18 cores 6 GB t 2t 3t 36 GB 33% improvement Dominant Resource Fairness (DRF) computes the dominant share (DS) of every user and seeks to maximize the minimum DS across all users 36<br>
37
1Time to finish a set of jobs Resources used 4 cores 6 GB 2 tasks 2 tasks 2 tasks 2 tasks t 2t 3t 4t Job Schedule 4 cores 6 GB 4 cores 6 GB 2 cores 4 GB Resources used 2 cores 4 GB 2 tasks 2 tasks 2 tasks 2 tasks t 2t 3t 4t Job Schedule 4 cores 6 GB 4 cores 6 GB 4 cores 6 GB Pack No Pack 29% improvement 37 Packing efficiency does not achieve everything Achieving packing efficiency does not necessarily improve job completion time<br>
38
38<br>