Distributed Systems CS 15-440 Hadoop Lecture 18, October 30, 2019 Mohammad Hammoud Today Last Session: MPI Todays Session: Hadoop Distributed File System and MapReduce Announcements: P3 is out. It is due on November 20 by midnight Quiz II
"Distributed Systems CS 15-440 Hadoop Lecture 18," is the property of its rightful owner. Permission is granted to
download and print the materials on this website for personal, non-commercial use only, and to display it
on your personal computer provided you do not modify the materials and that you retain all copyright
notices contained in the materials. By downloading content from our website, you accept the terms of this
agreement.
Presentation Transcript
01
Distributed SystemsCS 15-440 Hadoop
Lecture 18, October 30, 2019
Mohammad Hammoud<br>
02
Today Last Session:
MPI
Today’s Session:
Hadoop Distributed File System and MapReduce
Announcements:
P3 is out. It is due on November 20 by midnight
Quiz II is on Wednesday, November 13<br>
03
We Live in a World of Data…<br>
04
What Do We Do With Big Data? Store Access Encrypt We want to do all these seamlessly... Share Process …. and more!<br>
05
Where to Store Big Data? The underlying storage system is a key component for enabling Big Data querying/mining/analytics
Typically, the storage system would “partition” and “distribute” Big Data, using striping (or partitioning) and placement techniques
This allows for concurrent accesses to data
as well as improves fault-tolerance 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Logical File Stripe Size Striping Unit Server 1 Server 2 Server 3 Server 4 0 1 2 3 4 5 6 10 14 7 11 15 9 13 8 12<br>
06
Example: The Google File System GFS paritions large files into fixed-size blocks and distributes them randomly across cluster machines Server 0
(Writer) Blk 0 Blk 1 Blk 2 Blk 3 Blk 4 Blk 5 Blk 6 Server 1 Blk 0 Blk 2 Blk 3 Blk 3 0M 64M 128M 192M 256M 320M 384M Server 2 Server 3 Blk 1 Blk 2 Blk 4 Blk 6 Blk 0 Blk 1 Blk 4 Blk 5 Blk 6 Large File Blk 0 Blk 1 Blk 2 Blk 3 Blk 4 Blk 5 Blk 6 Blk 5<br>
07
Example: The Google File System GFS adopts a master-slave architecture GFS client Master Chunk Server Linux File System Chunk Server Linux File System Chunk Server Linux File System File name Contact address Chunk Id, range Chunk data<br>
08
How to Process Big Data? One alternative: Create a custom distributed system (or program) for each new algorithm
Cumbersome!
Another alternative: utilize modern distributed analytics frameworks, which:
Relieve programmers from concerns with many of the difficult aspects of developing distributed programs
Allow programmers to focus on ONLY the sequential parts of their programs
Examples:
Hadoop MapReduce
Google’s Pregel
CMU’s Distributed GraphLab<br>
09
Distributed Analytics Frameworks Hadoop MapReduce Introduction Programming Model Execution Model Architectural & Scheduling
Models<br>
10
Hadoop Hadoop is one of the most successful realizations of large-scale “data-parallel” distributed analytics frameworks
Hadoop MapReduce is an open source implementation of Google’s MapReduce
Hadoop uses Hadoop Distributed File System (HDFS) as a storage layer
Distributed Analytics Frameworks Hadoop MapReduce Introduction Programming Model Execution Model Architectural & Scheduling
Models<br>
13
The Programming Model Hadoop MapReduce employs a shared-based programming model, which entails that:
Tasks can interact (if needed) via reading and writing to a shared space
HDFS provides the shared space for all Map and Reduce tasks
Programmers write only sequential code, without defining functions that send/receive messages between tasks MT1 MT2 MT3 MT4 MT5 MT6 A Shared Address Space (Provided by HDFS) RT1 RT2 RT3 “Implicit” communication (provided by the MapReduce Engine)- Programmers do not write or call any communication routines A Shared Address Space (Provided by HDFS)<br>
14
Example: Word Count Mohammad is
delivering a
lecture to the
15-440 class
The course
name of 15-440
is Distributed Systems A Text File Mohammad is
delivering a
lecture to the
15-440 class The course
name of 15-440
is Distributed Systems A Chunk of File A Chunk of File A Map Function Parse &
Count Iterate&
Sum Parse &
Count A Map Function A Reduce
Function<br>
15
Distributed Analytics Frameworks Hadoop MapReduce Introduction Programming Model Execution Model Architectural & Scheduling
Models<br>
16
The Execution Model Hadoop MapReduce adopts a synchronous execution model
A distributed program (or system) is said to be synchronous if and only if its constituent tasks operate in a lock-step mode
No two tasks can run concurrently under two different iterations
In MapReduce:
Each iteration is treated as a MapReduce job
A job can encompass 1 or many Map tasks and 0 or many Reduce tasks
Programs with multiple iterations (i.e., iterative programs) are executed using multiple chained MapReduce jobs
When all Reduce tasks within job i are committed, a new job i + 1 is started (if any)
Hence, two different tasks cannot run in parallel under two different jobs (or iterations)<br>
17
Distributed Analytics Frameworks Hadoop MapReduce Introduction Programming Model Execution Model Architectural & Scheduling
Models<br>
18
The Architectural and Scheduling Models Hadoop MapReduce employs a master-slave architecture
A pull-based task scheduling strategy is used, whereby:
Map tasks are scheduled in proximity of HDFS blocks
Reduce tasks are scheduled anywhere Core Switch TaskTracker1 Request a Map Task Schedule a Map Task at an Empty Map Slot on TaskTracker1 Rack Switch 1 Rack Switch 2 TaskTracker2 TaskTracker3 TaskTracker4 TaskTracker5 JobTracker MT1 MT2 MT3 MT2 MT3 A slave The master<br>
19
The Architectural and Scheduling Models Hadoop MapReduce employs a master-slave architecture
With the above setup, how many Map tasks can run in parallel?
Each TaskTracker has by default two Map slots, thus can run two Map tasks concurrently
With 4 TaskTrackers and 2 Map slots on each TaskTracker, 8 Map tasks can be executed in parallel
The maximum number of Map tasks that can run in parallel is denoted as Map wave Core Switch TaskTracker1 Request a Map Task Schedule a Map Task at an Empty Map Slot on TaskTracker1 Rack Switch 1 Rack Switch 2 TaskTracker2 TaskTracker3 TaskTracker4 TaskTracker5 JobTracker MT1 MT2 MT3 MT2 MT3 A slave The master<br>
20
The Architectural and Scheduling Models Hadoop MapReduce employs a master-slave architecture
For a dataset with a size of 1024MB, how many Map waves are needed?
The size of each HDFS block is by default 64MB and each split encompasses by default 1 HDFS block
Hence, there will be a total of 1024/64 = 16 HDFS blocks or 16 splits
The input to each Map task is a single split, thus there will be a total of 16 Map tasks
Therefore, 16 tasks/8 slots = 2 Map waves will be needed Core Switch TaskTracker1 Request a Map Task Schedule a Map Task at an Empty Map Slot on TaskTracker1 Rack Switch 1 Rack Switch 2 TaskTracker2 TaskTracker3 TaskTracker4 TaskTracker5 JobTracker MT1 MT2 MT3 MT2 MT3 A slave The master<br>