Database Systems: Design, Implementation, and
SB
Published · 43 slides · 0 views
1 / 1
Description
Database Systems: Design, Implementation, and Management, 14e Module 14: Big Data and NoSQL Coronel, Carlos and Morris, Steven, Database Systems: Design, Implementation, and Management, 14 Edition. 2023 Cengage. All Rights Reserved. May
Related Topics
Share
Embed code
Download this presentation From Below
"Database Systems: Design, Implementation, and" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
01
Database Systems: Design, Implementation, and Management, 14e Module 14: Big Data and NoSQL Coronel, Carlos and Morris, Steven, Database Systems: Design, Implementation, and Management, 14 Edition. © 2023 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.<br>
02
Chapter Objectives (1 of 2) By the end of this chapter, you should be able to:
Explain the role of Big Data in modern business
Describe the primary characteristics of Big Data and how these go beyond the traditional “3 Vs”
Explain how the core components of the Hadoop framework operate
Identify the major components of the Hadoop ecosystem
Summarize the four major approaches of the NoSQL data model and how they differ from the relational model<br>
Explain the role of Big Data in modern business
Describe the primary characteristics of Big Data and how these go beyond the traditional “3 Vs”
Explain how the core components of the Hadoop framework operate
Identify the major components of the Hadoop ecosystem
Summarize the four major approaches of the NoSQL data model and how they differ from the relational model<br>
03
Chapter Objectives (2 of 2) By the end of this chapter, you should be able to (continued):
Describe the characteristics of NewSQL databases
Understand how to work with document databases using MongoDB
Understand how to work with graph databases using Neo4j<br>
Describe the characteristics of NewSQL databases
Understand how to work with document databases using MongoDB
Understand how to work with graph databases using Neo4j<br>
04
Big Data Big Data refers to a set of data that displays the characteristics of volume, velocity, and variety (the 3 Vs) to an extent that makes the data unsuitable for management by a relational DBMS
These characteristics can be defined as follows:
Volume – the quantity of data to be stored
Velocity – the speed at which data is entering the system
Variety – the variations in the structure of the data to be stored<br>
These characteristics can be defined as follows:
Volume – the quantity of data to be stored
Velocity – the speed at which data is entering the system
Variety – the variations in the structure of the data to be stored<br>
05
Volume Volume, the quantity of data to be stored, is a key characteristic of Big Data
Scaling up is keeping the same number of systems but migrating each one to a larger system
Scaling out means that when the workload exceeds server capacity, it is spread out across a number of servers<br>
Scaling up is keeping the same number of systems but migrating each one to a larger system
Scaling out means that when the workload exceeds server capacity, it is spread out across a number of servers<br>
06
Velocity (1 of 2) Velocity refers to the rate at which new data enters the system as well as the rate at which the data must be processed
The velocity of processing can be broken down into two categories:
Stream processing focuses on input processing and requires analysis of data stream as it enters the system
Scientists have created algorithms to decide ahead of time which data will be kept
Feedback loop processing refers to the analysis of the data to produce actionable results<br>
The velocity of processing can be broken down into two categories:
Stream processing focuses on input processing and requires analysis of data stream as it enters the system
Scientists have created algorithms to decide ahead of time which data will be kept
Feedback loop processing refers to the analysis of the data to produce actionable results<br>
07
Velocity (2 of 2) Figure 14.3 Feedback Loop Processing<br>
08
Variety Variety refers to the vast array of formats and structures in which data may be captured
Structured data is data that has been organized to fit a predefined data model
Unstructured data is data that is not organized to fit into a predefined data model
Semistructured data combines elements of both – some parts of the data fit a predefined model while other parts do not
Relational databases rely on structured data
One advantage of providing structure is the flexibility of being able to structure the data in different ways for different applications<br>
Structured data is data that has been organized to fit a predefined data model
Unstructured data is data that is not organized to fit into a predefined data model
Semistructured data combines elements of both – some parts of the data fit a predefined model while other parts do not
Relational databases rely on structured data
One advantage of providing structure is the flexibility of being able to structure the data in different ways for different applications<br>
09
Other Characteristics Variability refers to the changes in the meaning of data based on context
Sentimental analysis is a method of text analysis that attempts to determine if a statement conveys a positive, negative, or neutral attitude about a topic
Veracity refers to the trustworthiness of data
Value refers to the degree to which the data can be analyzed for meaningful insight
Visualization is the ability to graphically resent data to make it understandable
Polyglot persistence is the coexistence of a variety of data storage and management technologies within an organization’s infrastructure<br>
Sentimental analysis is a method of text analysis that attempts to determine if a statement conveys a positive, negative, or neutral attitude about a topic
Veracity refers to the trustworthiness of data
Value refers to the degree to which the data can be analyzed for meaningful insight
Visualization is the ability to graphically resent data to make it understandable
Polyglot persistence is the coexistence of a variety of data storage and management technologies within an organization’s infrastructure<br>
10
Hadoop De facto standard for most Big Data storage and processing
Hadoop is a Java-based framework for distributing and processing very large data sets across clusters of computers
The two most important components include the following:
Hadoop Distributed File System (HDFS) is a low-level distributed file processing system that can be used directly for data storage
MapReduce is a programming model that supports processing large data sets<br>
Hadoop is a Java-based framework for distributing and processing very large data sets across clusters of computers
The two most important components include the following:
Hadoop Distributed File System (HDFS) is a low-level distributed file processing system that can be used directly for data storage
MapReduce is a programming model that supports processing large data sets<br>
11
HDFS (1 of 3) The Hadoop Distributed File System (HDFS) approach to distributing is based on the following key assumptions:
High volume – Hadoop has a default block sizes is 64 MB and can be configured to even larger values
Write-once, read-many: this model simplifies concurrency issues and improves data throughput
Streaming access: Hadoop is optimized for batch processing of entire files as a continuous stream of data
Fault tolerance: Hadoop is designed to replicate data across many different devices so that when one fails, data is still available from another device<br>
High volume – Hadoop has a default block sizes is 64 MB and can be configured to even larger values
Write-once, read-many: this model simplifies concurrency issues and improves data throughput
Streaming access: Hadoop is optimized for batch processing of entire files as a continuous stream of data
Fault tolerance: Hadoop is designed to replicate data across many different devices so that when one fails, data is still available from another device<br>
12
HDFS (2 of 3) Hadoop uses several types of nodes, which are computers that perform one or more types of tasks within the system
Data nodes store the actual file data
The name node contains file system metadata
The client node makes requests to the file system as needed to support user applications
The data node communicates with the name node and sends block reports and heartbeats
A block report is sent every 6 hours and informs the name node which blocks are on that data node
A heartbeat is used to let the name node know that the data node is still available<br>
Data nodes store the actual file data
The name node contains file system metadata
The client node makes requests to the file system as needed to support user applications
The data node communicates with the name node and sends block reports and heartbeats
A block report is sent every 6 hours and informs the name node which blocks are on that data node
A heartbeat is used to let the name node know that the data node is still available<br>
13
HDFS (3 of 3) Figure 14.4 Hadoop Distributed File System (HDFS)<br>
14
MapReduce MapReduce is the computing framework used to process large data sets across clusters
A map function takes a collection of data and sorts and filters it into a set of key-value pairs
The map function is performed by a program called a mapper
A reduce function summarizes the results of the map function into a single result
The reduce function is performed by a program called a reducer
The implementation of MapReduce complements the HDFS structure
Job tracker is a central control program used to report on MapReduce processing jobs
Task tracker is a program responsible for running map and reduce tasks on a node
Batch processing runs tasks from beginning to end with no user interaction<br>
A map function takes a collection of data and sorts and filters it into a set of key-value pairs
The map function is performed by a program called a mapper
A reduce function summarizes the results of the map function into a single result
The reduce function is performed by a program called a reducer
The implementation of MapReduce complements the HDFS structure
Job tracker is a central control program used to report on MapReduce processing jobs
Task tracker is a program responsible for running map and reduce tasks on a node
Batch processing runs tasks from beginning to end with no user interaction<br>
15
Hadoop Ecosystem (1 of 3) Most organizations that use Hadoop also use a set of other related products that interact and complement each other to produce an entire ecosystem of applications and tools
Like any ecosystem, the interconnected pieces are constantly evolving and their relationships are changing, so it is a rather fluid situation
MapReduce Simplification Applications
Hive is a data warehousing system that sits on top of HDFS and supports its own SQL-like language
Pig is a tool for compiling a high-level scripting language, named Pig Latin, into MapReduce jobs for executing in Hadoop<br>
Like any ecosystem, the interconnected pieces are constantly evolving and their relationships are changing, so it is a rather fluid situation
MapReduce Simplification Applications
Hive is a data warehousing system that sits on top of HDFS and supports its own SQL-like language
Pig is a tool for compiling a high-level scripting language, named Pig Latin, into MapReduce jobs for executing in Hadoop<br>
16
Hadoop Ecosystem (2 of 3) Data Ingestion Applications
Flume is a component for ingesting data in Hadoop
Sqoop is a tool for converting data back and forth between a relational database and the HDFS
Direct Query Applications
Hbase is a column-oriented NoSQL database designed to sit on top of the HDFS that quickly processes sparse datasets
Impala was the first SQL on Hadoop application<br>
Flume is a component for ingesting data in Hadoop
Sqoop is a tool for converting data back and forth between a relational database and the HDFS
Direct Query Applications
Hbase is a column-oriented NoSQL database designed to sit on top of the HDFS that quickly processes sparse datasets
Impala was the first SQL on Hadoop application<br>
17
Hadoop Ecosystem (3 of 3) Figure 14.6 A Sample of the Hadoop Ecosystem<br>
18
Hadoop Pushback Many organizations benefit from having a customized Hadoop ecosystem that is tailored to their specific needs in a manner that no other solution can duplicate
However, the learning curve can be steep
Companies such as IBM and Cloudera offer out-of-the-box Hadoop ecosystems called data platforms
The perceived complications of Hadoop have helped to propel interest in alternative solutions, such as NoSQL databases<br>
However, the learning curve can be steep
Companies such as IBM and Cloudera offer out-of-the-box Hadoop ecosystems called data platforms
The perceived complications of Hadoop have helped to propel interest in alternative solutions, such as NoSQL databases<br>
19
Knowledge Check Activity 14-1 What is Big Data? Give a brief definition.<br>
20
Knowledge Check Activity 14-1: Answer What is Big Data? Give a brief definition.
Answer: Big Data is data of such volume, velocity, and/or variety that it is difficult for traditional relational database technologies to store and process it.<br>
Answer: Big Data is data of such volume, velocity, and/or variety that it is difficult for traditional relational database technologies to store and process it.<br>
21
NoSQL NoSQL is the name given to a broad array of nonrelational database technologies that have developed to address Big Data challenges
The name does not describe what the NoSQL technologies are, but rather what they are not
There are hundreds of products that can be considered as being under the broadly defined term NoSQL
Most fit into one of four categories: key-value data stores, document databases, column-oriented databases, and graph databases<br>
The name does not describe what the NoSQL technologies are, but rather what they are not
There are hundreds of products that can be considered as being under the broadly defined term NoSQL
Most fit into one of four categories: key-value data stores, document databases, column-oriented databases, and graph databases<br>
22
Key-Value Databases Key-value (KV) databases are conceptually the simplest of the NoSQL data models
A KV database is a NoSQL database that stores data as a collection of key-value pairs
Key-value pairs are typically organized into “bucket”
A bucket can roughly be thought of as the KV database equivalent of a table
A bucket is a logical grouping of keys<br>
A KV database is a NoSQL database that stores data as a collection of key-value pairs
Key-value pairs are typically organized into “bucket”
A bucket can roughly be thought of as the KV database equivalent of a table
A bucket is a logical grouping of keys<br>
23
Document Databases (1 of 2) Figure 14.7 Key-Value Database Storage Figure 14.8 Document Database Tagged Format<br>
24
Document Databases (2 of 2) Document databases are conceptually similar to key-value databases
A document database stores data in key-value pairs in which the value component is composed of a tag-encoded document
JSON (JavaScript Object Notation) is a human-readable text format for data interchange that defines attributes and values in a document
BSON (Binary JSON) is a computer-readable format for data interchange that expands the JSON format to include additional data types including binary objects
A collection, in document databases, is a logical storage unit that contains similar documents, roughly analogous to a table in a relational database<br>
A document database stores data in key-value pairs in which the value component is composed of a tag-encoded document
JSON (JavaScript Object Notation) is a human-readable text format for data interchange that defines attributes and values in a document
BSON (Binary JSON) is a computer-readable format for data interchange that expands the JSON format to include additional data types including binary objects
A collection, in document databases, is a logical storage unit that contains similar documents, roughly analogous to a table in a relational database<br>
25
Column-Oriented Databases (1 of 2) Column-oriented databases refers to the following two technologies:
Column-centric storage, which is a storage technique in which data is stored in blocks which hold data from a single column across many rows
Row-centric storage, which is a storage technique in which data is stored in blocks which hold data from all columns of a given set of rows
A column family database is a NoSQL database that organizes data in key-value pairs with keys mapped to a set of columns in the value component
A super column is a groups of columns that are logically related
In a column family database, a collection of columns or super columns related to a collection of rows are grouped together to create a column family<br>
Column-centric storage, which is a storage technique in which data is stored in blocks which hold data from a single column across many rows
Row-centric storage, which is a storage technique in which data is stored in blocks which hold data from all columns of a given set of rows
A column family database is a NoSQL database that organizes data in key-value pairs with keys mapped to a set of columns in the value component
A super column is a groups of columns that are logically related
In a column family database, a collection of columns or super columns related to a collection of rows are grouped together to create a column family<br>
26
Column-Oriented Databases (2 of 2) Figure 14.10 Column Family Database<br>
27
Graph Databases (1 of 2) A graph database is a NoSQL database based on graph theory to store data about relationship-rich environments
The primary components of graph databases are nodes, edges, and properties
The node is a specific instance of something we want to keep data about
An edge is a relationship between nodes
Properties are the attributes or characteristics of a node or edge that are of interest to the users
A query in a graph database is called a traversal<br>
The primary components of graph databases are nodes, edges, and properties
The node is a specific instance of something we want to keep data about
An edge is a relationship between nodes
Properties are the attributes or characteristics of a node or edge that are of interest to the users
A query in a graph database is called a traversal<br>
28
Graph Databases (2 of 2) Figure 14.11 Graph Database Representation<br>
29
Knowledge Check Activity 14-2 What are the four basic categories of NoSQL databases?<br>
30
Knowledge Check Activity 14-2: Answer What are the four basic categories of NoSQL databases?
Answer: Key-value database, document databases, column family databases, and graph databases.<br>
Answer: Key-value database, document databases, column family databases, and graph databases.<br>
31
Aggregate Awareness Key-value, document, and column family databases are aggregate aware
Aggregate aware means that the data is collected or aggregated around a central topic or entity
The aggregate aware database models achieve clustering efficiency by making each piece of data relatively independent
Graph databases, like relational databases, are aggregate ignorant
Aggregate ignorant models do not organize the data into collections based on a central entity<br>
Aggregate aware means that the data is collected or aggregated around a central topic or entity
The aggregate aware database models achieve clustering efficiency by making each piece of data relatively independent
Graph databases, like relational databases, are aggregate ignorant
Aggregate ignorant models do not organize the data into collections based on a central entity<br>
32
NewSQL Databases NewSQL is a database model that attempts to provide ACID-compliant transactions across a highly distributed infrastructure
Characteristics of NewSQL include the following:
Have no proven track record
Have been adopted by relatively few organizations
NewSQL databases support:
SQL as the primary interface
ACID-compliant transactions<br>
Characteristics of NewSQL include the following:
Have no proven track record
Have been adopted by relatively few organizations
NewSQL databases support:
SQL as the primary interface
ACID-compliant transactions<br>
33
Working with Document Databases Using MongoDB (1 of 3) MongoDB is a popular document database
Among the NoSQL databases currently available, MongoDB has been one of the most successful in penetrating the database market
MongoDB, comes from the word humongous as its developers intended their new product to support extremely large data sets
It is designed for the following:
High availability
High scalability
High performance<br>
Among the NoSQL databases currently available, MongoDB has been one of the most successful in penetrating the database market
MongoDB, comes from the word humongous as its developers intended their new product to support extremely large data sets
It is designed for the following:
High availability
High scalability
High performance<br>
34
Working with Document Databases Using MongoDB (2 of 3) Importing Documents in MongoDB
Refer to the text for an importation example and considerations
Example of a MongoDB Query Using find()
Methods are programed functions to manipulate objects
The find() method retrieves objects from a collection that match the restrictions provided
Refer to the text for a query example<br>
Refer to the text for an importation example and considerations
Example of a MongoDB Query Using find()
Methods are programed functions to manipulate objects
The find() method retrieves objects from a collection that match the restrictions provided
Refer to the text for a query example<br>
35
Working with Document Databases Using MongoDB (3 of 3) Figure 14.12 Example of MongoDB Document Query<br>
36
Working with Graph Databases Using Neo4j Even though Neo4j is not yet as widely adopted as MongoDB, it has been one of the fastest growing NoSQL databases
Graph databases still work with concepts similar to entities and relationships
The focus is on the relationships
Graph databases are used in environments with complex relationships among entities
Graph databases are heavily reliant on interdependence among their data
Neo4j provides several interface options
It was originally designed with Java programming in mind and optimized for interaction through a Java API<br>
Graph databases still work with concepts similar to entities and relationships
The focus is on the relationships
Graph databases are used in environments with complex relationships among entities
Graph databases are heavily reliant on interdependence among their data
Neo4j provides several interface options
It was originally designed with Java programming in mind and optimized for interaction through a Java API<br>
37
Creating Nodes in Neo4j Nodes in a graph database correspond to entity instances in a relational database
In Neo4j, a label is the closest thing to the concept of a table from the relational model
A label is a tag that is used to associate a collection of nodes as being of the same type or belonging to the same group
Cypher is the interactive, declarative query language in Neo4j
Nodes and relationships are created using a CREATE command<br>
In Neo4j, a label is the closest thing to the concept of a table from the relational model
A label is a tag that is used to associate a collection of nodes as being of the same type or belonging to the same group
Cypher is the interactive, declarative query language in Neo4j
Nodes and relationships are created using a CREATE command<br>
38
Retrieving Node Data with MATCH and WHERE Refer to the text for examples of the following:
Retrieving node data with MATCH and WHERE
Retrieving relationship data with MATCH and WHERE<br>
Retrieving node data with MATCH and WHERE
Retrieving relationship data with MATCH and WHERE<br>
39
Retrieving Relationship Data with MATCH and WHERE Figure 14.13 Neo4j Query Using MATCH/WHERE/RETURN<br>
40
Knowledge Check Activity 14-3 Explain what it means for a database to be aggregate aware.<br>
41
Knowledge Check Activity 14-3: Answer Explain what it means for a database to be aggregate aware.
Answer: Aggregate aware means that the designer of the database has to be aware of the way the data in the database will be used, and then design the database around whichever component would be central to that usage. Instead of decomposing the data structures to eliminate redundancy, an aggregate aware database is collects, or aggregates, all of the data around a central component to minimize the structures required during processing.<br>
Answer: Aggregate aware means that the designer of the database has to be aware of the way the data in the database will be used, and then design the database around whichever component would be central to that usage. Instead of decomposing the data structures to eliminate redundancy, an aggregate aware database is collects, or aggregates, all of the data around a central component to minimize the structures required during processing.<br>
42
Summary (1 of 2) Now that the lesson has ended, you should be able to:
Explain the role of Big Data in modern business
Describe the primary characteristics of Big Data and how these go beyond the traditional “3 Vs”
Explain how the core components of the Hadoop framework operate
Identify the major components of the Hadoop ecosystem
Summarize the four major approaches of the NoSQL data model and how they differ from the relational model<br>
Explain the role of Big Data in modern business
Describe the primary characteristics of Big Data and how these go beyond the traditional “3 Vs”
Explain how the core components of the Hadoop framework operate
Identify the major components of the Hadoop ecosystem
Summarize the four major approaches of the NoSQL data model and how they differ from the relational model<br>
43
Summary (2 of 2) Now that the lesson has ended, you should be able to (continued):
Describe the characteristics of NewSQL databases
Understand how to work with document databases using MongoDB
Understand how to work with graph databases using Neo4j<br>
Describe the characteristics of NewSQL databases
Understand how to work with document databases using MongoDB
Understand how to work with graph databases using Neo4j<br>