Big Data Technologies syllabus
Browse the units
Big Data Technologies
7 units
3 Distributed Storage and Processing of Big Data (12 LHs)
- HDFS
- Command-line interface
- Hadoop file system interfaces
- Basic HDFS commands. MapReduce
- MapReduce fundamentals
- Execution overview of MapReduce
- Basic MapReduce API concepts
- Setting up the development environment
- Writing unit tests with MRUnit
- Running locally on test data
- Running on a cluster
- Anatomy of a MapReduce job run
- MapReduce types and formats
5 Data Analytics with Spark (6 LHs)
6 Querying Big Data with Pig and Hive (6 LHs)
7 Laboratory Work
- Students will gain hands-on experience through practical exercises that help strengthen…
- Writing and running MapReduce programs
- Performing CRUD operations on MongoDB and HBase
- Analysing data using Spark SQL and Spark Streaming
- Running Pig and Hive queries
- And completing a project that applies multiple Big Data technologies to real-world datasets
1. Introduction to Big Data (4 LHs)
- Introduction to Big Data
- Introduction to Big Data
- Characteristics of Big Data (Volume, Velocity, Variety, Veracity, Value)
- Characteristics of Big Data (Volume, Velocity, Variety, Veracity, Value)
- Current trends and real-life applications of Big Data
- Current trends and real-life applications of Big Data
- Scope and challenges of Big Data
- Scope and challenges of Big Data
- Tools and technologies used in Big Data
- Tools and technologies used in Big Data.
2. Hadoop Ecosystem (10 LHs)
- Introduction to Hadoop
- Introduction to Hadoop
- History of Hadoop
- Hadoop ecosystem overview
- Hadoop ecosystem overview
- Core components of the Hadoop ecosystem
- Core components of the Hadoop ecosystem: Hadoop Common, Hadoop Distributed File System (HDFS), MapReduce Framework, and YARN (Yet Another Resource Negotiator)
- Hadoop master/slave architecture
- Hadoop master/slave architecture
- Hadoop daemons
- Hadoop configuration modes
- Hadoop configuration modes
- Hadoop cluster setup
- Hadoop cluster setup
- Hadoop Streaming.
3. Distributed Storage and Processing of Big Data (12 LHs)
- HDFS
- HDFS: The design of HDFS
- HDFS concepts
- Command-line interface
- command-line interface
- Hadoop file system interfaces
- Hadoop file system interfaces
- data flow
- Basic HDFS commands. MapReduce
- basic HDFS commands. MapReduce: Functional programming
- MapReduce fundamentals
- MapReduce fundamentals
- Execution overview of MapReduce
- execution overview of MapReduce
- Basic MapReduce API concepts
- basic MapReduce API concepts
- Setting up the development environment
- setting up the development environment
- Writing unit tests with MRUnit
- writing unit tests with MRUnit
- Running locally on test data
- running locally on test data
- Running on a cluster
- running on a cluster
- Anatomy of a MapReduce job run
- anatomy of a MapReduce job run
- failures
- shuffle and sort
- task execution
- MapReduce types and formats
- MapReduce types and formats.
4. NoSQL Databases (10 LHs)
- Types of data
- Types of data
- Introduction to NoSQL
- Introduction to NoSQL
- Need for NoSQL
- Types of NoSQL databases
- Types of NoSQL databases
- NoSQL vs. relational databases
- NoSQL vs. relational databases
- The CAP theorem. MongoDB
- The CAP theorem. MongoDB: Collections, documents, object IDs, queries on MongoDB, aggregation pipeline, nested documents. HBase: Overview, HBase vs. RDBMS, HBase vs. HDFS, HBase architecture, HBase data model, concept of row keys and column families, HBase regions, creating a table, writing queries to insert and retrieve data to and from HBase.
5. Data Analytics with Spark (6 LHs)
- Introduction to Spark
- Introduction to Spark
- Need for Spark
- Evolution of Spark
- Spark shell
- Spark context
- Resilient Distributed Dataset (RDD)
- Resilient Distributed Dataset (RDD)
- Transformations
- Programming with RDD
- Programming with RDD
- Spark Core
- Spark SQL
- MLlib
- Spark Streaming
- GraphX.
6. Querying Big Data with Pig and Hive (6 LHs)
- Pig
- Pig: Introduction to Pig
- Execution modes of Pig
- execution modes of Pig
- Comparison of Pig with databases
- comparison of Pig with databases
- Grunt
- Pig Latin
- User-defined functions
- user-defined functions
- user-defined functions.
- Data processing operators. Hive
- data processing operators. Hive: Hive shell
- Hive services
- Hive Metastore
- Comparison with traditional databases
- comparison with traditional databases
- HiveQL
- tables
- querying data
7. Laboratory Work
- Students will gain hands-on experience through practical exercises that help strengthen…
- Students will gain hands-on experience through practical exercises that help strengthen their understanding of Big Data technologies. This includes setting up Hadoop clusters and performing HDFS operations
- Writing and running MapReduce programs
- writing and running MapReduce programs
- Performing CRUD operations on MongoDB and HBase
- performing CRUD operations on MongoDB and HBase
- Analysing data using Spark SQL and Spark Streaming
- analysing data using Spark SQL and Spark Streaming
- Running Pig and Hive queries
- running Pig and Hive queries
- And completing a project that applies multiple Big Data technologies to real-world datasets
- and completing a project that applies multiple Big Data technologies to real-world datasets. These exercises aim to provide practical experience in implementing Big Data solutions across multiple platforms.