Skip to content

Big Data Technologies syllabus

MIT 5517 units · 44 topicsAcademic year 2083/84
Browse the units

Big Data Technologies

7 units

1. Introduction to Big Data (4 LHs)

  1. Introduction to Big Data
    1. Introduction to Big Data
  2. Characteristics of Big Data (Volume, Velocity, Variety, Veracity, Value)
    1. Characteristics of Big Data (Volume, Velocity, Variety, Veracity, Value)
  3. Current trends and real-life applications of Big Data
    1. Current trends and real-life applications of Big Data
  4. Scope and challenges of Big Data
    1. Scope and challenges of Big Data
  5. Tools and technologies used in Big Data
    1. Tools and technologies used in Big Data.

2. Hadoop Ecosystem (10 LHs)

  1. Introduction to Hadoop
    1. Introduction to Hadoop
    2. History of Hadoop
  2. Hadoop ecosystem overview
    1. Hadoop ecosystem overview
  3. Core components of the Hadoop ecosystem
    1. Core components of the Hadoop ecosystem: Hadoop Common, Hadoop Distributed File System (HDFS), MapReduce Framework, and YARN (Yet Another Resource Negotiator)
  4. Hadoop master/slave architecture
    1. Hadoop master/slave architecture
    2. Hadoop daemons
  5. Hadoop configuration modes
    1. Hadoop configuration modes
  6. Hadoop cluster setup
    1. Hadoop cluster setup
    2. Hadoop Streaming.

3. Distributed Storage and Processing of Big Data (12 LHs)

  1. HDFS
    1. HDFS: The design of HDFS
    2. HDFS concepts
  2. Command-line interface
    1. command-line interface
  3. Hadoop file system interfaces
    1. Hadoop file system interfaces
    2. data flow
  4. Basic HDFS commands. MapReduce
    1. basic HDFS commands. MapReduce: Functional programming
  5. MapReduce fundamentals
    1. MapReduce fundamentals
  6. Execution overview of MapReduce
    1. execution overview of MapReduce
  7. Basic MapReduce API concepts
    1. basic MapReduce API concepts
  8. Setting up the development environment
    1. setting up the development environment
  9. Writing unit tests with MRUnit
    1. writing unit tests with MRUnit
  10. Running locally on test data
    1. running locally on test data
  11. Running on a cluster
    1. running on a cluster
  12. Anatomy of a MapReduce job run
    1. anatomy of a MapReduce job run
    2. failures
    3. shuffle and sort
    4. task execution
  13. MapReduce types and formats
    1. MapReduce types and formats.

4. NoSQL Databases (10 LHs)

  1. Types of data
    1. Types of data
  2. Introduction to NoSQL
    1. Introduction to NoSQL
    2. Need for NoSQL
  3. Types of NoSQL databases
    1. Types of NoSQL databases
  4. NoSQL vs. relational databases
    1. NoSQL vs. relational databases
  5. The CAP theorem. MongoDB
    1. The CAP theorem. MongoDB: Collections, documents, object IDs, queries on MongoDB, aggregation pipeline, nested documents. HBase: Overview, HBase vs. RDBMS, HBase vs. HDFS, HBase architecture, HBase data model, concept of row keys and column families, HBase regions, creating a table, writing queries to insert and retrieve data to and from HBase.

5. Data Analytics with Spark (6 LHs)

  1. Introduction to Spark
    1. Introduction to Spark
    2. Need for Spark
    3. Evolution of Spark
    4. Spark shell
    5. Spark context
  2. Resilient Distributed Dataset (RDD)
    1. Resilient Distributed Dataset (RDD)
    2. Transformations
  3. Programming with RDD
    1. Programming with RDD
    2. Spark Core
    3. Spark SQL
    4. MLlib
    5. Spark Streaming
    6. GraphX.

6. Querying Big Data with Pig and Hive (6 LHs)

  1. Pig
    1. Pig: Introduction to Pig
  2. Execution modes of Pig
    1. execution modes of Pig
  3. Comparison of Pig with databases
    1. comparison of Pig with databases
    2. Grunt
    3. Pig Latin
  4. User-defined functions
    1. user-defined functions
    2. user-defined functions.
  5. Data processing operators. Hive
    1. data processing operators. Hive: Hive shell
    2. Hive services
    3. Hive Metastore
  6. Comparison with traditional databases
    1. comparison with traditional databases
    2. HiveQL
    3. tables
    4. querying data

7. Laboratory Work

  1. Students will gain hands-on experience through practical exercises that help strengthen…
    1. Students will gain hands-on experience through practical exercises that help strengthen their understanding of Big Data technologies. This includes setting up Hadoop clusters and performing HDFS operations
  2. Writing and running MapReduce programs
    1. writing and running MapReduce programs
  3. Performing CRUD operations on MongoDB and HBase
    1. performing CRUD operations on MongoDB and HBase
  4. Analysing data using Spark SQL and Spark Streaming
    1. analysing data using Spark SQL and Spark Streaming
  5. Running Pig and Hive queries
    1. running Pig and Hive queries
  6. And completing a project that applies multiple Big Data technologies to real-world datasets
    1. and completing a project that applies multiple Big Data technologies to real-world datasets. These exercises aim to provide practical experience in implementing Big Data solutions across multiple platforms.