Information Retrieval and Web Search syllabus
Browse the units
Information Retrieval and Web Search
6 units
1 Classical IR Foundations (12 LHs)
- Introduction and Text Processing
- Information Need VS Queries
- Stemming (Porter Stemmer)
- Stop Word Removal Inverted Indexing
- Zipf’s Law IR Models
- Inverse Document Frequency (IDF)
- Vector Space Model Probabilistic Model
- Probability Ranking Principle (PRP)
- BM25 Ranking Function
- Probabilistic VS likelihood Models Learning Outcome
2 Neural Information Retrieval (20 LHs)
- Foundation
- Large Language Models for IR
- BERT as Encoder Only Model Text Tokenization
- Unigrams Text Representations for Ranking
- Word Embedding (Word2Vec)
- Contextual Embedding (BERT Embedding)
- Sentence Embedding (Sentence BERT Embedding)
- Re - ranking Retrieval
- Embedding Models VS Re – ranker Models Interaction Focused Systems
- Fine – tuning interaction focus systems
- Dealing with Long Texts RAG (Retrieval Augmented Generation)
- Design of RAG Systems
- Indexing Pipeline (Data Loading, Data Chunking, Data Embedding, Vector Database)
- Generation Pipeline (Retrieval, Augmented, Generation)
- RAG Evaluations (Retrieval Metrics, RAG Specific Metrics)
- Hallucination as a Challenge Learning Outcome
- Real-time data sources?
5 Web Search (6 LHs)
- Search engines (working principle)
- Spidering (Structure of a spider, Simple spidering algorithm, multithreaded spidering, Bot)
- Directed spidering (Topic directed, Link directed)
- Crawlers (Basic crawler architecture)
- Link analysis (HITS, Page ranking)
- Handling “invisible” Web – Snippet generation Learning Outcome
- Lectures with demonstration
- Hands-on lab sessions
- Problem-based learning
- Guest lectures from tech industry experts
- Continuous assessment and feedback
- Multimedia presentations to visualize concepts
1. Classical IR Foundations (12 LHs)
- Introduction and Text Processing
- Introduction and Text Processing: IR Architecture
- Information Need VS Queries
- Information Need VS Queries
- Tokenization
- Lemmatization
- Stemming (Porter Stemmer)
- stemming (Porter Stemmer)
- Stop Word Removal Inverted Indexing
- Stop Word Removal Inverted Indexing: Inverted Index Construction
- Skip List
- Zipf’s Law IR Models
- Zipf’s Law IR Models: Boolean Retrieval Model
- Term Frequency (TF)
- Inverse Document Frequency (IDF)
- Inverse Document Frequency (IDF)
- TF – IDF
- Cosine Similarity
- Vector Space Model Probabilistic Model
- Vector Space Model Probabilistic Model: The Binary Independence Model (BIM)
- Probability Ranking Principle (PRP)
- Probability Ranking Principle (PRP)
- BM25 Ranking Function
- BM25 Ranking Function
- Probabilistic VS likelihood Models Learning Outcome
- Probabilistic VS likelihood Models Learning Outcome: Understand the basic underling foundation of information retrieval engine. Develop an idea how classical IR models are implemented for information retrieval.
2. Neural Information Retrieval (20 LHs)
- Foundation
- Foundation: General architecture of Transformer
- Large Language Models for IR
- Large Language Models for IR
- BERT as Encoder Only Model Text Tokenization
- BERT as Encoder Only Model Text Tokenization: BPE (Byte Pair Encoding)
- Word Piece
- Unigrams Text Representations for Ranking
- Unigrams Text Representations for Ranking: BOW Encodings
- Word Embedding (Word2Vec)
- Word Embedding (Word2Vec)
- Contextual Embedding (BERT Embedding)
- Contextual Embedding (BERT Embedding)
- Sentence Embedding (Sentence BERT Embedding)
- Sentence Embedding (Sentence BERT Embedding)
- Position Embedding
- Attention
- Re - ranking Retrieval
- Re - ranking Retrieval
- Embedding Models VS Re – ranker Models Interaction Focused Systems
- Embedding Models VS Re – ranker Models Interaction Focused Systems: Pre – trained language models
- Fine – tuning interaction focus systems
- Fine – tuning interaction focus systems
- Prompt Optimization
- Dealing with Long Texts RAG (Retrieval Augmented Generation)
- Dealing with Long Texts RAG (Retrieval Augmented Generation): Novelty of RAG
- Design of RAG Systems
- Design of RAG Systems
- Indexing Pipeline (Data Loading, Data Chunking, Data Embedding, Vector Database)
- Indexing Pipeline (Data Loading, Data Chunking, Data Embedding, Vector Database)
- Generation Pipeline (Retrieval, Augmented, Generation)
- Generation Pipeline (Retrieval, Augmented, Generation)
- RAG Evaluations (Retrieval Metrics, RAG Specific Metrics)
- RAG Evaluations (Retrieval Metrics, RAG Specific Metrics)
- Hallucination as a Challenge Learning Outcome
- Hallucination as a Challenge Learning Outcome: Describe the modern information retrieval model as large language models. How transformer based model can be used in text representation and ranking the documents? Learn the need of semantics over syntactic aspects. Why do we need RAG to improve the quality of Large Language Model (LLM) responses by grounding the model on external
- Real-time data sources?
- real-time data sources?
3. Evaluation Information Retrieval (5 LHs)
- Unranked Metrics (Precision, Recall)
- Unranked Metrics (Precision, Recall)
- Average Precision
- Discounted Cumulated Gain
- Discounted Cumulated Gain
- Ranked Metrics (Mean Reciprocal Rank, Normalized Discounted Cumulative Gain)
- Ranked Metrics (Mean Reciprocal Rank, Normalized Discounted Cumulative Gain)
- Faithfulness
- Relevance
- ANOVA (Analysis of Variance)
- ANOVA (Analysis of Variance)
- A/B Testing
- Interleaving Learning Outcome
- Interleaving Learning Outcome: Explain the evaluation methods to express the accuracy.
4. Adaptations and Concerns (5 LHs)
- CLIR (Cross Language Information Retrieval)
- CLIR (Cross Language Information Retrieval): Introduction
- Some use cases
- The Core Technology of CLIR Bias in Retrieval Systems
- The Core Technology of CLIR Bias in Retrieval Systems: Pre – existing bias
- Stakeholder bias
- Data bias
- Algorithmic bias Privacy in IR
- Algorithmic bias Privacy in IR: Privacy for Searchers
- Privacy for Search Engines
- Privacy for Search Engines
- Privacy for Document Owners Learning Outcome
- Privacy for Document Owners Learning Outcome: Why do we need multilingual machine translation model? Describe the different privacy issues related to information retrieval
5. Web Search (6 LHs)
- Search engines (working principle)
- Search engines (working principle)
- Spidering (Structure of a spider, Simple spidering algorithm, multithreaded spidering, Bot)
- Spidering (Structure of a spider, Simple spidering algorithm, multithreaded spidering, Bot)
- Directed spidering (Topic directed, Link directed)
- Directed spidering (Topic directed, Link directed)
- Crawlers (Basic crawler architecture)
- Crawlers (Basic crawler architecture)
- Link analysis (HITS, Page ranking)
- Link analysis (HITS, Page ranking)
- Query log analysis
- Handling “invisible” Web – Snippet generation Learning Outcome
- Handling “invisible” Web – Snippet generation Learning Outcome: Describe the working mechanism of search engine spider. Apply the Link analysis approach for ranking the pages. What are the advantages of having Query Logs? Pedagogical Strategies
- Lectures with demonstration
- Lectures with demonstration
- Hands-on lab sessions
- Hands-on lab sessions
- Problem-based learning
- Problem-based learning
- Guest lectures from tech industry experts
- Guest lectures from tech industry experts
- Continuous assessment and feedback
- Continuous assessment and feedback
- Multimedia presentations to visualize concepts
- Multimedia presentations to visualize concepts
- Mini project
6. Laboratory Works
- The laboratory work includes writing programs for implementing the concepts of text…
- The laboratory work includes writing programs for implementing the concepts of text normalization pipeline and construct an Inverted Index from raw text documents
- To implement classic ranking and scoring algorithms to return documents sorted by relevance
- to implement classic ranking and scoring algorithms to return documents sorted by relevance
- To build a multithreaded web crawler and implementing semantic machine learning models…
- to build a multithreaded web crawler and implementing semantic machine learning models and generative systems using a programming language like Python.