Complete Guide to Evaluating Large Language Models (LLMs)
7h 57mIntermediate2025-07-02
Authors

Pearson

Sinan Ozdemir
Course details
In this comprehensive course, AI and LLM expert Sinan Ozdemir shares with you the knowledge and skills to assess LLM performance effectively. Get a detailed introduction to the process of evaluating LLMs, Multimodal AI, and AI-powered applications like agents and RAG. Learn how to thoroughly assess and evaluate these powerful and often unwieldy AI tools so you can make sure they meet your real-world needs. This course prepares you to evaluate and optimize LLMs so you can produce cutting edge AI applications.
Learning objectives
Distinguish between generative and understanding tasks.
Apply key metrics for common tasks.
Evaluate multiple-choice tasks.
Evaluate free text response tasks.
Evaluate embedding tasks.
Evaluate classification tasks.
Build an LLM classifier with BERT and ChatGPT.
Evaluate LLMs with benchmarks.
Probe LLMs.
Fine-tune LLMs.
Evaluate and clean data.
Evaluate AI agents.
Evaluate retrieval-augmented generation (RAG) systems.
Evaluate a recommendation engine.
Use evaluation to combat AI drift.
Learning objectives
Distinguish between generative and understanding tasks.
Apply key metrics for common tasks.
Evaluate multiple-choice tasks.
Evaluate free text response tasks.
Evaluate embedding tasks.
Evaluate classification tasks.
Build an LLM classifier with BERT and ChatGPT.
Evaluate LLMs with benchmarks.
Probe LLMs.
Fine-tune LLMs.
Evaluate and clean data.
Evaluate AI agents.
Evaluate retrieval-augmented generation (RAG) systems.
Evaluate a recommendation engine.
Use evaluation to combat AI drift.
Skills covered
Natural Language Processing (NLP)Generative AIArtificial Intelligence FoundationsArtificial Intelligence (AI)One-Off
Concepts
0. Introduction
- 01 - Evaluating LLMs - Introduction
1. Foundations of LLM Evaluation
- 02 - Topics
- 03 - Introduction to evaluation - Why it matters
- 04 - Generative versus understanding tasks
- 05 - Key metrics for common tasks
2. Evaluating Generative Tasks
- 06 - Topics
- 07 - Evaluating multiple-choice tasks
- 08 - Evaluating free text response tasks, part 1
- 09 - Evaluating free text response tasks, part 2
- 10 - AIs supervising AIs - LLM as a judge
3. Evaluating Understanding Tasks
- 11 - Topics
- 12 - Evaluating embedding tasks
- 13 - Evaluating classification tasks
- 14 - Building an LLM classifier with BERT and GPT
4. Using Benchmarks Effectively
- 15 - Topics
- 16 - The role of benchmarks
- 17 - Interrogating common benchmarks
- 18 - Evaluating LLMs with benchmarks
5. Probing LLMs for a World Model
- 19 - Topics
- 20 - Probing LLMs for knowledge
- 21 - Probing LLMs to play games
6. Evaluating LLM Fine-Tuning
- 22 - Topics
- 23 - Fine-tuning objectives
- 24 - Metrics for fine-tuning success
- 25 - Practical demonstration - Evaluating fine-tuning
- 26 - Evaluating and cleaning data
7. Case Studies
- 27 - Topics
- 28 - Evaluating AI agents - Task automation and tool integration
- 29 - Measuring retrieval-augmented generation (RAG) systems
- 30 - Building and evaluating a recommendation engine using LLMs
- 31 - Using evaluation to combat AI drift
- 32 - Time-series regression
8. Summary of Evaluation and Looking Ahead
- 33 - Topics
- 34 - When and how to evaluate
- 35 - Looking ahead - Trends in LLM evaluation
Conclusion
- 36 - Evaluating LLMs - Summary