Foundational Math for Generative AI: Understanding LLMs and Transformers through Practical Applications

Foundational Math for Generative AI: Understanding LLMs and Transformers through Practical Applications

2h 59mIntermediate2025-02-03

Authors

Axel Sirota

Axel Sirota

Course details

Unlock the mysteries behind the models powering today’s most advanced AI applications. In this course, instructor Axel Sirota takes you beyond just using large language models (LLMs) like BERT or GPT and highlights the mathematical foundations of generative AI. Explore the challenge of sentiment analysis with simple recurrent neural networks (RNNs) and progressively evolve your approach as you gain a deep understanding of attention mechanisms, transformers, and models. Through intuitive explanations and hands-on coding exercises, Axel outlines why attention revolutionized natural language processing, and how transformers reshaped the field by eliminating the need for RNNs altogether. Along the way, get tips on fine-tuning pretrained models, applying cutting-edge techniques like low-rank adaptation (LoRA), and leveraging your newly acquired skills to build smarter, more efficient models and innovate in the fast-evolving world of AI.

Learning objectives
Gain an intuitive understanding of how and why LLMs and transformers work.
Learn how attention mechanisms evolved to solve key problems in RNN-based models.
Develop a sentiment analysis model using TensorFlow, Keras, and Hugging Face’s DistilBERT.
Enhance models progressively with mathematical insights applied in code, from word embeddings to attention and transformer layers.
Use visualizations to grasp how attention and optimization work.

Skills covered

Natural Language Processing (NLP)Generative AIData AnalysisArtificial Intelligence (AI)Data ScienceBusiness Analysis and StrategyBusiness Software and ToolsOne-Off

Concepts

Introduction

  • Intro to foundational math for generative AI
  • Getting the most out of this course
  • Version check

Introduction to Math for GenAI and Attention Basics

  • Why LLMs and attention matter
  • RNNs and the context bottleneck problem
  • Demo - Building a simple RNN model for sentiment analysis
  • Introduction to attention - Bahdanau s solution
  • Demo - Adding attention to an RNN model
  • Solution - Implement Bahdanau's attention

Transformers - Removing RNNs for More Efficient Models

  • From RNNs to transformers
  • Understanding self-attention in transformers
  • Multi-head attention and positional encoding
  • Building a transformer model for sentiment analysis
  • Solution - Build a two-layer transformer encoder

Deep Dive into LLMs and Model Fine-Tuning

  • The three types of LLMs
  • Special decoder-only models
  • Explaining encoder-only models like BERT
  • Fine-tuning DistilBERT for sentiment analysis
  • Attention masks in transformers
  • Solution - Detect irony and climate stance in TweetEval

Conclusion

  • Course summary and next steps
80,000 Toman