Hands-On Introduction to Transformers for Computer Vision
3h 46mIntermediate2025-07-09
Authors

Daniel Gural
Course details
Explore the groundbreaking features of vision transformers in this robust machine learning course. Begin with the foundations of transformers and their revolutionary applications in computer vision. Learn how to implement vision transformers, uncover techniques for training and fine-tuning state-of-the-art models, and assess them using rigorous evaluation metrics. Efficiently deploy models using quantization to support edge environments and uncover insights. Observe firsthand the transition from theory to practice through real-world dataset applications. Gain insights into ensuring model safety and reliability through evaluation and debugging techniques. Whether you're an AI researcher, a data engineer, or a developer, this course equips you with the tools to excel in the rapidly evolving landscape of machine learning.
Learning objectives
Understand why and how transformers are so popular in CV.
Implement transformers in your own projects.
Fine-tune to improve model performance.
Add explainability for responsible AI models.
Take a deep dive in Vision Transformers.
Learning objectives
Understand why and how transformers are so popular in CV.
Implement transformers in your own projects.
Fine-tune to improve model performance.
Add explainability for responsible AI models.
Take a deep dive in Vision Transformers.
Skills covered
PyTorchNeural Networks and Deep LearningArtificial Intelligence (AI)Open SourceOne-Off
Concepts
0. Introduction
- 01 - Upgrade your transformer and computer vision knowledge
1. Course Setup
- 02 - Getting comfortable with Codespaces
2. What Is a Transformer
- 03 - So, what is a transformer
- 04 - A brief history - Life before transformers
- 05 - Why transformers are revolutionary for computer vision
- 06 - Comparing convolutional neural networks (CNNs) to transformers
3. Getting Started with Vision Transformers
- 07 - Overview of PyTorch and Hugging Face transformers
- 08 - Building a dataset
- 09 - Running a pretrained model for classification
4. Implementing Vision Transformers from Scratch
- 10 - Implementing from scratch - Attention is all you need
- 11 - Tokenization of images - How vision transformers (ViTs) see
- 12 - Building a simple ViT
5. Fine-Tuning Transformers with Transfer Learning
- 13 - Transformers in the wild - How to find the right one for you
- 14 - Fine-tuning in practice - Adapting ViTs to new tasks
- 15 - Training strategies for vision transformers
- 16 - Fine-tuning Swin Transformers
6. Real-World Use Cases
- 17 - Choosing the right ViT for the job
- 18 - Evaluating ViTs
- 19 - Optimizing inference at deployment
7. Explainability and Attention Visualization
- 20 - What is an attention heatmap
- 21 - Visualizing attention
Conclusion
- 22 - Next steps