Hands-On Introduction to Transformers for Computer Vision
3h 46mIntermediate2025-07-09
Authors

Daniel Gural
Course details
Explore the groundbreaking features of vision transformers in this robust machine learning course. Begin with the foundations of transformers and their revolutionary applications in computer vision. Learn how to implement vision transformers, uncover techniques for training and fine-tuning state-of-the-art models, and assess them using rigorous evaluation metrics. Efficiently deploy models using quantization to support edge environments and uncover insights. Observe firsthand the transition from theory to practice through real-world dataset applications. Gain insights into ensuring model safety and reliability through evaluation and debugging techniques. Whether you're an AI researcher, a data engineer, or a developer, this course equips you with the tools to excel in the rapidly evolving landscape of machine learning.
Learning objectives
Understand why and how transformers are so popular in CV.
Implement transformers in your own projects.
Fine-tune to improve model performance.
Add explainability for responsible AI models.
Take a deep dive in Vision Transformers.
Learning objectives
Understand why and how transformers are so popular in CV.
Implement transformers in your own projects.
Fine-tune to improve model performance.
Add explainability for responsible AI models.
Take a deep dive in Vision Transformers.
Skills covered
Model Training and EvaluationPyTorchTraditional AI and Machine LearningArtificial Intelligence (AI)Open SourceOne-Off
Concepts
Introduction
- Upgrade your transformer and computer vision knowledge
Course Setup
- Getting comfortable with Codespaces
What Is a Transformer
- So, what is a transformer
- A brief history - Life before transformers
- Why transformers are revolutionary for computer vision
- Comparing convolutional neural networks (CNNs) to transformers
Getting Started with Vision Transformers
- Overview of PyTorch and Hugging Face transformers
- Building a dataset
- Running a pretrained model for classification
Implementing Vision Transformers from Scratch
- Implementing from scratch - Attention is all you need
- Tokenization of images - How vision transformers (ViTs) see
- Building a simple ViT
Fine-Tuning Transformers with Transfer Learning
- Transformers in the wild - How to find the right one for you
- Fine-tuning in practice - Adapting ViTs to new tasks
- Training strategies for vision transformers
- Fine-tuning Swin Transformers
Real-World Use Cases
- Choosing the right ViT for the job
- Evaluating ViTs
- Optimizing inference at deployment
Explainability and Attention Visualization
- What is an attention heatmap
- Visualizing attention
Conclusion
- Next steps