Hands-On Introduction to Transformers for Computer Vision

Hands-On Introduction to Transformers for Computer Vision

3h 46mIntermediate2025-07-09

Authors

Daniel Gural

Daniel Gural

Course details

Explore the groundbreaking features of vision transformers in this robust machine learning course. Begin with the foundations of transformers and their revolutionary applications in computer vision. Learn how to implement vision transformers, uncover techniques for training and fine-tuning state-of-the-art models, and assess them using rigorous evaluation metrics. Efficiently deploy models using quantization to support edge environments and uncover insights. Observe firsthand the transition from theory to practice through real-world dataset applications. Gain insights into ensuring model safety and reliability through evaluation and debugging techniques. Whether you're an AI researcher, a data engineer, or a developer, this course equips you with the tools to excel in the rapidly evolving landscape of machine learning.

Learning objectives
Understand why and how transformers are so popular in CV.
Implement transformers in your own projects.
Fine-tune to improve model performance.
Add explainability for responsible AI models.
Take a deep dive in Vision Transformers.

Skills covered

Model Training and EvaluationPyTorchTraditional AI and Machine LearningArtificial Intelligence (AI)Open SourceOne-Off

Concepts

Introduction

  • Upgrade your transformer and computer vision knowledge

Course Setup

  • Getting comfortable with Codespaces

What Is a Transformer

  • So, what is a transformer
  • A brief history - Life before transformers
  • Why transformers are revolutionary for computer vision
  • Comparing convolutional neural networks (CNNs) to transformers

Getting Started with Vision Transformers

  • Overview of PyTorch and Hugging Face transformers
  • Building a dataset
  • Running a pretrained model for classification

Implementing Vision Transformers from Scratch

  • Implementing from scratch - Attention is all you need
  • Tokenization of images - How vision transformers (ViTs) see
  • Building a simple ViT

Fine-Tuning Transformers with Transfer Learning

  • Transformers in the wild - How to find the right one for you
  • Fine-tuning in practice - Adapting ViTs to new tasks
  • Training strategies for vision transformers
  • Fine-tuning Swin Transformers

Real-World Use Cases

  • Choosing the right ViT for the job
  • Evaluating ViTs
  • Optimizing inference at deployment

Explainability and Attention Visualization

  • What is an attention heatmap
  • Visualizing attention

Conclusion

  • Next steps
80,000 Toman