Multimodal AI Essentials: Merging Text, Image, and Audio for Next-Generation AI Applications
5h 33mIntermediate2025-07-02
Authors

Pearson

Sinan Ozdemir
Course details
This course shows you how combining modalities like text, audio, video, and images can enable AI systems to achieve remarkable capabilities. Gain hands-on experience building visual question-and-answer models, generating personalized images with diffusion, designing end to end multimodal applications, and even fine-tuning multimodal models for specific tasks. This course gives you the tools, knowledge, and confidence to design and deploy your own state-of-the-art multimodal AI systems.
Learning objectives
Apply multimodal AI concepts.
Build a voice-to-voice app.
Apply visual question answering (VQA) concepts and architecture.
Construct, fine-tune, and evaluate diffusion models with DreamBooth.
Fine-tune a text-to-speech model with SpeechT5.
Build visual agents from the ground up.
Evaluate the performance of multimodal models.
Extend multimodal systems with advanced techniques like computer use.
Learning objectives
Apply multimodal AI concepts.
Build a voice-to-voice app.
Apply visual question answering (VQA) concepts and architecture.
Construct, fine-tune, and evaluate diffusion models with DreamBooth.
Fine-tune a text-to-speech model with SpeechT5.
Build visual agents from the ground up.
Evaluate the performance of multimodal models.
Extend multimodal systems with advanced techniques like computer use.
Skills covered
Neural Networks and Deep LearningAI Productivity ToolsArtificial Intelligence FoundationsArtificial Intelligence for BusinessArtificial Intelligence (AI)Business Software and ToolsOne-Off
Concepts
0. Introduction
- 01 - Multimodal AI essentials - Introduction
1. Introduction to Multimodal AI
- 02 - Topics
- 03 - Overview of multimodal AI concepts
- 04 - Types of data in multimodal systems
- 05 - Building a voice-to-voice app
2. Building Visual Question Answering (VQA) Models
- 06 - Topics
- 07 - Understanding VQA - Concepts and architecture
- 08 - Fusing modalities to perform VQA, part 1
- 09 - Fusing modalities to perform VQA, part 2
- 10 - Fusing modalities to perform VQA, part 3
- 11 - Blending modalities to perform VQA, part 1
- 12 - Blending modalities to perform VQA, part 2
3. Exploring Diffusion Models
- 13 - Topics
- 14 - Introduction to diffusion models
- 15 - Hands-on - Implementing diffusion models with DreamBooth
4. Developing Multimodal AI Systems
- 16 - Topics
- 17 - Designing multimodal AI systems
- 18 - Fine-tuning a text-to-speech model with T5
- 19 - Building visual agents
5. Evaluating and Testing Multimodal AI Systems
- 20 - Topics
- 21 - Evaluating multimodal models - Accuracy and performance
- 22 - Bias and ethics in multimodality
6. Expanding and Applying Multimodal AI
- 23 - Topics
- 24 - Extending multimodal systems with advanced techniques
- 25 - Future trends and innovations in multimodal AI
Conclusion
- 26 - Multimodal AI essentials - Summary