Azure AI for Developers: Azure AI Speech
55mIntermediate2025-04-15
Authors

Marco Casalaina
Course details
Using pre-built or customizable speech models, Azure AI Speech allows developers to build multimodal, multilingual, voice-enabled AI apps. In this course, instructor Marco Casalaina begins by outlining the basic features and capabilities of Azure Speech and identifies the most common use cases. Then, through hands-on instruction, he covers speech to text models and transcriptions, text to speech tools and voices, and avatar creation. The course wraps up with coverage of advanced Azure Speech capabilities.
Learning objectives
Identify common use cases for Azure AI Speech.
Customize speech to text models to fit specific needs.
Build and test text to speech audio content.
Build custom avatars and integrate gestures for enhanced communication.
Learning objectives
Identify common use cases for Azure AI Speech.
Customize speech to text models to fit specific needs.
Build and test text to speech audio content.
Build custom avatars and integrate gestures for enhanced communication.
Skills covered
Azure AI ServicesAI Development Tools and PlatformsProgramming FoundationsCloud AdministrationBuilding with AICloud PlatformsCloud ComputingMicrosoftSoftware DevelopmentDeep Dive (X:Y)
Concepts
Introduction
- What this course is about
- What you should know
Azure Speech in Action - Common Use Cases
- Common scenarios for Azure AI Speech
Speech to Text and Transcription
- How speech to text works
- Transcription
- Customizing speech to text
- Choosing between the OpenAI Whisper and Azure Speech models
- Speech translation
Text to Speech
- Text to speech - Azure Voice Gallery
- Audio content creation
- Custom voices
Avatars
- Combining speech with avatars
- Building custom avatars
- Live chat avatars
Advanced Speech Capabilities
- Video translation
- Pronunciation assessment
- Using Azure Content Understanding for audio and video
- Azure Speech vs. real-time LLMs
Conclusion
- More resources on Azure Speech