Microsoft Azure Data Engineer Associate (DP-203) Cert Prep: 2 Design and Develop Data Processing
3h 28mIntermediate2023-08-14
Authors

Microsoft Learn
Build Skills That Open Doors

Tim Warner
Technical Trainer and Content Developer
Course details
The latest professional certifications from Azure are aligned with specific industry roles. Earning your Azure certification helps to validate your unique Azure skill set and increase your value in today's IT job market. Join Microsoft MVP and Microsoft Certified Azure Solutions Architect Tim Warner for an overview of the core concepts and skills required to pass the Microsoft Azure Data Engineer Associate (DP-203) certification exam. In this course, the second in a four-part certification prep series, explore the fundamentals of how to design and develop data processing solutions. Learn how to ingest and transform data with tools like Apache Spark, Transact SQL, Data Factory, and Stream Analytics. Gather insights for working with transformed data, troubleshooting transformations, and designing, developing, configuring, and troubleshooting solutions for both batch and stream processing. Along the way, find out how to manage batches and pipelines for consistent and successful delivery.
Skills covered
Data EngineeringAzureNetwork AdministrationCloud PlatformsCert PrepNetwork and System AdministrationCloud ComputingData ScienceMicrosoft
Concepts
Ingest and Transform Data
- Learning objectives
- Transform data by using Apache Spark
- Transform data by using Transact-SQL
- Transform data by using Data Factory
- Transform data by using Azure Synapse pipelines
- Transform data by using Stream Analytics
Work with Transformed Data
- Learning objectives
- Cleanse data
- Split data
- Shred JSON
- Encode and decode data
Troubleshoot Data Transformations
- Learning objectives
- Configure error handling for the transformation
- Normalize and denormalize values
- Transform data by using Scala
- Perform data exploratory analysis
Design a Batch Processing Solution
- Learning objectives
- Develop batch processing solutions by using Data Factory, Data Lake, Spark, Azure Synapse pipelines, PolyBase, and Azure Databricks
- Create data pipelines
- Design and implement incremental data loads
- Design and develop slowly changing dimensions
- Handle security and compliance requirements
- Scale resources
Develop a Batch Processing Solution
- Learning objectives
- Configure the batch size
- Design and create tests for data pipelines
- Integrate Jupyter and Python Notebooks into a data pipeline
- Handle duplicate data
- Handle missing data
- Handle late-arriving data
Configure a Batch Processing Solution
- Learning objectives
- Upsert data
- Regress to a previous state
- Design and configure exception handling
- Configure batch retention
- Revisit batch processing solution design
- Debug Spark jobs by using the Spark UI
Design a Stream Processing Solution
- Learning objective
- Develop a stream processing solution by using Stream Analytics, Azure Databricks, and Azure Event Hubs
- Process data by using Spark structured streaming
- Monitor for performance and functional regressions
- Design and create windowed aggregates
- Handle schema drift
Process Data in a Stream Processing Solution
- Learning objectives
- Process time series data
- Process across partitions
- Process within one partition
- Configure checkpoints and watermarking during processing
- Scale resources
- Design and create tests for data pipelines
- Optimize pipelines for analytical or transactional purposes
Troubleshoot a Stream Processing Solution
- Learning objectives
- Handle interruptions
- Design and configure exception handling
- Upsert data
- Replay archived stream data
- Design a stream processing solution
Manage Batches and Pipelines
- Learning objectives
- Trigger batches
- Handle failed batch loads
- Validate batch loads
- Manage data pipelines in Data Factory and Synapse pipelines
- Schedule data pipelines in Data Factory and Synapse pipelines
- Implement version control for pipeline artifacts
- Manage Spark jobs in a pipeline