Microsoft Azure Data Engineer Associate (DP-203) Cert Prep: 4 Monitor and Optimize Data Storage and Data Processing

Microsoft Azure Data Engineer Associate (DP-203) Cert Prep: 4 Monitor and Optimize Data Storage and Data Processing

1h 9mIntermediate2023-08-14

Authors

Microsoft Learn

Microsoft Learn

Build Skills That Open Doors

Tim Warner

Tim Warner

Technical Trainer and Content Developer

Course details

The latest professional certifications from Azure are aligned with specific industry roles. Earning your Azure certification helps to validate your unique Azure skill set and increase your value in today's IT job market. Join Microsoft MVP and Microsoft Certified Azure Solutions Architect Tim Warner for an overview of the core concepts and skills required to pass the Microsoft Azure Data Engineer Associate (DP-203) certification exam. In this course, the last in a four-part certification prep series, discover how to monitor and optimize data storage and data processing solutions. Along the way, get tips on fine-tuning strategies for using compact small files, rewriting UDRs in Azure, handling data skew and spill, tuning shuffle partitions, finding shuffling in a pipeline, and optimizing overall resource management.

Skills covered

Cloud StorageData EngineeringAzureNetwork AdministrationCloud PlatformsCert PrepNetwork and System AdministrationCloud ComputingData ScienceMicrosoft

Concepts

Monitor Data Storage

  • Learning objectives
  • Implement logging used by Azure Monitor
  • Configure monitoring services
  • Measure performance of data movement
  • Monitor and update statistics about data across a system
  • Monitor data pipeline performance
  • Measure query performance

Monitor Data Processing

  • Learning objectives
  • Monitor cluster performance
  • Understand custom logging options
  • Schedule and monitor pipeline tests
  • Interpret Azure Monitor metrics and logs
  • Interpret a Spark Directed Acyclic Graph (DAG)

Tune Data Storage

  • Learning objectives
  • Compact small files
  • Rewrite user-defined functions (UDFs)
  • Handle skew in data
  • Handle data spill
  • Tune shuffle partitions
  • Find shuffling in a pipeline
  • Optimize resource management

Optimize and Troubleshoot Data Processing

  • Learning objectives
  • Tune queries by using indexers
  • Tune queries by using cache
  • Optimize pipelines for analytical or transactional purposes
  • Optimize pipeline for descriptive versus analytical workloads
  • Troubleshoot failed Spark jobs
  • Troubleshoot failed pipeline runs

Conclusion

  • Summary
40,000 Toman