Microsoft Azure Data Engineer Associate (DP-203) Cert Prep: 1 Design and Implement Data Storage

Microsoft Azure Data Engineer Associate (DP-203) Cert Prep: 1 Design and Implement Data Storage

3h 1mIntermediate2023-08-14

Authors

Microsoft Learn

Microsoft Learn

Build Skills That Open Doors

Tim Warner

Tim Warner

Technical Trainer and Content Developer

Course details

The latest professional certifications from Azure are aligned with specific industry roles. Earning your Azure certification helps to validate your unique Azure skill set and increase your value in today's IT job market. Join Microsoft MVP and Microsoft Certified Azure Solutions Architect Tim Warner for an overview of the core concepts and skills required to pass the Microsoft Azure Data Engineer Associate (DP-203) certification exam. In this course, the first in a four-part certification prep series, explore the fundamentals of how to design and implement data storage solutions. Ensure that data pipelines and data stores are high performing, efficient, organized, and reliable, given a set of business requirements and constraints specific to your organization. Learn how to deal with unanticipated issues swiftly, minimize data loss, and design, implement, monitor, and optimize data platforms and structures to meet data pipeline needs.

Skills covered

Cloud StorageData EngineeringAzureNetwork AdministrationCloud PlatformsCert PrepNetwork and System AdministrationCloud ComputingData ScienceMicrosoft

Concepts

Introduction

  • Introduction

Design and Implement Data Storage

  • Learning objectives
  • Design an Azure Data Lake solution
  • Recommend file types for storage
  • Recommend file types for analytical queries
  • Design for efficient querying

Design for Data Pruning

  • Learning objectives
  • Design a folder structure that represents levels of data transformation
  • Design a distribution strategy
  • Design a data archiving solution

Design a Partition Strategy

  • Learning objectives
  • Design a partition strategy for files
  • Design a partition strategy for analytical workloads
  • Design a partition strategy for efficiency and performance
  • Design a partition strategy for Azure Synapse Analytics
  • Identify when partitioning is needed in Azure Data Lake Storage Gen2

Design the Serving Layer

  • Learning objectives
  • Design star schemas
  • Design slowly changing dimensions
  • Design a dimensional hierarchy
  • Design a solution for temporal data
  • Design for incremental loading
  • Design analytical stores
  • Design metastores in Azure Synapse Analytics and Azure Databricks

Implement Physical Data Storage Structures

  • Learning objectives
  • Implement compression
  • Implement partitioning
  • Implement sharding
  • Implement different table geometries with Azure Synapse Analytics pools
  • Implement data redundancy
  • Implement distributions
  • Implement data archiving

Implement Logical Data Structures

  • Learning objectives
  • Build a temporal data solution
  • Build a slowly changing dimension
  • Build a logical folder structure
  • Build external tables
  • Implement file and folder structures for efficient querying and data pruning

Implement the Serving Layer

  • Learning objectives
  • Deliver data in a relational star schema
  • Deliver data in Parquet files
  • Maintain metadata
  • Implement a dimensional hierarchy
80,000 Toman