Complete Guide to Databricks for Data Engineering
6h 10mIntermediate2025-02-28
Authors

Deepak Goyal
Course details
In this course, master Databricks to become an ace data engineer. Learn how to expertly debug, process, and analyze huge amounts of data and build scalable solutions as instructor Deepak Goyal guides you through a deep dive on how the Databricks platform works. Explore PySpark transformation and Spark SQL in Databricks, along with how to read and write a DataFrame in Databricks. Plus, learn about Delta Lake, join optimizations, notebook scheduling, cluster management, workflows, and more.
Learning objectives
Learn how Databricks works end to end.
Understand the Delta Lake in Databricks.
Perform transformations in PySpark.
Use Spark SQL in Databricks.
Write DataFrames in PySpark.
Understand cluster management in Databricks.
Explore the Unity catalog in Databricks.
Learning objectives
Learn how Databricks works end to end.
Understand the Delta Lake in Databricks.
Perform transformations in PySpark.
Use Spark SQL in Databricks.
Write DataFrames in PySpark.
Understand cluster management in Databricks.
Explore the Unity catalog in Databricks.
Skills covered
DatabricksData EngineeringData ScienceOne-Off
Concepts
0. Introduction
- 01 - Course introduction
- 02 - What you should know
1. Introduction to Databricks
- 03 - What is Databricks
- 04 - Setting up a Databricks workspace
- 05 - Navigating the Databricks interface
- 06 - Introduction to Databricks notebooks
- 07 - Create a single-node cluster for practice
2. Getting Started with Databricks
- 08 - Understanding the Databricks File System (DBFS)
- 09 - Load sample data in DBFS
- 10 - Browse and explore data in DBFS
3. Read Data with Databricks
- 11 - Understand DataFrames
- 12 - Read a CSV file in Databricks
- 13 - Use a schema to read a file in Databricks
- 14 - Read a JSON file in Databricks
- 15 - Read a Parquet file in Databricks
- 16 - Handle nested JSON data in Databricks
4. PySpark Transformation in Databricks
- 17 - Use filter and where transformations in PySpark
- 18 - Add or remove columns in PySpark
- 19 - Use the select function in PySpark
- 20 - Use UNION and DISTINCT in PySpark
- 21 - Handle nulls in PySpark
- 22 - Use sortBy and orderBy in PySpark
- 23 - Use groupBy and aggregation in PySpark
- 24 - Manipulate strings in PySpark
- 25 - Handle date manipulation in PySpark
- 26 - Handle timestamp manipulation in PySpark
5. Write a DataFrame in Databricks
- 27 - Write a DataFrame as a file in DBFS
- 28 - Write a DataFrame as using partitioning
6. Spark SQL in Databricks
- 29 - What is Spark SQL
- 30 - Create temporary views in Databricks
- 31 - Create global temp views in Databricks
- 32 - Use Spark SQL transformations
- 33 - Write DataFrames as managed tables in PySpark
- 34 - Write a DataFrame as external table in PySpark
7. Delta Lake and Delta Tables in Databricks
- 35 - What is Delta Lake and its benefits
- 36 - Create Delta tables
- 37 - Handle DML operations in Delta tables
- 38 - Time travel using Delta Lake
8. Join Optimizations in Databricks
- 39 - Handle multiple types of join
- 40 - Broadcast join
- 41 - Bucketing in PySpark
9. Scheduling the Notebook
- 42 - Notebook job scheduling
10. Cluster Management in Databricks
- 43 - Understand the interactive cluster
- 44 - Explore cluster configuration and the UI
- 45 - Understand job clusters
11. Workflows in Databricks
- 46 - Understand workflows in Databricks
- 47 - Create a workflow in Databricks
12. dbutils in Databricks
- 48 - What is dbutils in Databricks
- 49 - dbutils fs commands
- 50 - dbutils mounting
- 51 - dbutils notebook
13. Unity Catalog in Databricks
- 52 - Understand the Unity Catalog
14. Capstone Project
- 53 - Project use case
- 54 - Solution
Conclusion
- 55 - Next steps