PySpark Essential Training: Introduction to Building Data Pipelines
1h 8mIntermediate2025-08-07
Authors

Sam Bail
Course details
PySpark is a powerful library that brings Apache Spark’s distributed computing capabilities to Python, making it a key tool for processing large-scale data efficiently. In this course, data engineer and analyst Sam Bail provides a structured and hands-on introduction to PySpark, starting with an overview of Apache Spark, its architecture, and its ecosystem. Learn about Spark’s core concepts, such as the DataFrame API, transformations, lazy evaluations, and actions, before setting up a lab environment and working with a real dataset. Plus, gain insights into how PySpark fits into a broader data engineering ecosystem and best practices on running PySpark in a production environment.
Learning objectives
Build your understanding of the core concepts of Spark and PySpark.
Understand how to install PySpark, load, manipulate, and analyze large datasets in a notebook environment.
Gain an understanding of how PySpark fits into a wider data engineering ecosystem.
Understand best practices about executing PySpark in a production environment.
Learning objectives
Build your understanding of the core concepts of Spark and PySpark.
Understand how to install PySpark, load, manipulate, and analyze large datasets in a notebook environment.
Gain an understanding of how PySpark fits into a wider data engineering ecosystem.
Understand best practices about executing PySpark in a production environment.
Skills covered
Data EngineeringPythonEssential TrainingData ScienceOpen Source
Concepts
0. Introduction
- 01 - Course overview
- 02 - Prerequisites
- 03 - Using GitHub repo
1. Introduction to Spark and PySpark
- 04 - Introduction to Apache Spark - The foundation of PySpark
- 05 - The Apache Spark ecosystem
- 06 - Spark vs. PySpark
2. Setting Up PySpark
- 07 - Google Colab notebook setup
- 08 - Downloading a dataset
3. Working with PySpark DataFrames
- 09 - Introduction to PySpark DataFrames
- 10 - Data formats and loading data
- 11 - Schema and data types
- 12 - Basic querying (select, filter, and sort)
- 13 - Challenge - Querying a DataFrame
- 14 - Solution - Querying a DataFrame
4. Essential PySpark Data Manipulation
- 15 - Handling missing data
- 16 - Creating new columns
- 17 - Unions and joins
- 18 - Aggregating
- 19 - Writing data
- 20 - Challenge - Essential data manipulation
- 21 - Solution - Essential data manipulation
5. PySpark SQL
- 22 - What is PySpark SQL
- 23 - Creating temporary views
- 24 - Using SQL queries
- 25 - Challenge - PySpark SQL
- 26 - Solution - PySpark SQL
6. PySpark in a Production Environment
- 27 - Production environment requirements
- 28 - Example production environment setup
- 29 - A typical PySpark production workflow
- 30 - Cloud services
Conclusion
- 31 - Recap of key concepts and next steps