Special offers now — see discounted courses.
day
:
hour
:
min
:
sec
See special offers
PySpark Essential Training: Introduction to Building Data Pipelines

PySpark Essential Training: Introduction to Building Data Pipelines

1h 8mIntermediate2025-08-07

Authors

Sam Bail

Sam Bail

Course details

PySpark is a powerful library that brings Apache Spark’s distributed computing capabilities to Python, making it a key tool for processing large-scale data efficiently. In this course, data engineer and analyst Sam Bail provides a structured and hands-on introduction to PySpark, starting with an overview of Apache Spark, its architecture, and its ecosystem. Learn about Spark’s core concepts, such as the DataFrame API, transformations, lazy evaluations, and actions, before setting up a lab environment and working with a real dataset. Plus, gain insights into how PySpark fits into a broader data engineering ecosystem and best practices on running PySpark in a production environment.

Learning objectives
Build your understanding of the core concepts of Spark and PySpark.
Understand how to install PySpark, load, manipulate, and analyze large datasets in a notebook environment.
Gain an understanding of how PySpark fits into a wider data engineering ecosystem.
Understand best practices about executing PySpark in a production environment.

Skills covered

Data EngineeringPythonEssential TrainingData ScienceOpen Source

Concepts

0. Introduction

  • 01 - Course overview
  • 02 - Prerequisites
  • 03 - Using GitHub repo

1. Introduction to Spark and PySpark

  • 04 - Introduction to Apache Spark - The foundation of PySpark
  • 05 - The Apache Spark ecosystem
  • 06 - Spark vs. PySpark

2. Setting Up PySpark

  • 07 - Google Colab notebook setup
  • 08 - Downloading a dataset

3. Working with PySpark DataFrames

  • 09 - Introduction to PySpark DataFrames
  • 10 - Data formats and loading data
  • 11 - Schema and data types
  • 12 - Basic querying (select, filter, and sort)
  • 13 - Challenge - Querying a DataFrame
  • 14 - Solution - Querying a DataFrame

4. Essential PySpark Data Manipulation

  • 15 - Handling missing data
  • 16 - Creating new columns
  • 17 - Unions and joins
  • 18 - Aggregating
  • 19 - Writing data
  • 20 - Challenge - Essential data manipulation
  • 21 - Solution - Essential data manipulation

5. PySpark SQL

  • 22 - What is PySpark SQL
  • 23 - Creating temporary views
  • 24 - Using SQL queries
  • 25 - Challenge - PySpark SQL
  • 26 - Solution - PySpark SQL

6. PySpark in a Production Environment

  • 27 - Production environment requirements
  • 28 - Example production environment setup
  • 29 - A typical PySpark production workflow
  • 30 - Cloud services

Conclusion

  • 31 - Recap of key concepts and next steps

About us

LyndaKade is a leading learning platform that helps people learn business, software, technology, and creative skills to achieve personal and professional goals.

Phone numberAparat ChannelTelegram SupportTelegram ChannelInstagram Page

All rights to this site belong to LyndaKade.

Terms of Service|Privacy Policy

نماد الکترونیک enamad در صورت اتصال با آی‌پی داخل کشور، نمایش داده خواهد شد.
logo-samandehi - لوگو ساماندهی
Zarinpal
Zibal