Cloud Hadoop: Scaling Apache Spark
3h 16mBeginner2025-02-07
Authors

Lynn Langit
Cloud Architect
Course details
Apache Hadoop and Spark make it possible to generate genuine business insights from big data. The Amazon cloud is natural home for this powerful toolset, providing a variety of services for running large-scale data-processing workflows. Learn to implement your own Apache Hadoop and Spark workflows on AWS in this course with big data architect Lynn Langit. Explore deployment options for production-scaled jobs using virtual machines with EC2, managed Spark clusters with EMR, or containers with EKS. Learn how to configure and manage Hadoop clusters and Spark jobs with Databricks, and use Python or the programming language of your choice to import data and execute jobs. Plus, learn how to use Spark libraries for machine learning, genomics, and streaming. Each lesson helps you understand which deployment option is best for your workload.
Learning objectives
File systems for Hadoop and Spark
Working with Databricks
Loading data into tables
Setting up Hadoop and Spark clusters on the cloud
Running Spark jobs
Importing and exporting Python notebooks
Executing Spark jobs in Databricks using Python and Scala
Importing data into Spark clusters
Coding and executing Spark transformations and actions
Data caching
Spark libraries: Spark SQL, SparkR, Spark ML, and more
Spark streaming
Scaling Spark with AWS and GCP
Learning objectives
File systems for Hadoop and Spark
Working with Databricks
Loading data into tables
Setting up Hadoop and Spark clusters on the cloud
Running Spark jobs
Importing and exporting Python notebooks
Executing Spark jobs in Databricks using Python and Scala
Importing data into Spark clusters
Coding and executing Spark transformations and actions
Data caching
Spark libraries: Spark SQL, SparkR, Spark ML, and more
Spark streaming
Scaling Spark with AWS and GCP
Skills covered
Apache SparkApacheData EngineeringData ScienceDeep Dive (X:Y)
Concepts
0. Introduction
- 01 - Scaling Apache Hadoop and Spark
1. Hadoop and Spark Fundamentals
- 02 - Modern Hadoop and Spark
- 03 - File systems used with Hadoop and Spark
- 04 - Apache or commercial Hadoop distros
- 05 - Hadoop and Spark libraries
- 06 - Hadoop on Google Cloud Platform
- 07 - Spark Job on Google Cloud Platform
2. AWS Cloud Spark Environments
- 08 - Sign up for Databricks Community Edition
- 09 - Add Hadoop libraries
- 10 - Databricks AWS Community Edition
- 11 - Load data into tables
- 12 - Hadoop and Spark cluster on AWS EMR
- 13 - Run Spark job on AWS EMR
- 14 - Review batch architecture for ETL on AWS
3. Spark Basics
- 15 - Apache Spark libraries
- 16 - Spark data interfaces
- 17 - Select your programming language
- 18 - Spark session objects
- 19 - Spark shell
4. Using Spark
- 20 - Tour the Databricks Environment
- 21 - Tour the notebook
- 22 - Import and export notebooks
- 23 - Calculate Pi on Spark
- 24 - Run WordCount of Spark with Scala
- 25 - Import data
- 26 - Transformations and actions
- 27 - Caching and the DAG
- 28 - Architecture - Streaming for prediction
5. Spark Libraries
- 29 - Spark SQL
- 30 - SparkR
- 31 - Spark ML - Preparing data
- 32 - Spark ML - Building the model
- 33 - Spark ML - Evaluating the model
- 34 - Advanced machine learning on Spark
- 35 - MXNet
- 36 - Spark with ADAM for genomics
- 37 - Spark architecture for genomics
6. Spark Streaming
- 38 - Reexamine streaming pipelines
- 39 - Spark Streaming
- 40 - Streaming ingest services
- 41 - Advanced Spark Streaming with MLeap
7. Scaling Spark on AWS and GCP
- 42 - Scale Spark on the cloud by example
- 43 - Build a quick start with Databricks AWS
- 44 - Scale Spark cloud compute with VMs
- 45 - Optimize cloud Spark virtual machines
- 46 - Use AWS EKS containers and data lake
- 47 - Optimize Spark cloud data tiers on Kubernetes
- 48 - Build reproducible cloud infrastructure
- 49 - Scale on GCP Dataproc or on Terra.bio
- 50 - Serverless Spark with Dataproc Notebook
Conclusion
- 51 - Continue learning for scaling