Apache Spark Essential Training: Big Data Engineering

Apache Spark Essential Training: Big Data Engineering

1h 5mIntermediate2024-01-01

Authors

Kumaran Ponnambalam

Kumaran Ponnambalam

Working with data for 20+ years

Course details

Data engineering is the foundation for building analytics and data science applications in the new Big Data world. Data engineering requires combining multiple big data technologies to construct data pipelines and networks to stream, process, and store data. This course focuses on building full-fledged solutions that combine Apache Spark with other big data tools to create end-to-end data pipelines. Instructor Kumaran Ponnambalam begins by defining data engineering, its functions, and its concepts. Next, Kumaran goes over how Spark capabilities such as parallel processing, execution plans, state management options, and machine learning work with extract, transform, load (ETL). He introduces you to batch processing use cases and processes, as well as real-time processing pipelines. After taking you through several useful best practices, Kumaran concludes with an end-to-end exercise project.

Skills covered

Apache SparkApacheData EngineeringData AnalysisData ScienceBusiness Analysis and StrategyBusiness Software and ToolsOne-Off

Concepts

Introduction

  • Driving big data engineering with Apache Spark
  • Course prerequisites
  • Setting up the exercise files

Data Engineering Concepts

  • What is data engineering
  • Data engineering vs. data analytics vs. data science
  • Data engineering functions
  • Batch vs. real-time processing
  • Data engineering with Spark

Spark Capabilities for ETL

  • Spark architecture review
  • Parallel processing with Spark
  • Spark execution plan
  • Stateful stream processing
  • Spark analytics and ML

Batch Processing Pipelines

  • Batch processing use case - Problem statement
  • Batch processing use case - Design
  • Setting up the local DB
  • Uploading stock to a central store
  • Aggregating stock across warehouses

Real-Time Processing Pipelines

  • Real-time use case - Problem
  • Real-time use case - Design
  • Generating a visits data stream
  • Building a website analytics job
  • Executing the real-time pipeline

Data Engineering with Spark - Best Practices

  • Batch vs. real-time options
  • Scaling extraction and loading operations
  • Scaling processing operations
  • Building resiliency

End-to-End Exercise Project

  • Project exercise requirements
  • Solution design
  • Extracting long last actions
  • Building a scorecard

Conclusion

  • More about Apache Spark
40,000 Toman