Learning Hadoop
1h 53mIntermediate2023-10-20
Authors

Lynn Langit
Cloud Architect
Course details
Hadoop is indispensable when it comes to processing big data—as necessary to understanding your information as servers are to storing it. In this course, cloud architect Lynn Langit provides a thorough introduction to Hadoop. Find out how to set up Cloud Hadoop and learn about core components like JVMs, the HDFS file system, AWS S3, and cluster components. Step through the process of setting up and verifying your development environment. Explore ways you can use MapReduce with Hadoop and learn how to tune each MapReduce job. Go over scaling VM-based Hadoop clusters on GCP Dataproc HDFS. Learn how to select appropriate NoSQL options for Hadoop with Hive, HBase, and Pig. Plus, dive into Apache Spark architecture and how to run an Apache Spark job on a Hadoop cluster.
Skills covered
HadoopApacheData EngineeringLearningData Science
Concepts
0. Introduction
- 01 - What and why Hadoop
- 02 - What you should know
- 03 - Use cloud services
1. Set Up Cloud Hadoop
- 04 - What is Hadoop
- 05 - Review Hadoop distributions and cloud services
- 06 - Set up GCP Dataproc Metastore and VM cluster
- 07 - Verify GCP Dataproc VM cluster
2. Understand Hadoop Core Components
- 08 - Understand Hadoop components
- 09 - Understand Java virtual machines (JVMs)
- 10 - Explore Hadoop file systems - HDFS
- 11 - Explore Hadoop file systems - AWS S3
- 12 - Review Hadoop cluster components
3. Set Up and Verify Development Environment
- 13 - Review test jobs
- 14 - Review job output
- 15 - Verify Hadoop web interfaces in your test environment
- 16 - Verify Hadoop Spark web interfaces in your test environment
- 17 - Use the Jupyter interface for Hadoop
4. Understand MapReduce
- 18 - What is MapReduce
- 19 - What is MapReduce word count
- 20 - Review MapReduce word count job
- 21 - Prepare for MapReduce Java coding
- 22 - Review MapReduce WordCount job code
5. Tune MapReduce
- 23 - Tune by physical methods
- 24 - Tune a Mapper
- 25 - Understanding data types
- 26 - Tune a Reducer
- 27 - Use MR 2.0 and 3.0
- 28 - Review MR optimization examples
6. Scale Cloud Hadoop
- 29 - Migrate to Cloud Hadoop
- 30 - Scale VM-based Clusters
- 31 - Use autoscale policies
- 32 - Scale Kubernetes Spark clusters
7. Use Hive, Pig, and Spark
- 33 - Understand Hive and HBase
- 34 - Create and query tables with Hive
- 35 - Understand Pig
- 36 - Run WordCount using Pig
- 37 - Review Spark architecture
- 38 - Scale a Spark job to calculate Pi
Conclusion
- 39 - Learn more about using Hadoop