Manage and Optimize Big Data with Apache Iceberg
1h 24mIntermediate2024-08-16
Authors

Deepak Goyal
Course details
Big data is only getting bigger, and Apache Iceberg has become one of the most popular formats for advanced management of large data sets. In this course, Deepak Goyal covers the architecture, schema management, and table operations of Iceberg, equipping you with the skills to efficiently manage scalable data lakes, ensuring data integrity and optimizing query performance. The course also contains practical examples and demo sessions for you to get hands-on experience and take your data skills to the next level.
Skills covered
Data Resource ManagementSQLData EngineeringDatabase ManagementData AnalysisData ScienceBusiness Analysis and StrategyBusiness Software and ToolsOpen SourceOne-Off
Concepts
0. Introduction
- 01 - Apache Iceberg and big data
- 02 - What you should know
1. Getting Started with Apache Iceberg
- 03 - Apache Iceberg introduction
- 04 - Role of Iceberg in modern data architecture
- 05 - Key features and advantages
- 06 - Setting up your Iceberg environment on Databricks
- 07 - Creating your first Iceberg table - Practical
- 08 - Challenge - Create an Iceberg table
- 09 - Solution - Create an Iceberg table
2. Deep Dive into the Core Concepts of Iceberg
- 10 - Iceberg table structure and metadata
- 11 - Schema management and evolution
- 12 - Implement schema changes
- 13 - Partitioning in Iceberg
- 14 - Implementing partitioning
- 15 - Challenge - Change the table schema
- 16 - Solution - Change the table schema
3. Advanced Data Management with Apache Iceberg
- 17 - Time travel and snapshot management
- 18 - Implementing time travel
- 19 - Transactional data operations and concurrency control
4. Optimizing and Scaling with Apache Iceberg
- 20 - Optimizing file layout and size for performance
- 21 - Advanced catalog integration
- 22 - Performance tuning strategies
- 23 - Iceberg alternatives
Conclusion
- 24 - Next steps