Complete Guide to Data Lakes and Lakehouses
4h 24mAdvanced2024-08-30
Authors

Thalia Barrera
Course details
In this course, data engineer and technical writer Thalia Barrera offers an introductory yet comprehensive overview of data lakes. Learn about key concepts like data lake architecture, operation, and integration with existing data systems. Delve into how data lakes are integral to AI and machine learning workflows. Go over the differences between data lakes, data warehouses, and databases. Explore various data formats and their applicability in a data lake environment. Use included hands-on exercises to practice setting up a basic data lake and performing simple data operations. When you finish this course, you will be equipped to make informed decisions about implementing and managing data lakes in your organization.
Skills covered
Artificial Intelligence FoundationsData EngineeringArtificial Intelligence (AI)Data ScienceOne-Off
Concepts
Introduction
- Data lakes, lakehouses, and more
- What you should know
- Capstone project preview
Introduction to Data Lakes
- What is a data lake
- Origins and evolution
- Architecture core components
- Data lake vs. data warehouse
- Data lake vs. data mesh
Storage In Data Lakes
- Storage types
- Storage hosting
- Storage solutions - S3, GCS and Azure Blob Storage and HDFS
- Folder structures
- File formats
- Data compression
- Data partitioning
Data Ingestion in Data Lakes
- Data ingestion methods
- ETL vs. ELT
- Data transformation
- Data quality
- Error handling, logging, and monitoring
- Orchestration
- Data ingestion platforms
Data Management and Governance in Data Lakes
- Introduction to data management and governance
- Metadata management
- Data cataloging
- Data lineage
- Data security, privacy, and compliance
- Data management tools and platforms
Introduction to Data Lakehouses
- What is a data lakehouse
- ACID transactions
- Schema management
- Table formats - Delta Lake, Apache Iceberg, Apache Hudi
Data Consumption and Query Engines in Lakes and Lakehouses
- Introduction to data consumption
- Unified data analysis - Spark
- SQL on Hadoop - Hive and Impala
- Interactive query engines - Presto and Trino
- Data indexing
- Optimizing query performance
- Data consumption security considerations
Advanced Data Platforms for Lakes and Lakehouses
- Unified analytics platforms - Databricks and Snowflake
- Cloud data warehouses - BigQuery, Azure Synapse, and Redshift
- Self-service data platforms - Dremio and Starburst
- Interactive notebooks - Jupyter, Zeppelin, Databricks
- BI tools - Tableau, Power BI, Superset, Metabase
- APIs and services for data consumption
Capstone - Building a Data Lakehouse
- Capstone project overview
- Data model overview
- Project installation and code walkthrough
- Infrastructure setup
- Raw data ingestion
- Transformation models overview
- Solution - Build a data model with SQL
- Executing data transformations
- Data orchestration
Capstone - BI, Advanced Analytics, and ML in the Lakehouse
- Dremio walkthrough
- Executing queries and creating virtual datasets
- Creating complex virtual datasets using SQL
- Connecting Dremio to Apache Superset
- Creating a marketing dashboard
- Connecting Dremio to Jupyter Notebook
- Advanced product reviews analytics
- Solution - Vehicle health analytics in Jupyter
Capstone - Generative AI in the Lakehouse
- Introduction to LLMs and vector embeddings - Llama
- Introduction to RAG (retrieval-augmented generation)
- Introduction to vector databases - Chroma
- What is Langchain
- Generative AI project overview - Sales copilot
- Installation and code walkthrough
- Project execution - Using the copilot
Conclusion
- Recap and key takeaways
- Next steps on your data journey