Using Large Datasets with pandas
37mIntermediate2024-02-13
Authors

Miki Tebeka
CEO at 353Solutions
Course details
As data grows in size and complexity, most enterprises start to think about how to migrate to a larger-format data system such as Spark. However, this move can be quite painful, and you’ll most likely need to learn an entirely new set of tools. In this course, join instructor Miki Tebeka to learn how to get started working with large datasets using pandas, the fast, powerful, flexible, and easy-to-use data analysis tool built on top of the Python programming language. Find out how to navigate storage formats, tips for saving memory, efficient memory computation strategies, and more. Along the way, Miki also demonstrates how to leverage a handful of alternatives to pandas that still use the same API, such as Dask, Polars, and Beefy VM.
Skills covered
pandasData EngineeringData AnalysisData ScienceBusiness Analysis and StrategyBusiness Software and ToolsOpen SourceOne-Off
Concepts
0. Introduction
- 01 - Larget datasets with pandas
- 02 - What you should know
- 03 - Setting up
1. Big Data
- 04 - The problem with big data
- 05 - Big data is dead
- 06 - Big data systems
2. Using Less Memory
- 07 - Load only what you need
- 08 - Types
- 09 - Categorical data
- 10 - Arrow
- 11 - Challenge - Daily passengers
- 12 - Solution - Daily passengers
3. Memory Efficient Computation
- 13 - Iterative computation
- 14 - SQL
- 15 - Faster calculations
- 16 - Challenge - Maximum passengers
- 17 - Solution - Maximum passengers
4. Other Options
- 18 - dask
- 19 - polars
- 20 - Beefy VM
- 21 - Challenge - Median ride distance
- 22 - Solution - Median ride distance
Conclusion
- 23 - Next steps