Cleaning Data for Effective Data Science: Data Ingestion, Anomaly Detection, Value Imputation, and Feature Engineering
4h 50mIntermediate2025-07-11
Authors

Pearson

David Mertz
Course details
The course introduces the tools and techniques needed for data ingestion, anomaly detection, value imputation, and feature engineering. Numerous ingested formats are addressed, including JSON, CSV, SQL RDBMS, HDF5, NoSQL databases, and binary serialized data structures. Instructor David Mertz outlines why some problems are peculiar to data representation, while others link to the data in itself. To address untidiness in data, learn how and when to impute missing values, detect unreliable data and statistical anomalies, and generate synthetic features that are necessary for successful data analysis and visualization goals. By the end of this course, you’ll be equipped with highly marketable and in-demand skills in data analysis, machine learning, and data integrity troubleshooting.
Learning objectives
Understand and process tabular and hierarchical data.
Identify and remediate bias and data anomalies.
Ingest data from diverse formats.
Impute values in a manner suitable to a purpose.
Engineer features for machine learning models.
Learning objectives
Understand and process tabular and hierarchical data.
Identify and remediate bias and data anomalies.
Ingest data from diverse formats.
Impute values in a manner suitable to a purpose.
Engineer features for machine learning models.
Skills covered
Data Science FoundationsData EngineeringData AnalysisData ScienceBusiness Analysis and StrategyBusiness Software and ToolsOne-Off
Concepts
0. Introduction
- 01 - Cleaning data for effective data science
- 02 - Introduction
- 03 - Doing the other 80 of the work
- 04 - Types of grime
- 05 - Nomenclature
- 06 - Visual rendering
- 07 - Data hygiene
1. Data Ingestion - Tabular Formats
- 08 - Topics
- 09 - CSV
- 10 - Spreadsheets considered harmful
- 11 - Other formats
2. Data Ingestion - Hierarchical Formats
- 12 - Topics
- 13 - XML
- 14 - JSON
- 15 - NoSQL databases
3. Data Ingestion - Repurposing Data Sources
- 16 - Topics
- 17 - Web scraping
- 18 - Portable Document Format
- 19 - Image formats
4. Anomaly Detection
- 20 - Topics
- 21 - Missing data
- 22 - SQL
- 23 - Hierarchical formats
- 24 - Sentinels
- 25 - Miscoded data
- 26 - Fixed bounds
- 27 - Outliers
5. Data Quality
- 28 - Topics
- 29 - Missing data
- 30 - Biasing trends
- 31 - Benford s law
- 32 - Class imbalance
- 33 - Normalization and scaling
6. Value Imputation
- 34 - Topics
- 35 - Typical value imputation
- 36 - Trend imputation
- 37 - Sampling
Conclusion
- 38 - Summary