Data Cleaning Fundamentals: Preparing Data for Analytics and AI

Data Cleaning Fundamentals: Preparing Data for Analytics and AI

1h 29mBeginner2026-08-17

Authors

Susan Walsh

Susan Walsh

Course details

Poor data quietly undermines reporting, decision-making, and business performance, and its impact becomes even more critical as organizations invest in analytics and AI. In this course, you’ll learn why data quality matters, how to recognize the most common types of dirty data across sales, marketing, finance, and procurement, and what to do about it. You’ll be introduced to the COAT framework (Consistent, Organized, Accurate, Trustworthy) and shown how to apply it to assess and improve the data you rely on every day. Through practical Excel techniques, you’ll learn how to identify, clean, standardize, and structure messy datasets so they become usable and reliable. By the end, you’ll be able to build a simple, repeatable approach to data cleaning that supports better reporting, stronger decisions, and AI-ready data.

Learning objectives
Explain why data quality matters and how poor data impacts reporting, decision-making, business performance and AI.
Recognize the most common types of dirty data across sales, marketing, finance, and procurement datasets.
Apply the COAT framework (Consistent, Organized, Accurate, Trustworthy) to assess and improve the quality of your data.
Use Excel techniques to identify, clean, standardize, and structure messy datasets.
Create a practical, repeatable approach to cleaning data that supports analytics, automation, and AI readiness.

Concepts

Introduction

  • Building AI-ready, trustworthy data

What Is Dirty Data and Why Does It Matter

  • What is dirty data and why does it matter
  • AI amplifies dirty data and requires more governance
  • Eliminating variations that break analysis and AI
  • Structuring data for scale and reuse
  • Reducing errors and assumptions
  • Justifying the data cleaning investment to your organization

Improve Data Quality with the COAT Framework

  • Introduction to the COAT framework
  • Consistent - Standardising data for accurate analysis and AI
  • Organized - Structuring data for scale and reuse
  • Accurate - Reducing errors and assumptions
  • Trustworthy - Knowing when data is fit for purpose
  • Maintenance - Keeping the data COAT on
  • Assessing real datasets using the COAT framework

Buiding a Repeatable AI-Ready Data Cleaning Process

  • Choose the right approach based on scale
  • Setting realistic expectations with stakeholders
  • Good habits - Normalization and classification spot checking
  • Good habits - Address spot checking
  • Stop cleaning the same data twice - Build master lists

Improving Data Quality using Excel

  • Working safely and efficiently in Excel for data cleaning
  • Excel formulas that support repeatable data cleaning
  • Excel functions that support repeatable data cleaning
  • Using Copilot in Excel to identify data quality issues

Final Project - Preparing an AI-Ready Dataset

  • Final project introduction
  • Final project walkthrough

Conclusion

  • Better data means better decisions
40,000 Toman