Special offers now — see discounted courses.
day
:
hour
:
min
:
sec
See special offers
Data Ingestion with Python

Data Ingestion with Python

1h 24mIntermediate2023-12-22

Authors

Miki Tebeka

Miki Tebeka

CEO at 353Solutions

Course details

A sizable portion of a data scientist's day is often spent fetching and cleaning the data they need to train their algorithms. In this course, learn how to use Python tools and techniques to get the relevant, high-quality data you need. Instructor Miki Tebeka covers reading files, including how to work with CSV, XML, and JSON files. He also discusses calling APIs, web scraping (and why it should be a last resort), and validating and cleaning data. Plus, discover how to establish and monitor key performance indicators (KPIs) that help you monitor your data pipeline.

Learning objectives
Describe the characteristics of different data types and the work of data scientists.
Describe different data serialization formats and explain how to use them in Python.
Define APIs and explain how to use them with Python to make http calls, interpret JSON, and utilize message queues.
Explain what web scraping is and describe ways to do it.
Define what a schema is and describe characteristics of schemas and how they influence operations.
Describe the characteristics of different types of databases.
Categorize types of errors and explain how to correct them.
Explain design criteria for data systems and describe how to monitor performance using KPIs.

Skills covered

Data EngineeringPythonProgramming LanguagesData ScienceOpen SourceSoftware DevelopmentOne-Off

Concepts

0. Introduction

  • 01 - Why is data ingestion important
  • 02 - What you should know
  • 03 - Using the exercise files
  • 04 - Using the Coderpad quizzes

1. Data Ingestion Overview

  • 05 - Overview of data scientists work
  • 06 - Where does data come from
  • 07 - Different types of data
  • 08 - The data pipeline (ETL)
  • 09 - Final destination (data lake)

2. Reading Files

  • 10 - Working in CSV
  • 11 - Working in XML
  • 12 - Working in Parquet, Avro, and ORC
  • 13 - Unstructured text
  • 14 - JSON
  • 15 - Solution - CSV to JSON

3. Calling APIs

  • 16 - Working with JSON
  • 17 - Making HTTP calls
  • 18 - Processing event-based data
  • 19 - Solution - Location from IP

4. Web Scraping

  • 20 - Try to find an API
  • 21 - Working with Beautiful Soup
  • 22 - Working with Scrapy
  • 23 - Working with Selenium
  • 24 - Other considerations
  • 25 - Solution - Get stock information from HTML

5. Schema

  • 26 - What are schemas
  • 27 - Working with ontologies
  • 28 - What should be in schema
  • 29 - Schema changes
  • 30 - Schema validations

6. Working with Databases

  • 31 - Types of databases
  • 32 - Hosted and cost of ops
  • 33 - Working with relational databases
  • 34 - Working with key or value databases
  • 35 - Working with document databases
  • 36 - Working with graph databases
  • 37 - Solution - ETL

7. Troubleshooting Data

  • 38 - Data is never 100 okay
  • 39 - Causes of errors
  • 40 - Filling missing values
  • 41 - Finding outliers (manual)
  • 42 - Finding outliers (ML)
  • 43 - Solution - Clean rides dataset

8. Data KPIs and Process

  • 44 - Design your data
  • 45 - KPIs
  • 46 - What to monitor

Conclusion

  • 47 - Next steps

About us

LyndaKade is a leading learning platform that helps people learn business, software, technology, and creative skills to achieve personal and professional goals.

Phone numberAparat ChannelTelegram SupportTelegram ChannelInstagram Page

All rights to this site belong to LyndaKade.

Terms of Service|Privacy Policy

نماد الکترونیک enamad در صورت اتصال با آی‌پی داخل کشور، نمایش داده خواهد شد.
logo-samandehi - لوگو ساماندهی
Zarinpal
Zibal