Complete Guide to Generative AI for Data Analysis and Data Science
10h 17mIntermediate2024-09-27
Authors

Dan Sullivan
Enterprise Architect, Big Data Expert
Course details
GenAI has the potential to enable many more people to work with and analyze data, but to succeed, you need a solid foundation in data management, statistics, and machine learning. This course provides that foundation. Instructor Dan Sullivan teaches how to break down business questions and data science questions into components that can be addressed programmatically and then how to use genAI to create programs and scripts to implement a solution. This course focuses on the three pillars needed to be a successful data analyst or data scientist: problem solving skills, an understanding of statistics and machine learning, and practical experience with data management procedures.
Learning objectives
Explore data analytics and data science problems, starting with an overview of data analysis/data science methodology and processes.
Learn about statistics topics that include data types, descriptive statistics, inferential statistics, correlation analysis, regression, data wrangling, and data exploration.
Learn machine learning topics that include supervised and unsupervised learning, classification and regression, clustering, tree-based algorithms, neural networks and deep learning, and time series analysis and anomaly detection.
Learn about data management practices, including working with files and databases, data quality evaluation, data cleaning, metadata management, and data lifecycle management.
Learning objectives
Explore data analytics and data science problems, starting with an overview of data analysis/data science methodology and processes.
Learn about statistics topics that include data types, descriptive statistics, inferential statistics, correlation analysis, regression, data wrangling, and data exploration.
Learn machine learning topics that include supervised and unsupervised learning, classification and regression, clustering, tree-based algorithms, neural networks and deep learning, and time series analysis and anomaly detection.
Learn about data management practices, including working with files and databases, data quality evaluation, data cleaning, metadata management, and data lifecycle management.
Skills covered
ClaudeAnthropicPostgreSQLChatGPTSQLOpenAIGenerative AIPythonData AnalysisArtificial Intelligence (AI)Data ScienceBusiness Analysis and StrategyBusiness Software and ToolsOpen SourceOne-Off
Concepts
0. Introduction
- 01 - Getting started
1. Demystifying Data - Data Analysis and Data Science
- 02 - Asking questions
- 03 - Collecting and obtaining data
- 04 - Cleaning and preparing data
- 05 - Analyzing data
- 06 - Predictive modeling
- 07 - Machine learning
- 08 - Interpret the results
2. Tools of the Trade
- 09 - Problem-solving
- 10 - Statistics
- 11 - Machine learning algorithms
- 12 - Spreadsheets
- 13 - Python
- 14 - SQL and relational databases
- 15 - Statistics platforms
- 16 - Machine learning libraries
3. Thinking About Data
- 17 - Quantitative and qualitative data
- 18 - Discrete vs. continuous data
- 19 - Categorical data
4. Techniques for Describing Data
- 20 - Measures of central tendency
- 21 - Measures of spread
- 22 - Visualizing data distribution
- 23 - Describing a dataset using generative AI
- 24 - Challenge - Describing data
- 25 - Solution - Describing data
5. Distributions of Data
- 26 - Distributions of data
- 27 - Visualizing a normal distribution in a spreadsheet
- 28 - Jupyter Notebook and Colab
- 29 - Generating a normal distribution
- 30 - Visualizing a normal distribution in Python
- 31 - Visualizing a uniform distribution in Python
- 32 - Visualizing a bimodal distribution in Python
- 33 - Challenge - Distributions of data
- 34 - Solution - Distribution of data
6. Sampling Data
- 35 - Sampling and large populations
- 36 - Creating samples
- 37 - Saving samples to a file
- 38 - Comparing population to sample statistics
- 39 - Challenge - Sampling data
- 40 - Solution - Sampling data
7. Making Inferences from Data
- 41 - Inferential statistics
- 42 - Hypothesis testing methodology
- 43 - Analyzing customer preferences
- 44 - Type I and type II errors
- 45 - ANOVA tests for comparing means
- 46 - Generating Python scripts for ANOVA
- 47 - Testing independence of categorical variables
- 48 - Generating Python Scripts for Chi-squared tests
- 49 - Correlation analysis
- 50 - Testing for normality
- 51 - Generating Python for testing normality
- 52 - Generating Python for correlation analysis
- 53 - Challenge - Making inferences from data
- 54 - Solution - Making inferences from data
8. Visualizing Data
- 55 - Visualizing data
- 56 - Visualizing trends
- 57 - Visualizing correlations
- 58 - Visualizing composition
- 59 - Visualizing distributions
- 60 - Challenge - Visualizing data
- 61 - Solution - Visualizing data
9.Regression
- 62 - Linear regression
- 63 - Evaluating linear regression models
- 64 - Visualizing sales data
- 65 - Building a linear regression model
- 66 - Evaluating a sales linear regression model
- 67 - Challenge - Building a regression model
- 68 - Solution - Building a regression model
10. Analyzing Data in Files
- 69 - Data files
- 70 - Using spreadsheets with CSV files
- 71 - Reviewing an example JSON file
- 72 - Using jq with JSON files
- 73 - Generating jq commands using AI
- 74 - Dataframes in Python
- 75 - Loading CSV data into dataframes
- 76 - Loading JSON into dataframes
- 77 - Inspecting dataframes
- 78 - Data quality and data cleansing
- 79 - Using AI for data quality and data cleansing
- 80 - Challenge - Missing data
- 81 - Solution - Missing data
11. Analyzing Data in Databases
- 82 - Relational databases
- 83 - NoSQL databases
- 84 - Extraction, transformation, and loading data into databases
- 85 - Introduction to SQL
- 86 - Creating tables and inserting data
- 87 - Querying data with SQL
- 88 - Joining data with SQL
- 89 - Descriptiive statistics in SQL
- 90 - Generating synthetic data sets for a relational database
- 91 - Generating a star schema, synthetic data, and queries
- 92 - Challenge - Generate a relational data model
- 93 - Solution - Generate a relational data model
12. Introduction to Machine Learning
- 94 - Supervised and unsupervised learning
- 95 - Classification
- 96 - Regression
- 97 - Clustering
- 98 - Machine learning lifecycle
- 99 - Feature engineering
- 100 - Model evaluation
13. Building Machine Learning Models - Classification
- 101 - Simple classification model
- 102 - Handling missing data
- 103 - Comparing multiple algorithms
- 104 - Classification with neural networks
- 105 - Hyperparameter tuning
- 106 - Evaluating feature importance
- 107 - Challenge - Predicting consumer intent
- 108 - Solution - Predicting consumer intent
14. Building Machine Learning Models - Clustering
- 109 - Clustering with k-means
- 110 - Clustering with DBSCAN
- 111 - Clustering with hierarchical clustering
- 112 - Challenge - Customer segmentation
- 113 - Solution - Customer segmentation
15. Open Access ML Datasets
- 114 - Open access ML datasets
16. Network Analysis
- 115 - Introduction to graph theory
- 116 - NetworkX
- 117 - Analyzing a social network
- 118 - Supply chains and network analysis
- 119 - Generating a synthetic supply chain
- 120 - Visualizing a complex supply chain
- 121 - Finding highest betweenness scores
- 122 - Advanced topics in supply chain analysis
- 123 - Challenge - Analyzing a social network
- 124 - Solution - Analyzing a social network
17. Simulations
- 125 - Introduction to simulations
- 126 - Types of simulations
- 127 - Modeling inventory management
- 128 - Agent-based modeling
- 129 - Modeling the spread of infectious diseases
- 130 - Agent-base infectious diseases modeling
- 131 - Challenge - Simulating forest fires
- 132 - Solution - Simulating forest fires
18. Capstone Project
- 133 - Capstone project requirements
- 134 - Capstone project solution
19. Continuing Your AI Learning Journey
- 135 - Next steps and additional resources