Analyzing Big Data with Hive
1h 53mIntermediate2017-01-20
Authors

Ben Sullins
Data Geek, Tech Consultant
Course details
Businesses thrive by making informed decisions that target the needs of their customers and users. To make such strategic decisions, they rely on data. Hive is a tool of choice for many data scientists because it allows them to work with SQL, a familiar syntax, to derive insights from Hadoop, reflecting the information that businesses seek to plan effectively.
This course shows how to use Hive to process data. Instructor Ben Sullins starts by showing you how to structure and optimize your data. Next, he explains how to get Hue, the Hadoop user interface, to leverage HiveQL when analyzing data. Using the newly configured option, he then demonstrates how to load data, create aggregate tables for fast query access, and run advanced analytics. He also takes you through managing tables and putting functions to use. This course is designed to help you find new ways to work with datasets so you can answer the tough data science questions that come your way.
Learning objectives
Defining data structures in Hive
Selecting data
Joining tables
Manipulating data
Filtering results
Aggregating data
Using built-in aggregate functions
Mastering built-in table-generating functions
Using CUBE and ROLLUP
Using clauses: WHERE and HAVING
Using LIKE, JOIN, and SEMI JOIN
Using functions: String, math, date, and conditional
This course shows how to use Hive to process data. Instructor Ben Sullins starts by showing you how to structure and optimize your data. Next, he explains how to get Hue, the Hadoop user interface, to leverage HiveQL when analyzing data. Using the newly configured option, he then demonstrates how to load data, create aggregate tables for fast query access, and run advanced analytics. He also takes you through managing tables and putting functions to use. This course is designed to help you find new ways to work with datasets so you can answer the tough data science questions that come your way.
Learning objectives
Defining data structures in Hive
Selecting data
Joining tables
Manipulating data
Filtering results
Aggregating data
Using built-in aggregate functions
Mastering built-in table-generating functions
Using CUBE and ROLLUP
Using clauses: WHERE and HAVING
Using LIKE, JOIN, and SEMI JOIN
Using functions: String, math, date, and conditional
Skills covered
HiveData EngineeringProjectData Science
Concepts
0. Introduction
- 01 - Welcome
- 02 - What you should know before watching this course
- 03 - Using the exercise files
1. Hive Concepts and Setup
- 04 - Why use Hive
- 05 - How Hive works
- 06 - Setting up our demo environment
2. Working with Data in Hive
- 07 - Understanding table structures in Hive
- 08 - Creating tables in Hive
- 09 - Handling CSV files in Hive
- 10 - Partitioning tables
3. Retrieving Data from Hive
- 11 - Simple SELECT statement
- 12 - Retrieving data from complex structures
4. Aggregating Data
- 13 - Simple aggregations
- 14 - Enhanced aggregations with grouping sets
- 15 - Using CUBE and ROLLUP
5. Filtering Results
- 16 - Simple filter with the WHERE clause
- 17 - Filtering aggregates with HAVING clause
- 18 - Finding similar values with LIKE
6. Joining Tables
- 19 - Combining tables with JOIN
- 20 - When to use SEMI JOIN
- 21 - Joining multiple tables together
7. Manipulating Data
- 22 - Types of data manipulation functions
- 23 - String functions
- 24 - Math functions
- 25 - Date functions
- 26 - Conditional functions
Conclusion
- 27 - Next steps