Special offers now — see discounted courses.
day
:
hour
:
min
:
sec
See special offers
Scala Essential Training for Data Science (2017)

Scala Essential Training for Data Science (2017)

1h 52mIntermediate2017-09-17

Authors

Dan Sullivan

Dan Sullivan

Enterprise Architect, Big Data Expert

Course details

Discover how to leverage Scala—the popular language that combines object-oriented design with functional programming—in your data science work. In this course, learn about the Scala features most useful to data scientists, including custom functions, parallel processing, and programming Spark with Scala. Dan Sullivan kicks off the course with an introduction for non-Scala programmers. Next, he describes how to use SQL from Scala—a particularly useful concept for data scientists, since they often have to extract data from relational databases. He then covers parallel processing constructs in Scala, sharing techniques that are useful for medium-sized data sets that can be analyzed on a single server with multiple cores.

Dan also focuses on using Scala with Spark, a distributed processing platform. He first describes how to work with Resilient Distributed Datasets (RDDs)—a fundamental Spark data structure—and then explains how to use Scala with Spark DataFrames, a new class of data structure specially designed for analytic processing. He wraps up the course by providing a summary of advantages of using Scala for data science.

Learning objectives
The advantages of Scala for data science
Scala data types
Scala arrays, vectors, and ranges
Parallel processing in Scala
Mapping functions over parallel collections
When and when not to use parallel collections
Using SQL in Scala
Scala and Spark RDDs
Scala and Spark DataFrames
Creating DataFrames

Skills covered

ScalaEssential TrainingProgramming LanguagesOpen SourceSoftware Development

Concepts

0. Introduction

  • 01 - Welcome
  • 02 - What you should know
  • 03 - Using the exercise files

1. Introduction to Scala

  • 04 - The advantages of Scala for data science
  • 05 - Installing Scala
  • 06 - Scala data types
  • 07 - Scala collections
  • 08 - Scala sets Scala arrays, vectors, and ranges
  • 09 - Scala maps
  • 10 - Scala expressions
  • 11 - Scala functions
  • 12 - Scala objects

2. Parallel Processing in Scala

  • 13 - Advantages of parallel collections
  • 14 - Creating parallel collections
  • 15 - Mapping functions over parallel collections
  • 16 - Filtering parallel collections
  • 17 - When and when not to use parallel collections

3. Using SQL in Scala

  • 18 - Installing PostgreSQL
  • 19 - Loading data into PostgreSQL
  • 20 - Connecting to PostgreSQL
  • 21 - Querying with SQL strings
  • 22 - Querying with prepared statements
  • 23 - Summary of SQL in Scala

4. Scala and Spark RDDs

  • 24 - Introduction to Spark
  • 25 - Installing Spark
  • 26 - Getting Started with Spark RDDs
  • 27 - Mapping Functions over RDDs
  • 28 - Statistics over RDDs
  • 29 - Summary of Scala and Spark RDDs

5. Scala and Spark DataFrames

  • 30 - Creating DataFrames
  • 31 - Grouping and filtering on DataFrames
  • 32 - Joining DataFrames
  • 33 - Working with JSON files
  • 34 - Summary of Scala and Spark DataFrames

Conclusion

  • 35 - Review of Scala for data science

About us

LyndaKade is a leading learning platform that helps people learn business, software, technology, and creative skills to achieve personal and professional goals.

Phone numberAparat ChannelTelegram SupportTelegram ChannelInstagram Page

All rights to this site belong to LyndaKade.

Terms of Service|Privacy Policy

نماد الکترونیک enamad در صورت اتصال با آی‌پی داخل کشور، نمایش داده خواهد شد.
logo-samandehi - لوگو ساماندهی
Zarinpal
Zibal