Introduction to NLP Using R
2h 32mAdvanced2023-05-19
Authors

Mark Niemann-Ross
Technologist experienced in hardware, software, and science fiction
Course details
Natural language processing (NLP) is one of the most important components of artificial intelligence. It allows you to process, analyze, and understand large amounts of data in the form of natural language. In this course, instructor Mark Niemann-Ross shows you how to get started implementing NLP algorithms using R, the popular programming language for statistical computing and graphics.
Explore the basics of manipulating matrices and producing statistics, both of which are core to successful NLP. Learn how to use tools and text-mining frameworks such as tm, quanteda, and tidytext, as well as work with corpora, sources, and other types of NLP document metadata. Mark covers the best practices for preprocessing text in preparation for NLP, creating structured data, applying statistics to text, performing sentiment analysis, visualizing datasets, and more.
Explore the basics of manipulating matrices and producing statistics, both of which are core to successful NLP. Learn how to use tools and text-mining frameworks such as tm, quanteda, and tidytext, as well as work with corpora, sources, and other types of NLP document metadata. Mark covers the best practices for preprocessing text in preparation for NLP, creating structured data, applying statistics to text, performing sentiment analysis, visualizing datasets, and more.
Skills covered
RStatisticsNatural Language Processing (NLP)Artificial Intelligence (AI)Programming LanguagesData ScienceOpen SourceSoftware DevelopmentOne-Off
Concepts
0. Introduction
- 01 - Welcome to natural language processing with R
- 02 - Skills and tools you ll need to be successful in this course
1. Up and Running with tm
- 03 - What is tm and why do you need it
- 04 - tm documentation walk-through
- 05 - Real-world NLP with tm
- 06 - Real-world NLP with quanteda
- 07 - Real-world NLP with tidytext
2. Corpora and Sources
- 08 - Understanding corpora and sources
- 09 - Examining corpora
- 10 - Examining sources
- 11 - Custom sources
- 12 - Combining and subsetting corpora
3. Working with NLP Metadata
- 13 - Working with document metadata
- 14 - Make useful metadata
- 15 - Finding and filtering based on metadata
4. Preprocessing Text in Preparation for NLP
- 16 - Transformations
- 17 - Stop words
- 18 - Stemming
- 19 - Lemmatization
- 20 - Tokenization
- 21 - Ngrams
- 22 - Part of speech tagging
5. Create Structured Data
- 23 - Understanding the document-term matrix
- 24 - Create the document-term matrix
- 25 - Weighting the document-term matrix
- 26 - Focus the document-term matrix
6. Apply Statistics to Text
- 27 - Word and document frequency
- 28 - Hierarchical clustering
- 29 - Associated terms
7. Sentiment Analysis
- 30 - What is sentiment analysis
- 31 - Real-world example of sentiment analysis
- 32 - Sentiment datasets
- 33 - Sentiment tools
8. Visualizing Natural Language Processing
- 34 - Plotting text mining
- 35 - Plotting Zipf s and Heap s Law
- 36 - Word clouds
Conclusion
- 37 - Your next steps in NLP