NLTK Essentials

NLTK Essentials

By : Nitin Hardeniya

Buy this Book

NLTK Essentials

By: Nitin Hardeniya

Buy this Book

Overview of this book

Natural Language Processing (NLP) is the field of artificial intelligence and computational linguistics that deals with the interactions between computers and human languages. With the instances of human-computer interaction increasing, it’s becoming imperative for computers to comprehend all major natural languages. Natural Language Toolkit (NLTK) is one such powerful and robust tool. You start with an introduction to get the gist of how to build systems around NLP. We then move on to explore data science-related tasks, following which you will learn how to create a customized tokenizer and parser from scratch. Throughout, we delve into the essential concepts of NLP while gaining practical insights into various open source tools and libraries available in Python for NLP. You will then learn how to analyze social media sites to discover trending topics and perform sentiment analysis. Finally, you will see tools which will help you deal with large scale text. By the end of this book, you will be confident about NLP and data science concepts and know how to apply them in your day-to-day work.

NLTK Essentials

Credits

About the Author

About the Reviewers

www.PacktPub.com

Preface

Free Chapter

Introduction to Natural Language Processing

Why learn NLP?

Let's start playing with Python!

Diving into NLTK

Your turn

Summary

Text Wrangling and Cleansing

What is text wrangling?

Summary

Part of Speech Tagging

What is Part of speech tagging

Named Entity Recognition (NER)

Your Turn

Summary

Parsing Structure in Text

Shallow versus deep parsing

The two approaches in parsing

Why we need parsing

Different types of parsers

Dependency parsing

Chunking

Information extraction

Summary

NLP Applications

Building your first NLP application

Other NLP applications

Summary

Text Classification

Machine learning

Text classification

Sampling

The Random forest algorithm

Text clustering

Topic modeling in text

References

Summary

Web Crawling

Web crawlers

Writing your first crawler

Summary

Using NLTK with Other Python Libraries

NumPy

SciPy

pandas

matplotlib

External references

Summary

Social Media Mining in Python

Data collection

Data extraction

Geovisualization

Summary

Text Mining at Scale

Different ways of using Python on Hadoop

NLTK on Hadoop

Scikit-learn on Hadoop

PySpark

Summary

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Lemmatization

Lemmatization is a more methodical way of converting all the grammatical/inflected forms of the root of the word. Lemmatization uses context and part of speech to determine the inflected form of the word and applies different normalization rules for each part of speech to get the root word (lemma):

>>>from nltk.stem import WordNetLemmatizer
>>>wlem = WordNetLemmatizer()
>>>wlem.lemmatize("ate")
eat

Here, WordNetLemmatizer is using wordnet, which takes a word and searches wordnet, a semantic dictionary. It also uses a morph analysis to cut to the root and search for the specific lemma (variation of the word). Hence, in our example it is possible to get eat for the given variation ate, which was never possible with stemming.

Can you explain what the difference is between Stemming and lemmatization?
Can you come up with a Porter stemmer (Rule-based) for your native language?
Why would it be harder to implement a stemmer for languages like Chinese?

NLTK Essentials

By : Nitin Hardeniya

NLTK Essentials

By: Nitin Hardeniya

Overview of this book

Related Content you might be interested in

Current Title:

NLTK Essentials

Lemmatization