Python 3 Text Processing with NLTK 3 Cookbook

Book Image

Python 3 Text Processing with NLTK 3 Cookbook

By : Jacob Perkins

Book Image

Python 3 Text Processing with NLTK 3 Cookbook

By: Jacob Perkins

Overview of this book

Python 3 Text Processing with NLTK 3 Cookbook

Python 3 Text Processing with NLTK 3 Cookbook

Credits

About the Author

About the Author

About the Reviewers

About the Reviewers

www.PacktPub.com

www.PacktPub.com

Preface

Free Chapter

Tokenizing Text and WordNet Basics

Tokenizing Text and WordNet Basics

Tokenizing text into sentences

Tokenizing sentences into words

Tokenizing sentences using regular expressions

Training a sentence tokenizer

Filtering stopwords in a tokenized sentence

Looking up Synsets for a word in WordNet

Looking up lemmas and synonyms in WordNet

Calculating WordNet Synset similarity

Discovering word collocations

Replacing and Correcting Words

Replacing and Correcting Words

Lemmatizing words with WordNet

Replacing words matching regular expressions

Removing repeating characters

Spelling correction with Enchant

Replacing synonyms

Replacing negations with antonyms

Creating Custom Corpora

Creating Custom Corpora

Setting up a custom corpus

Creating a wordlist corpus

Creating a part-of-speech tagged word corpus

Creating a chunked phrase corpus

Creating a categorized text corpus

Creating a categorized chunk corpus reader

Lazy corpus loading

Creating a custom corpus view

Creating a MongoDB-backed corpus reader

Corpus editing with file locking

Part-of-speech Tagging

Part-of-speech Tagging

Default tagging

Training a unigram part-of-speech tagger

Combining taggers with backoff tagging

Training and combining ngram taggers

Creating a model of likely word tags

Tagging with regular expressions

Training a Brill tagger

Training the TnT tagger

Using WordNet for tagging

Tagging proper names

Classifier-based tagging

Training a tagger with NLTK-Trainer

Extracting Chunks

Extracting Chunks

Chunking and chinking with regular expressions

Merging and splitting chunks with regular expressions

Expanding and removing chunks with regular expressions

Partial parsing with regular expressions

Training a tagger-based chunker

Classification-based chunking

Extracting named entities

Extracting proper noun chunks

Extracting location chunks

Training a named entity chunker

Training a chunker with NLTK-Trainer

Transforming Chunks and Trees

Transforming Chunks and Trees

Filtering insignificant words from a sentence

Correcting verb forms

Swapping verb phrases

Swapping noun cardinals

Swapping infinitive phrases

Singularizing plural nouns

Chaining chunk transformations

Converting a chunk tree to text

Flattening a deep tree

Creating a shallow tree

Converting tree labels

Text Classification

Text Classification

Bag of words feature extraction

Training a Naive Bayes classifier

Training a decision tree classifier

Training a maximum entropy classifier

Training scikit-learn classifiers

Measuring precision and recall of a classifier

Calculating high information words

Combining classifiers with voting

Classifying with multiple binary classifiers

Training a classifier with NLTK-Trainer

Distributed Processing and Handling Large Datasets

Distributed Processing and Handling Large Datasets

Distributed tagging with execnet

Distributed chunking with execnet

Parallel list processing with execnet

Storing a frequency distribution in Redis

Storing a conditional frequency distribution in Redis

Storing an ordered dictionary in Redis

Distributed word scoring with Redis and execnet

Parsing Specific Data Types

Parsing Specific Data Types

Parsing dates and times with dateutil

Timezone lookup and conversion

Extracting URLs from HTML with lxml

Cleaning and stripping HTML

Converting HTML entities with BeautifulSoup

Detecting and converting character encodings

Penn Treebank Part-of-speech Tags

Penn Treebank Part-of-speech Tags

Index

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Default tagging

Default tagging provides a baseline for part-of-speech tagging. It simply assigns the same part-of-speech tag to every token. We do this using the DefaultTagger class. This tagger is useful as a last-resort tagger, and provides a baseline to measure accuracy improvements.

Getting ready

We're going to use the treebank corpus for most of this chapter because it's a common standard and is quick to load and test. But everything we do should apply equally well to brown, conll2000, and any other part-of-speech tagged corpus.

How to do it...

The DefaultTagger class takes a single argument, the tag you want to apply. We'll give it NN, which is the tag for a singular noun. DefaultTagger is most useful when you choose the most common part-of-speech tag. Since nouns tend to be the most common types of words, a noun tag is recommended.

>>> from nltk.tag import DefaultTagger
>>> tagger = DefaultTagger('NN')
>>> tagger.tag(['Hello', 'World'])
[('Hello', 'NN'), ('World...