Python Web Scraping Cookbook

By : Michael Heydt

Python Web Scraping Cookbook

By: Michael Heydt

Overview of this book

Python Web Scraping Cookbook is a solution-focused book that will teach you techniques to develop high-performance scrapers and deal with crawlers, sitemaps, forms automation, Ajax-based sites, caches, and more. You'll explore a number of real-world scenarios where every part of the development/product life cycle will be fully covered. You will not only develop the skills needed to design and develop reliable performance data flows, but also deploy your codebase to AWS. If you are involved in software engineering, product development, or data mining (or are interested in building data-driven products), you will find this book useful as each recipe has a clear purpose and objective. Right from extracting data from the websites to writing a sophisticated web crawler, the book's independent recipes will be a godsend. This book covers Python libraries, requests, and BeautifulSoup. You will learn about crawling, web spidering, working with Ajax websites, paginated items, and more. You will also learn to tackle problems such as 403 errors, working with proxy, scraping images, and LXML. By the end of this book, you will be able to scrape websites more efficiently and able to deploy and operate your scraper in the cloud.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

Getting Started with Scraping

Introduction

Setting up a Python development environment

Scraping Python.org with Requests and Beautiful Soup

Scraping Python.org in urllib3 and Beautiful Soup

Scraping Python.org with Scrapy

Scraping Python.org with Selenium and PhantomJS

Data Acquisition and Extraction

Introduction

How to parse websites and navigate the DOM using BeautifulSoup

Searching the DOM with Beautiful Soup's find methods

Querying the DOM with XPath and lxml

Querying data with XPath and CSS selectors

Using Scrapy selectors

Loading data in unicode / UTF-8

Processing Data

Introduction

Working with CSV and JSON data

Storing data using AWS S3

Storing data using MySQL

Storing data using PostgreSQL

Storing data in Elasticsearch

How to build robust ETL pipelines with AWS SQS

Working with Images, Audio, and other Assets

Introduction

Downloading media content from the web

Parsing a URL with urllib to get the filename

Determining the type of content for a URL

Determining the file extension from a content type

Downloading and saving images to the local file system

Downloading and saving images to S3

Generating thumbnails for images

Taking a screenshot of a website

Taking a screenshot of a website with an external service

Performing OCR on an image with pytesseract

Creating a Video Thumbnail

Ripping an MP4 video to an MP3

Scraping - Code of Conduct

Introduction

Scraping legality and scraping politely

Respecting robots.txt

Crawling using the sitemap

Crawling with delays

Using identifiable user agents

Setting the number of concurrent requests per domain

Using auto throttling

Using an HTTP cache for development

Scraping Challenges and Solutions

Introduction

Retrying failed page downloads

Supporting page redirects

Waiting for content to be available in Selenium

Limiting crawling to a single domain

Processing infinitely scrolling pages

Controlling the depth of a crawl

Controlling the length of a crawl

Handling paginated websites

Handling forms and forms-based authorization

Handling basic authorization

Preventing bans by scraping via proxies

Randomizing user agents

Caching responses

Text Wrangling and Analysis

Introduction

Installing NLTK

Performing sentence splitting

Performing tokenization

Performing stemming

Performing lemmatization

Determining and removing stop words

Calculating the frequency distributions of words

Identifying and removing rare words

Removing punctuation marks

Piecing together n-grams

Scraping a job listing from StackOverflow

Reading and cleaning the description in the job listing

Searching, Mining and Visualizing Data

Introduction

Geocoding an IP address

How to collect IP addresses of Wikipedia edits

Visualizing contributor location frequency on Wikipedia

Creating a word cloud from a StackOverflow job listing

Crawling links on Wikipedia

Visualizing page relationships on Wikipedia

Calculating degrees of separation

Creating a Simple Data API

Introduction

Creating a REST API with Flask-RESTful

Integrating the REST API with scraping code

Adding an API to find the skills for a job listing

Storing data in Elasticsearch as the result of a scraping request

Checking Elasticsearch for a listing before scraping

Creating Scraper Microservices with Docker

Introduction

Installing Docker

Installing a RabbitMQ container from Docker Hub

Running a Docker container (RabbitMQ)

Creating and running an Elasticsearch container

Stopping/restarting a container and removing the image

Creating a generic microservice with Nameko

Creating a scraping microservice

Creating a scraper container

Creating an API container

Composing and running the scraper locally with docker-compose

Making the Scraper as a Service Real

Introduction

Creating and configuring an Elastic Cloud trial account

Accessing the Elastic Cloud cluster with curl

Connecting to the Elastic Cloud cluster with Python

Performing an Elasticsearch query with the Python API

Using Elasticsearch to query for jobs with specific skills

Modifying the API to search for jobs by skill

Storing configuration in the environment

Creating an AWS IAM user and a key pair for ECS

Configuring Docker to authenticate with ECR

Pushing containers into ECR

Creating an ECS cluster

Creating a task to run our containers

Starting and accessing the containers in AWS

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Scraping Python.org with Requests and Beautiful Soup

In this recipe we will install Requests and Beautiful Soup and scrape some content from www.python.org. We'll install both of the libraries and get some basic familiarity with them. We'll come back to them both in subsequent chapters and dive deeper into each.

Getting ready...

In this recipe, we will scrape the upcoming Python events from https://www.python.org/events/pythonevents. The following is an an example of The Python.org Events Page (it changes frequently, so your experience will differ):

We will need to ensure that Requests and Beautiful Soup are installed. We can do that with the following:

pywscb $ pip install requests
Downloading/unpacking requests
 Downloading requests-2.18.4-py2.py3-none-any.whl (88kB): 88kB downloaded
Downloading/unpacking certifi>=2017.4.17 (from requests)
 Downloading certifi-2018.1.18-py2.py3-none-any.whl (151kB): 151kB downloaded
Downloading/unpacking idna>=2.5,<2.7 (from requests)
 Downloading idna-2.6-py2.py3-none-any.whl (56kB): 56kB downloaded
Downloading/unpacking chardet>=3.0.2,<3.1.0 (from requests)
 Downloading chardet-3.0.4-py2.py3-none-any.whl (133kB): 133kB downloaded
Downloading/unpacking urllib3>=1.21.1,<1.23 (from requests)
 Downloading urllib3-1.22-py2.py3-none-any.whl (132kB): 132kB downloaded
Installing collected packages: requests, certifi, idna, chardet, urllib3
Successfully installed requests certifi idna chardet urllib3
Cleaning up...
pywscb $ pip install bs4
Downloading/unpacking bs4
 Downloading bs4-0.0.1.tar.gz
 Running setup.py (path:/Users/michaelheydt/pywscb/env/build/bs4/setup.py) egg_info for package bs4

How to do it...

Now let's go and learn to scrape a couple events. For this recipe we will start by using interactive python.

Start it with the ipython command:

$ ipython
Python 3.6.1 |Anaconda custom (x86_64)| (default, Mar 22 2017, 19:25:17)
Type "copyright", "credits" or "license" for more information.
IPython 5.1.0 -- An enhanced Interactive Python.
? -> Introduction and overview of IPython's features.
%quickref -> Quick reference.
help -> Python's own help system.
object? -> Details about 'object', use 'object??' for extra details.
In [1]:

Next we import Requests

In [1]: import requests

We now use requests to make a GET HTTP request for the following url: https://www.python.org/events/python-events/ by making a GET request:

In [2]: url = 'https://www.python.org/events/python-events/'
In [3]: req = requests.get(url)

That downloaded the page content but it is stored in our requests object req. We can retrieve the content using the .text property. This prints the first 200 characters.

req.text[:200]
Out[4]: '<!doctype html>\n<!--[if lt IE 7]> <html class="no-js ie6 lt-ie7 lt-ie8 lt-ie9"> <![endif]-->\n<!--[if IE 7]> <html class="no-js ie7 lt-ie8 lt-ie9"> <![endif]-->\n<!--[if IE 8]> <h'

We now have the raw HTML of the page. We can now use beautiful soup to parse the HTML and retrieve the event data.

First import Beautiful Soup

In [5]: from bs4 import BeautifulSoup

Now we create a BeautifulSoup object and pass it the HTML.

In [6]: soup = BeautifulSoup(req.text, 'lxml')

Now we tell Beautiful Soup to find the main <ul> tag for the recent events, and then to get all the <li> tags below it.

In [7]: events = soup.find('ul', {'class': 'list-recent-events'}).findAll('li')

And finally we can loop through each of the <li> elements, extracting the event details, and print each to the console:

In [13]: for event in events:
 ...: event_details = dict()
 ...: event_details['name'] = event_details['name'] = event.find('h3').find("a").text
 ...: event_details['location'] = event.find('span', {'class', 'event-location'}).text
 ...: event_details['time'] = event.find('time').text
 ...: print(event_details)
 ...:
{'name': 'PyCascades 2018', 'location': 'Granville Island Stage, 1585 Johnston St, Vancouver, BC V6H 3R9, Canada', 'time': '22 Jan. – 24 Jan. 2018'}
{'name': 'PyCon Cameroon 2018', 'location': 'Limbe, Cameroon', 'time': '24 Jan. – 29 Jan. 2018'}
{'name': 'FOSDEM 2018', 'location': 'ULB Campus du Solbosch, Av. F. D. Roosevelt 50, 1050 Bruxelles, Belgium', 'time': '03 Feb. – 05 Feb. 2018'}
{'name': 'PyCon Pune 2018', 'location': 'Pune, India', 'time': '08 Feb. – 12 Feb. 2018'}
{'name': 'PyCon Colombia 2018', 'location': 'Medellin, Colombia', 'time': '09 Feb. – 12 Feb. 2018'}
{'name': 'PyTennessee 2018', 'location': 'Nashville, TN, USA', 'time': '10 Feb. – 12 Feb. 2018'}

This entire example is available in the 01/01_events_with_requests.py script file. The following is its content and it pulls together all of what we just did step by step:

import requests
from bs4 import BeautifulSoup

def get_upcoming_events(url):
    req = requests.get(url)

    soup = BeautifulSoup(req.text, 'lxml')

    events = soup.find('ul', {'class': 'list-recent-events'}).findAll('li')

    for event in events:
        event_details = dict()
        event_details['name'] = event.find('h3').find("a").text
        event_details['location'] = event.find('span', {'class', 'event-location'}).text
        event_details['time'] = event.find('time').text
        print(event_details)

get_upcoming_events('https://www.python.org/events/python-events/')

You can run this using the following command from the terminal:

$ python 01_events_with_requests.py
{'name': 'PyCascades 2018', 'location': 'Granville Island Stage, 1585 Johnston St, Vancouver, BC V6H 3R9, Canada', 'time': '22 Jan. – 24 Jan. 2018'}
{'name': 'PyCon Cameroon 2018', 'location': 'Limbe, Cameroon', 'time': '24 Jan. – 29 Jan. 2018'}
{'name': 'FOSDEM 2018', 'location': 'ULB Campus du Solbosch, Av. F. D. Roosevelt 50, 1050 Bruxelles, Belgium', 'time': '03 Feb. – 05 Feb. 2018'}
{'name': 'PyCon Pune 2018', 'location': 'Pune, India', 'time': '08 Feb. – 12 Feb. 2018'}
{'name': 'PyCon Colombia 2018', 'location': 'Medellin, Colombia', 'time': '09 Feb. – 12 Feb. 2018'}
{'name': 'PyTennessee 2018', 'location': 'Nashville, TN, USA', 'time': '10 Feb. – 12 Feb. 2018'}

How it works...

We will dive into details of both Requests and Beautiful Soup in the next chapter, but for now let's just summarize a few key points about how this works. The following important points about Requests:

Requests is used to execute HTTP requests. We used it to make a GET verb request of the URL for the events page.
The Requests object holds the results of the request. This is not only the page content, but also many other items about the result such as HTTP status codes and headers.
Requests is used only to get the page, it does not do an parsing.

We use Beautiful Soup to do the parsing of the HTML and also the finding of content within the HTML.

To understand how this worked, the content of the page has the following HTML to start the Upcoming Events section:

We used the power of Beautiful Soup to:

Find the <ul> element representing the section, which is found by looking for a <ul> with the a class attribute that has a value of list-recent-events.
From that object, we find all the <li> elements.

Each of these <li> tags represent a different event. We iterate over each of those making a dictionary from the event data found in child HTML tags:

The name is extracted from the <a> tag that is a child of the <h3> tag
The location is the text content of the <span> with a class of event-location
And the time is extracted from the datetime attribute of the <time> tag.

Python Web Scraping Cookbook

By : Michael Heydt

Python Web Scraping Cookbook

By: Michael Heydt

Overview of this book

Related Content you might be interested in

Current Title:

Python Web Scraping Cookbook

Python Web Scraping

Hands-On Web Scraping with Python