Reinforcement Learning Algorithms with Python

By : Andrea Lonza

Reinforcement Learning Algorithms with Python

By: Andrea Lonza

Overview of this book

Reinforcement Learning (RL) is a popular and promising branch of AI that involves making smarter models and agents that can automatically determine ideal behavior based on changing requirements. This book will help you master RL algorithms and understand their implementation as you build self-learning agents. Starting with an introduction to the tools, libraries, and setup needed to work in the RL environment, this book covers the building blocks of RL and delves into value-based methods, such as the application of Q-learning and SARSA algorithms. You'll learn how to use a combination of Q-learning and neural networks to solve complex problems. Furthermore, you'll study the policy gradient methods, TRPO, and PPO, to improve performance and stability, before moving on to the DDPG and TD3 deterministic algorithms. This book also covers how imitation learning techniques work and how Dagger can teach an agent to drive. You'll discover evolutionary strategies and black-box optimization techniques, and see how they can improve RL algorithms. Finally, you'll get to grips with exploration approaches, such as UCB and UCB1, and develop a meta-algorithm called ESBAS. By the end of the book, you'll have worked with key RL algorithms to overcome challenges in real-world applications, and be part of the RL research community.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

Section 1: Algorithms and Environments

The Landscape of Reinforcement Learning

An introduction to RL

Summary

Implementing RL Cycle and OpenAI Gym

Setting up the environment

OpenAI Gym and RL cycles

Development of ML models using TensorFlow

Introducing TensorBoard

Types of RL environments

Summary

Questions

Further reading

Solving Problems with Dynamic Programming

MDP

Categorizing RL algorithms

Dynamic programming

Summary

Questions

Further reading

Section 2: Model-Free RL Algorithms

Q-Learning and SARSA Applications

Learning without a model

TD learning

SARSA

Applying SARSA to Taxi-v2

Q-learning

Applying Q-learning to Taxi-v2

Summary

Questions

Deep Q-Network

Deep neural networks and Q-learning

DQN

Summary

Learning Stochastic and PG Optimization

Policy gradient methods

Understanding the REINFORCE algorithm

REINFORCE with baseline

Learning the AC algorithm

Summary

Questions

Further reading

TRPO and PPO Implementation

Roboschool

Natural policy gradient

Trust region policy optimization

Proximal Policy Optimization

Summary

Questions

Further reading

DDPG and TD3 Applications

Combining policy gradient optimization with Q-learning

Deep deterministic policy gradient

Twin delayed deep deterministic policy gradient (TD3)

Summary

Questions

Further reading

Section 3: Beyond Model-Free Algorithms and Improvements

Model-Based RL

Model-based methods

Combining model-based with model-free learning

ME-TRPO applied to an inverted pendulum

Summary

Questions

Further reading

Imitation Learning with the DAgger Algorithm

Technical requirements

The imitation approach

Playing Flappy Bird

Understanding the dataset aggregation algorithm

IRL

Summary

Questions

Further reading

Understanding Black-Box Optimization Algorithms

Beyond RL

The core of EAs

Scalable evolution strategies

Applying scalable ES to LunarLander

Summary

Questions

Further reading

Developing the ESBAS Algorithm

Exploration versus exploitation

Approaches to exploration

Epochal stochastic bandit algorithm selection

Summary

Questions

Further reading

Practical Implementation for Resolving RL Challenges

Best practices of deep RL

Challenges in deep RL

Advanced techniques

RL in the real world

Future of RL and its impact on society

Summary

Questions

Further reading

Assessments

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

To get the most out of this book

Working knowledge of Python is necessary. Knowledge of RL and the various tools used for it will also be beneficial.

Download the example code files

You can download the example code files for this book from your account at www.packt.com. If you purchased this book elsewhere, you can visit www.packtpub.com/support and register to have the files emailed directly to you.

You can download the code files by following these steps:

Log in or register at www.packt.com.
Select the Support tab.
Click on Code Downloads.
Enter the name of the book in the Search box and follow the onscreen instructions.

Once the file is downloaded, please make sure that you unzip or extract the folder using the latest version of:

WinRAR/7-Zip for Windows
Zipeg/iZip/UnRarX for Mac
7-Zip/PeaZip for Linux

The code bundle for the book is also hosted on GitHub at https://github.com/PacktPublishing/Reinforcement-Learning-Algorithms-with-Python. In case there's an update to the code, it will be updated on the existing GitHub repository.

We also have other code bundles from our rich catalog of books and videos available at https://github.com/PacktPublishing/. Check them out!

Download the color images

We also provide a PDF file that has color images of the screenshots/diagrams used in this book. You can download it here: http://www.packtpub.com/sites/default/files/downloads/9781789131116_ColorImages.pdf .

Conventions used

There are a number of text conventions used throughout this book.

CodeInText: Indicates code words in text, database table names, folder names, filenames, file extensions, pathnames, dummy URLs, user input, and Twitter handles. Here is an example: "In this book, we use Python 3.7, but all versions above 3.5 should work. We also assume that you've already installed numpy and matplotlib."

A block of code is set as follows:

import gym

# create the environment 
env = gym.make("CartPole-v1")
# reset the environment before starting
env.reset()

# loop 10 times
for i in range(10):
    # take a random action
    env.step(env.action_space.sample())
    # render the game
   env.render()

# close the environment
env.close()

Any command-line input or output is written as follows:

$ git clone https://github.com/pybox2d/pybox2d
$ cd pybox2d
$ pip install -e .

Bold: Indicates a new term, an important word, or words that you see onscreen. For example, words in menus or dialog boxes appear in the text like this. Here is an example: "In reinforcement learning (RL), the algorithm is called the agent, and it learns from the data provided by an environment."

Warnings or important notes appear like this.

Tips and tricks appear like this.

Reinforcement Learning Algorithms with Python

By : Andrea Lonza

Reinforcement Learning Algorithms with Python

By: Andrea Lonza

Overview of this book

Related Content you might be interested in

Current Title:

Reinforcement Learning Algorithms with Python

Mastering Reinforcement Learning with Python

TensorFlow Reinforcement Learning Quick Start Guide

Hands-On Intelligent Agents with OpenAI Gym