Book Image

The Reinforcement Learning Workshop

By : Alessandro Palmas, Emanuele Ghelfi, Dr. Alexandra Galina Petre, Mayur Kulkarni, Anand N.S., Quan Nguyen, Aritra Sen, Anthony So, Saikat Basak

Book Image

The Reinforcement Learning Workshop

By: Alessandro Palmas, Emanuele Ghelfi, Dr. Alexandra Galina Petre, Mayur Kulkarni, Anand N.S., Quan Nguyen, Aritra Sen, Anthony So, Saikat Basak

Overview of this book

Various intelligent applications such as video games, inventory management software, warehouse robots, and translation tools use reinforcement learning (RL) to make decisions and perform actions that maximize the probability of the desired outcome. This book will help you to get to grips with the techniques and the algorithms for implementing RL in your machine learning models. Starting with an introduction to RL, youÔÇÖll be guided through different RL environments and frameworks. YouÔÇÖll learn how to implement your own custom environments and use OpenAI baselines to run RL algorithms. Once youÔÇÖve explored classic RL techniques such as Dynamic Programming, Monte Carlo, and TD Learning, youÔÇÖll understand when to apply the different deep learning methods in RL and advance to deep Q-learning. The book will even help you understand the different stages of machine-based problem-solving by using DARQN on a popular video game Breakout. Finally, youÔÇÖll find out when to use a policy-based method to tackle an RL problem. By the end of The Reinforcement Learning Workshop, youÔÇÖll be equipped with the knowledge and skills needed to solve challenging problems using reinforcement learning.

Preface

1. Introduction to Reinforcement Learning

1. Introduction to Reinforcement Learning

Learning Paradigms

Fundamentals of Reinforcement Learning

Reinforcement Learning Frameworks

Applications of Reinforcement Learning

Free Chapter

2. Markov Decision Processes and Bellman Equations

2. Markov Decision Processes and Bellman Equations

Markov Processes

3. Deep Learning in Practice with TensorFlow 2

3. Deep Learning in Practice with TensorFlow 2

An Introduction to TensorFlow and Keras

How to Implement a Neural Network Using TensorFlow

Simple Regression Using TensorFlow

Simple Classification Using TensorFlow

TensorBoard – How to Visualize Data Using TensorBoard

4. Getting Started with OpenAI and TensorFlow for Reinforcement Learning

4. Getting Started with OpenAI and TensorFlow for Reinforcement Learning

OpenAI Universe – Complex Environment

TensorFlow for Reinforcement Learning

OpenAI Baselines

Training an RL Agent to Solve a Classic Control Problem

5. Dynamic Programming

5. Dynamic Programming

Solving Dynamic Programming Problems

Identifying Dynamic Programming Problems

Dynamic Programming in RL

6. Monte Carlo Methods

6. Monte Carlo Methods

The Workings of Monte Carlo Methods

Understanding Monte Carlo with Blackjack

Types of Monte Carlo Methods

Exploration versus Exploitation Trade-Off

Importance Sampling

Solving Frozen Lake Using Monte Carlo

7. Temporal Difference Learning

7. Temporal Difference Learning

Introduction to TD Learning

TD(0) – SARSA and Q-Learning

N-Step TD and TD(λ) Algorithms

The Relationship between DP, Monte-Carlo, and TD Learning

8. The Multi-Armed Bandit Problem

8. The Multi-Armed Bandit Problem

Formulation of the MAB Problem

The Python Interface

The Greedy Algorithm

The Explore-then-Commit Algorithm

The ε-Greedy Algorithm

The UCB algorithm

Thompson Sampling

Contextual Bandits

9. What Is Deep Q-Learning?

9. What Is Deep Q-Learning?

Basics of Deep Learning

Basics of PyTorch

The Action-Value Function (Q Value Function)

Deep Q Learning

Challenges in DQN

10. Playing an Atari Game with Deep Recurrent Q-Networks

10. Playing an Atari Game with Deep Recurrent Q-Networks

Understanding the Breakout Environment

CNNs in TensorFlow

Combining a DQN with a CNN

RNNs in TensorFlow

Building a DRQN

Introduction to the Attention Mechanism and DARQN

11. Policy-Based Methods for Reinforcement Learning

11. Policy-Based Methods for Reinforcement Learning

Policy Gradients

Deep Deterministic Policy Gradients

Improving Policy Gradients

12. Evolutionary Strategies for RL

12. Evolutionary Strategies for RL

Problems with Gradient-Based Methods

Introduction to Genetic Algorithms

Appendix

1. Introduction to Reinforcement Learning

2. Markov Decision Processes and Bellman Equations

3. Deep Learning in Practice with TensorFlow 2

4. Getting started with OpenAI and TensorFlow for Reinforcement Learning

5. Dynamic Programming

6. Monte Carlo Methods

7. Temporal Difference Learning

8. The Multi-Armed Bandit Problem

9. What Is Deep Q-Learning?

10. Playing an Atari Game with Deep Recurrent Q-Networks

11. Policy-Based Methods for Reinforcement Learning

12. Evolutionary Strategies for RL

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

12. Evolutionary Strategies for RL

Activity 12.01: Cart-Pole Activity

Import the required packages as follows:

import gym 
import numpy as np 
import math 
import tensorflow as tf
from matplotlib import pyplot as plt
from random import randint
from statistics import median, mean

Initialize the environment and the state and action space shapes:

env = gym.make('CartPole-v0')
no_states = env.observation_space.shape[0]
no_actions = env.action_space.n

Create a function to generate randomly selected initial network parameters:

def initial(run_test):
    #initialize arrays
    i_w = []
    i_b = []
    h_w = []
    o_w = []
    no_input_nodes = 8
    no_hidden_nodes = 4
    
    for r in range(run_test):
        input_weight = np.random.rand(no_states...