Hands-On Reinforcement Learning for Games

By : Micheal Lanham

Hands-On Reinforcement Learning for Games

By: Micheal Lanham

Overview of this book

With the increased presence of AI in the gaming industry, developers are challenged to create highly responsive and adaptive games by integrating artificial intelligence into their projects. This book is your guide to learning how various reinforcement learning techniques and algorithms play an important role in game development with Python. Starting with the basics, this book will help you build a strong foundation in reinforcement learning for game development. Each chapter will assist you in implementing different reinforcement learning techniques, such as Markov decision processes (MDPs), Q-learning, actor-critic methods, SARSA, and deterministic policy gradient algorithms, to build logical self-learning agents. Learning these techniques will enhance your game development skills and add a variety of features to improve your game agent’s productivity. As you advance, you’ll understand how deep reinforcement learning (DRL) techniques can be used to devise strategies to help agents learn from their actions and build engaging games. By the end of this book, you’ll be ready to apply reinforcement learning techniques to build a variety of projects and contribute to open source applications.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Section 1: Exploring the Environment

Free Chapter

Understanding Rewards-Based Learning

Technical requirements

Understanding rewards-based learning

Introducing the Markov decision process

Using value learning with multi-armed bandits

Exploring Q-learning with contextual bandits

Summary

Questions

Dynamic Programming and the Bellman Equation

Introducing DP

Understanding the Bellman equation

Building policy iteration

Building value iteration

Playing with policy versus value iteration

Exercises

Summary

Monte Carlo Methods

Understanding model-based and model-free learning

Introducing the Monte Carlo method

Adding RL

Playing the FrozenLake game

Using prediction and control

Exercises

Summary

Temporal Difference Learning

Understanding the TCA problem

Introducing TDL

Applying TDL to Q-learning

Exploring TD(0) in Q-learning

Running off- versus on-policy

Exercises

Summary

Exploring SARSA

Exploring SARSA on-policy learning

Using continuous spaces with SARSA

Extending continuous spaces

Working with TD (λ) and eligibility traces

Understanding SARSA (λ)

Exercises

Summary

Section 2: Exploiting the Knowledge

Going Deep with DQN

DL for RL

Using PyTorch for DL

Building neural networks with Torch

Understanding DQN in PyTorch

Exercising DQN

Exercises

Summary

Going Deeper with DDQN

Understanding visual state

Introducing CNNs

Working with a DQN on Atari

Introducing DDQN

Extending replay with prioritized experience replay

Exercises

Summary

Policy Gradient Methods

Understanding policy gradient methods

Introducing REINFORCE

Using advantage actor-critic

Building a deep deterministic policy gradient

Exploring trust region policy optimization

Exercises

Summary

Optimizing for Continuous Control

Understanding continuous control with Mujoco

Introducing proximal policy optimization

Using PPO with recurrent networks

Deciding on synchronous and asynchronous actors

Building actor-critic with experience replay

Exercises

Summary

All about Rainbow DQN

Rainbow – combining improvements in deep reinforcement learning

Using TensorBoard

Introducing distributional RL

Understanding noisy networks

Unveiling Rainbow DQN

Exercises

Summary

Exploiting ML-Agents

Installing ML-Agents

Building a Unity environment

Training a Unity environment with Rainbow

Creating a new environment

Advancing RL with ML-Agents

Exercises

Summary

DRL Frameworks

Choosing a framework

Introducing Google Dopamine

Playing with Keras-RL

Exploring RL Lib

Using TF-Agents

Exercises

Summary

Section 3: Reward Yourself

3D Worlds

Reasoning on 3D worlds

Training a visual agent

Generalizing 3D vision

Challenging the Unity Obstacle Tower Challenge

Exploring Habitat – embodied agents by FAIR

Exercises

Summary

From DRL to AGI

Learning meta learning

Introducing meta reinforcement learning

Using hindsight experience replay

Imagination and reasoning in RL

Understanding imagination-augmented agents

Exercises

Summary

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Understanding rewards-based learning

Machine learning is quickly becoming a broad and growing category, with many forms of learning systems addressed. We categorize learning based on the form of a problem and how we need to prepare it for a machine to process. In the case of supervised machine learning, data is first labeled before it is fed into the machine. Examples of this type of learning are simple image classification systems that are trained to recognize a cat or dog from a prelabeled set of cat and dog images. Supervised learning is the most popular and intuitive type of learning system. Other forms of learning that are becoming increasingly powerful are unsupervised and semi-supervised learning. Both of these methods eliminate the need for labels or, in the case of semi-supervised learning, require the labels to be defined more abstractly. The following diagram shows these learning methods and how they process data:

Variations of supervised learning

A couple of recent papers on arXiv.org (pronounced archive.org) suggest the use of semi-supervised learning to solve RL tasks. While the papers suggest no use of external rewards, they do talk about internal updates or feedback signals. This suggests a method of using internal reward RL, which, as we mentioned before, is a thing.

While this family of supervised learning methods has made impressive progress in just the last few years, they still lack the necessary planning and intelligence we expect from a truly intelligent machine. This is where RL picks up and differentiates itself. RL systems learn from interacting and making selections in the environment the agent resides in. The classic diagram of an RL system is shown here:

An RL system

In the preceding diagram, you can identify the main components of an RL system: the Agent and Environment, where the Agent represents the RL system, and the Environment could be representative of a game board, game screen, and/or possibly streaming data. Connecting these components are three primary signals, the State, Reward, and Action. The State signal is essentially a snapshot of the current state of Environment. The Reward signal may be externally provided by the Environment and provides feedback to the agent, either bad or good. Finally, the Action signal is the action the Agent selects at each time step in the environment. An action could be as simple as jump or a more complex set of controls operating servos. Either way, another key difference in RL is the ability for the agent to interact with, and change, the Environment.

Now, don't worry if this all seems a little muddled still—early researchers often encountered trouble differentiating between supervised learning and RL.

In the next section, we look at more RL terminology and explore the basic elements of an RL agent.

The elements of RL

Every RL agent is comprised of four main elements. These are policy, reward function, value function, and, optionally, model. Let's now explore what each of these terms means in more detail:

The policy: A policy represents the decision and planning process of the agent. The policy is what decides the actions the agent will take during a step.
The reward function: The reward function determines what amount of reward an agent receives after completing a series of actions or an action. Generally, a reward is given to an agent externally but, as we will see, there are internal reward systems as well.
The value function: A value function determines the value of a state over the long term. Determining the value of a state is fundamental to RL and our first exercise will be determining state values.
The model: A model represents the environment in full. In the case of a game of tic-tac-toe, this may represent all possible game states. For more advanced RL algorithms, we use the concept of a partially observable state that allows us to do away with a full model of the environment. Some environments that we will tackle in this book have more states than the number of atoms in the universe. Yes, you read that right. In massive environments like that, we could never hope to model the entire environment state.

We will spend the next several chapters covering each of these terms in excruciating detail, so don't worry if things feel a bit abstract still. In the next section, we will take a look at the history of RL.

The history of RL

An Introduction to RL, by Sutton and Barto (1998), discusses the origins of modern RL being derived from two main threads with a later joining thread. The two main threads are trial and error-based learning and dynamic programming, with the third thread arriving later in the form of temporal difference learning. The primary thread founded by Sutton, trial and error, is based on animal psychology. As for the other methods, we will look at each in far more detail in their respective chapters. A diagram showing how these three threads converged to form modern RL is shown here:

The history of modern RL

Dr. Richard S. Sutton, a distinguished research scientist for DeepMind and renowned professor from the University of Alberta, is considered the father of modern RL.

Lastly, before we jump in and start unraveling RL, let's look at why it makes sense to use this form of learning with games in the next section.

Why RL in games?

Various forms of machine learning systems have been used in gaming, with supervised learning being the primary choice. While these methods can be made to look intelligent, they are still limited by working on labeled or categorized data. While generative adversarial networks (GANs) show a particular promise in level and other asset generation, these families of algorithms cannot plan and make sense of long-term decision making. AI systems that replicate planning and interactive behavior in games are now typically done with hardcoded state machine systems such as finite state machines or behavior trees. Being able to develop agents that can learn for themselves the best moves or actions for an environment is literally game-changing, not only for the games industry, of course, but this should surely cause repercussions in every industry globally.

In the next section, we take a look at the foundation of the RL system, the Markov decision process.

Hands-On Reinforcement Learning for Games

By : Micheal Lanham

Hands-On Reinforcement Learning for Games

By: Micheal Lanham

Overview of this book

Related Content you might be interested in

Current Title:

Hands-On Reinforcement Learning for Games

Hands-On Deep Learning for Games

Learn Unity ML-Agents - Fundamentals of Unity Machine Learning

Hands-On Reinforcement Learning with Python