Book Image

Mastering Ceph

By : Nick Fisk

Book Image

Mastering Ceph

By: Nick Fisk

Overview of this book

Mastering Ceph covers all that you need to know to use Ceph effectively. Starting with design goals and planning steps that should be undertaken to ensure successful deployments, you will be guided through to setting up and deploying the Ceph cluster, with the help of orchestration tools. Key areas of Ceph including Bluestore, Erasure coding and cache tiering will be covered with help of examples. Development of applications which use Librados and Distributed computations with shared object classes are also covered. A section on tuning will take you through the process of optimisizing both Ceph and its supporting infrastructure. Finally, you will learn to troubleshoot issues and handle various scenarios where Ceph is likely not to recover on its own. By the end of the book, you will be able to successfully deploy and operate a resilient high performance Ceph cluster.

Preface

What this book covers

What you need for this book

Who this book is for

Reader feedback

Customer support

Free Chapter

Planning for Ceph

Planning for Ceph

How Ceph works?

Specific use cases

Infrastructure design

How to plan a successful Ceph implementation

Deploying Ceph

Preparing your environment with Vagrant and VirtualBox

A very simple playbook

Adding the Ceph Ansible modules

Change and configuration management

BlueStore

What is BlueStore?

Why was it needed?

How BlueStore works

How to use BlueStore

Erasure Coding for Better Storage Efficiency

Erasure Coding for Better Storage Efficiency

What is erasurecoding?

How does erasure coding work in Ceph?

Algorithms and profiles

Where can I use erasure coding?

Creating an erasure-coded pool

Developing with Librados

Developing with Librados

What is librados?

How to use librados?

Example librados application

Distributed Computation with Ceph RADOS Classes

Distributed Computation with Ceph RADOS Classes

Example applications and the benefits of using RADOS classes

Writing a simple RADOS class in Lua

Writing a RADOS class that simulates distributed computing

RADOS class caveats

Monitoring Ceph

Monitoring Ceph

Why it is important to monitor Ceph

What should be monitored

PG states -the good, the bad, and the ugly

Monitoring Ceph with collectd

Tiering with Ceph

Tiering with Ceph

Tiering versus caching

What is a bloom filter

Creating tiers in Ceph

Promotion throttling

Tuning Ceph

Recommended tunings

Troubleshooting

Troubleshooting

Repairing inconsistent objects

Slow performance

Extremely slow performance or no IO

Investigating PGs in a down state

Large monitor databases

Disaster Recovery

Disaster Recovery

What is a disaster?

Avoiding data loss

What can cause an outage or data loss?

Lost objects and inactive PGs

Recovering from a complete monitor failure

Using the Cephs object store tool

Investigating asserts

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

What can cause an outage or data loss?

The majority of outages and cases of data loss will be directly caused by the loss of a number of OSDs that exceed the replication level in a short period of time. If these OSDs do not come back online, be it due to a software or hardware failure and Ceph was not able to recover objects in-between OSD failures, then these objects are now lost.

If an OSD has failed due to a failed disk, then it is unlikely that recovery will be possible unless costly disk recovery services are utilized, and there is no guarantee that any recovered data will be in a consistent state. This chapter will not cover recovering from physical disk failures and will simply suggest that the default replication level of 3 should be used to protect you against multiple disk failures.

If an OSD has failed due to a software bug, the outcome is possibly a lot more positive, but the process is complex and time...