Book Image

OpenStack Sahara Essentials

By : Omar Khedher
Book Image

OpenStack Sahara Essentials

By: Omar Khedher

Overview of this book

The Sahara project is a module that aims to simplify the building of data processing capabilities on OpenStack. The goal of this book is to provide a focused, fast paced guide to installing, configuring, and getting started with integrating Hadoop with OpenStack, using Sahara. The book should explain to users how to deploy their data-intensive Hadoop and Spark clusters on top of OpenStack. It will also cover how to use the Sahara REST API, how to develop applications for Elastic Data Processing on Openstack, and setting up hadoop or spark clusters on Openstack.
Table of Contents (14 chapters)

Summary


In this chapter, we have installed OpenStack using the Packstack tool. Sahara has been successfully integrated and it is possible to instruct the Elastic Data Processing component in OpenStack using either the command-line interface or via Horizon.

Before walking through the rest of the chapters, it might be essential to have the OpenStack environment up and running without any issues. As has been described in Chapter 1, The Essence of Big Data in the Cloud, the Sahara service depends on other services and components of OpenStack. To run a Hadoop cluster, spin nodes, and run jobs, different services in OpenStack are involved. This will be covered in more detail in the next chapter, where we will examine the workflow of creating a Hadoop cluster in OpenStack using Sahara.