Hadoop 2.x Administration Cookbook

Hadoop 2.x Administration Cookbook

By : Aman Singh

Buy this Book

Hadoop 2.x Administration Cookbook

By: Aman Singh

Buy this Book

Overview of this book

Hadoop enables the distributed storage and processing of large datasets across clusters of computers. Learning how to administer Hadoop is crucial to exploit its unique features. With this book, you will be able to overcome common problems encountered in Hadoop administration. The book begins with laying the foundation by showing you the steps needed to set up a Hadoop cluster and its various nodes. You will get a better understanding of how to maintain Hadoop cluster, especially on the HDFS layer and using YARN and MapReduce. Further on, you will explore durability and high availability of a Hadoop cluster. You’ll get a better understanding of the schedulers in Hadoop and how to configure and use them for your tasks. You will also get hands-on experience with the backup and recovery options and the performance tuning aspects of Hadoop. Finally, you will get a better understanding of troubleshooting, diagnostics, and best practices in Hadoop administration. By the end of this book, you will have a proper understanding of working with Hadoop clusters and will also be able to secure, encrypt it, and configure auditing for your Hadoop clusters.

Hadoop 2.x Administration Cookbook

Credits

About the Author

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Hadoop Architecture and Deployment

Introduction

Building and compiling Hadoop

Installation methods

Setting up host resolution

Installing a single-node cluster - HDFS components

Installing a single-node cluster - YARN components

Installing a multi-node cluster

Configuring the Hadoop Gateway node

Decommissioning nodes

Adding nodes to the cluster

Maintaining Hadoop Cluster HDFS

Introduction

Configuring HDFS block size

Setting up Namenode metadata location

Loading data in HDFS

Configuring HDFS replication

HDFS balancer

Quota configuration

HDFS health and FSCK

Configuring rack awareness

Recycle or trash bin configuration

Distcp usage

Control block report storm

Configuring Datanode heartbeat

Maintaining Hadoop Cluster – YARN and MapReduce

Introduction

Running a simple MapReduce program

Hadoop streaming

Configuring YARN history server

Job history web interface and metrics

Configuring ResourceManager components

YARN containers and resource allocations

ResourceManager Web UI and JMX metrics

Preserving ResourceManager states

High Availability

Introduction

Namenode HA using shared storage

ZooKeeper configuration

Namenode HA using Journal node

Resourcemanager HA using ZooKeeper

Rolling upgrade with HA

Configure shared cache manager

Configure HDFS cache

HDFS snapshots

Configuring storage based policies

Configuring HA for Edge nodes

Schedulers

Introduction

Configuring users and groups

Fair Scheduler configuration

Fair Scheduler pools

Configuring job queues

Job queue ACLs

Configuring Capacity Scheduler

Queuing mappings in Capacity Scheduler

YARN and Mapred commands

YARN label-based scheduling

YARN SLS

Backup and Recovery

Introduction

Initiating Namenode saveNamespace

Using HDFS Image Viewer

Fetching parameters which are in-effect

Configuring HDFS and YARN logs

Backing up and recovering Namenode

Configuring Secondary Namenode

Promoting Secondary Namenode to Primary

Namenode recovery

Namenode roll edits – online mode

Namenode roll edits – offline mode

Datanode recovery – disk full

Configuring NFS gateway to serve HDFS

Recovering deleted files

Data Ingestion and Workflow

Introduction

Hive server modes and setup

Using MySQL for Hive metastore

Operating Hive with ZooKeeper

Loading data into Hive

Partitioning and Bucketing in Hive

Hive metastore database

Designing Hive with credential store

Configuring Flume

Configure Oozie and workflows

Performance Tuning

Tuning the operating system

Configuring YARN for performance

Configuring MapReduce for performance

Hive performance tuning

Benchmarking Hadoop cluster

HBase Administration

Introduction

Setting up single node HBase cluster

Setting up multi-node HBase cluster

Inserting data into HBase

Integration with Hive

HBase administration commands

HBase backup and restore

Tuning HBase

HBase upgrade

Migrating data from MySQL to HBase using Sqoop

Cluster Planning

Introduction

Disk space calculations

Nodes needed in the cluster

Memory requirements

Sizing the cluster as per SLA

Network design

Estimating the cost of the Hadoop cluster

Hardware and software options

Troubleshooting, Diagnostics, and Best Practices

Introduction

Namenode troubleshooting

Datanode troubleshooting

Resourcemanager troubleshooting

Diagnose communication issues

Parse logs for errors

Hive troubleshooting

HBase troubleshooting

Hadoop best practices

Security

Introduction

Encrypting disk using LUKS

Configuring Hadoop users

HDFS encryption at Rest

Configuring SSL in Hadoop

In-transit encryption

Enabling service level authorization

Securing ZooKeeper

Configuring auditing

Configuring Kerberos server

Configuring and enabling Kerberos for Hadoop

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Building and compiling Hadoop

The pre-build Hadoop binary available at www.apache.org, is a 32-bit version and is not suitable for the 64-bit hardware as it will not be able to utilize the entire addressable memory. Although, for lab purposes, we can use the 32-bit version, it will keep on giving warnings about the "not being built for the native library", which can be safely ignored.

In production, we will always be running Hadoop on hardware which is a 64-bit version and can support larger amounts of memory. To properly utilize memory higher than 4 GB on any node, we need the 64-bit complied version of Hadoop.

Getting ready

To step through the recipes in this chapter, or indeed the entire book, you will need at least one preinstalled Linux instance. You can use any distribution of Linux, such as Ubuntu, CentOS, or any other Linux flavor that the reader is comfortable with. The recipes are very generic and are expected to work with all distributions, although, as stated before, one may need to use distro-specific commands. For example, for package installation in CentOS we use yum package installer, or in Debian-based systems we use apt-get, and so on. The user is expected to know basic Linux commands and should know how to set up package repositories such as the yum repository. The user should also know how the DNS resolution is configured. No other prerequisites are required.

How to do it...

ssh to the Linux instance using any of the ssh clients. If you are on Windows, you need PuTTY. If you are using a Mac or Linux, there is a default terminal available to use ssh. The following command connects to the host with an IP of 10.0.0.4. Change it to whatever the IP is in your case:
```
$ ssh [email protected]
```
Change to the user root or any other privileged user:
```
$ sudo su -
```

Install the dependencies to build Hadoop:

# yum install gcc gcc-c++ openssl-devel make cmake jdk-1.7u45(minimum)

Download and install Maven:

wget mirrors.gigenet.com/apache/maven/maven-3/3.3.9/binaries/apache-maven-3.3.9-bin.tar.gz

Untar Maven:

# tar -zxf apache-maven-3.3.9-bin.tar.gz -C /opt/

Set up the Maven environment:

# cat /etc/profile.d/maven.sh
export JAVA_HOME=/usr/java/latest
export M3_HOME=/opt/apache-maven-3.3.9
export PATH=$JAVA_HOME/bin:/opt/apache-maven-3.3.9/bin:$PATH

Download and set up protobuf:

# wget https://github.com/google/protobuf/releases/download/v2.5.0/protobuf-2.5.0.tar.gz
# tar -xzf protobuf-2.5.0.tar.gz -C /root
# cd /opt/protobuf-2.5.0/
# ./configure
# make;make install

Download the latest Hadoop stable source code. At the time of writing, the latest Hadoop version is 2.7.3:

# wget apache.uberglobalmirror.com/hadoop/common/stable2/hadoop-2.7.3-src.tar.gz
# tar -xzf hadoop-2.7.3-src.tar.gz -C /opt/
# cd /opt/hadoop-2.7.2-src
# mvn package -Pdist,native -DskipTests -Dtar

You will see a tarball in the folder hadoop-2.7.3-src/hadoop-dist/target/.

How it works...

The tarball package created will be used for the installation of Hadoop throughout the book. It is not mandatory to build a Hadoop from source, but by default the binary packages provided by Apache Hadoop are 32-bit versions. For production, it is important to use a 64-bit version so as to fully utilize the memory beyond 4 GB and to gain other performance benefits.

Hadoop 2.x Administration Cookbook

By : Aman Singh

Hadoop 2.x Administration Cookbook

By: Aman Singh

Overview of this book

Related Content you might be interested in

Current Title:

Hadoop 2.x Administration Cookbook

Mastering Hadoop 3

Apache Hadoop 3 Quick Start Guide

HBase High Performance Cookbook

Building and compiling Hadoop

Getting ready

How to do it...

How it works...