Hadoop Real-World Solutions Cookbook

Hadoop Real-World Solutions Cookbook - Second Edition

By : Tanmay Deshpande

Buy this Book

Hadoop Real-World Solutions Cookbook - Second Edition

By: Tanmay Deshpande

Buy this Book

Overview of this book

Big data is the current requirement. Most organizations produce huge amount of data every day. With the arrival of Hadoop-like tools, it has become easier for everyone to solve big data problems with great efficiency and at minimal cost. Grasping Machine Learning techniques will help you greatly in building predictive models and using this data to make the right decisions for your organization. Hadoop Real World Solutions Cookbook gives readers insights into learning and mastering big data via recipes. The book not only clarifies most big data tools in the market but also provides best practices for using them. The book provides recipes that are based on the latest versions of Apache Hadoop 2.X, YARN, Hive, Pig, Sqoop, Flume, Apache Spark, Mahout and many more such ecosystem tools. This real-world-solution cookbook is packed with handy recipes you can apply to your own everyday issues. Each chapter provides in-depth recipes that can be referenced easily. This book provides detailed practices on the latest technologies such as YARN and Apache Spark. Readers will be able to consider themselves as big data experts on completion of this book. This guide is an invaluable tutorial if you are planning to implement a big data warehouse for your business.

Hadoop Real-World Solutions Cookbook Second Edition

Credits

About the Author

Acknowledgements

About the Reviewer

www.PacktPub.com

Preface

Free Chapter

Getting Started with Hadoop 2.X

Introduction

Installing a single-node Hadoop Cluster

Installing a multi-node Hadoop cluster

Adding new nodes to existing Hadoop clusters

Executing the balancer command for uniform data distribution

Entering and exiting from the safe mode in a Hadoop cluster

Decommissioning DataNodes

Performing benchmarking on a Hadoop cluster

Exploring HDFS

Introduction

Loading data from a local machine to HDFS

Exporting HDFS data to a local machine

Changing the replication factor of an existing file in HDFS

Setting the HDFS block size for all the files in a cluster

Setting the HDFS block size for a specific file in a cluster

Enabling transparent encryption for HDFS

Importing data from another Hadoop cluster

Recycling deleted data from trash to HDFS

Saving compressed data in HDFS

Mastering Map Reduce Programs

Introduction

Writing the Map Reduce program in Java to analyze web log data

Executing the Map Reduce program in a Hadoop cluster

Adding support for a new writable data type in Hadoop

Implementing a user-defined counter in a Map Reduce program

Map Reduce program to find the top X

Map Reduce program to find distinct values

Map Reduce program to partition data using a custom partitioner

Writing Map Reduce results to multiple output files

Performing Reduce side Joins using Map Reduce

Unit testing the Map Reduce code using MRUnit

Data Analysis Using Hive, Pig, and Hbase

Introduction

Storing and processing Hive data in a sequential file format

Storing and processing Hive data in the ORC file format

Storing and processing Hive data in the Parquet file format

Performing FILTER By queries in Pig

Performing Group By queries in Pig

Performing Order By queries in Pig

Performing JOINS in Pig

Writing a user-defined function in Pig

Analyzing web log data using Pig

Performing the Hbase operation in CLI

Performing Hbase operations in Java

Executing the MapReduce programming with an Hbase Table

Advanced Data Analysis Using Hive

Introduction

Processing JSON data in Hive using JSON SerDe

Processing XML data in Hive using XML SerDe

Processing Hive data in the Avro format

Writing a user-defined function in Hive

Performing table joins in Hive

Executing map side joins in Hive

Performing context Ngram in Hive

Call Data Record Analytics using Hive

Twitter sentiment analysis using Hive

Implementing Change Data Capture using Hive

Multiple table inserting using Hive

Data Import/Export Using Sqoop and Flume

Introduction

Importing data from RDMBS to HDFS using Sqoop

Exporting data from HDFS to RDBMS

Using query operator in Sqoop import

Importing data using Sqoop in compressed format

Performing Atomic export using Sqoop

Importing data into Hive tables using Sqoop

Importing data into HDFS from Mainframes

Incremental import using Sqoop

Creating and executing Sqoop job

Importing data from RDBMS to Hbase using Sqoop

Importing Twitter data into HDFS using Flume

Importing data from Kafka into HDFS using Flume

Importing web logs data into HDFS using Flume

Automation of Hadoop Tasks Using Oozie

Introduction

Implementing a Sqoop action job using Oozie

Implementing a Map Reduce action job using Oozie

Implementing a Java action job using Oozie

Implementing a Hive action job using Oozie

Implementing a Pig action job using Oozie

Implementing an e-mail action job using Oozie

Executing parallel jobs using Oozie (fork)

Scheduling a job in Oozie

Machine Learning and Predictive Analytics Using Mahout and R

Introduction

Setting up the Mahout development environment

Creating an item-based recommendation engine using Mahout

Creating a user-based recommendation engine using Mahout

Using Predictive analytics on Bank Data using Mahout

Clustering text data using K-Means

Performing Population Data Analytics using R

Performing Twitter Sentiment Analytics using R

Performing Predictive Analytics using R

Integration with Apache Spark

Introduction

Running Spark standalone

Running Spark on YARN

Olympics Athletes analytics using the Spark Shell

Creating Twitter trending topics using Spark Streaming

Analyzing Parquet files using Spark

Analyzing JSON data using Spark

Processing graphs using Graph X

Conducting predictive analytics using Spark MLib

Hadoop Use Cases

Introduction

Call Data Record analytics

Web log analytics

Sensitive data masking and encryption using Hadoop

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Adding new nodes to existing Hadoop clusters

Sometimes, it may happen that an existing Hadoop cluster's capacity is not adequate enough to handle all the data you may want to process. In this case, you can add new nodes to the existing Hadoop cluster without any downtime for the existing cluster. Hadoop supports horizontal scalability.

Getting ready

To perform this recipe, you should have a Hadoop cluster running. Also, you will need one more machine. If you are using AWS EC2, then you can launch an EC2 instance that's similar to what we did in the previous recipes. You will also need the same security group configurations in order to make the installation process smooth.

How to do it...

To add a new instance to an existing cluster, simply install and configure Hadoop the way we did for the previous recipe. Make sure that you put the same configurations in core-site.xml and yarn-site.xml, which will point to the correct master node.

Once all the configurations are done, simply execute commands to start the newly added datanode and nodemanager:

/usr/local/hadoop/sbin/hadoop-daemon.sh start datanode
/usr/local/hadoop/sbin/yarn-daemon.sh start nodemanager

If you take a look at the cluster again, you will find that the new node is registered. You can use the dfsadmin command to take a look at the number of nodes and amount of capacity that's been used:

hdfs dfsadmin -report

Here is a sample output for the preceding command:

How it works...

Hadoop supports horizontal scalability. If the resources that are being used are not enough, we can always go ahead and add new nodes to the existing cluster without hiccups. In Hadoop, it's always the slave that reports to the master. So, while making configurations, we always configure the details of the master and do nothing about the slaves. This architecture helps achieve horizontal scalability as at any point of time, we can add new nodes by only providing the configurations of the master, and everything else is taken care of by the Hadoop cluster. As soon as the daemons start, the master node realizes that a new node has been added and it becomes part of the cluster.

Hadoop Real-World Solutions Cookbook - Second Edition

By : Tanmay Deshpande

Hadoop Real-World Solutions Cookbook - Second Edition

By: Tanmay Deshpande

Overview of this book

Related Content you might be interested in

Current Title:

Hadoop Real-World Solutions Cookbook - Second Edition

Adding new nodes to existing Hadoop clusters

Getting ready

How to do it...

How it works...