In this recipe, we will configure the Oozie workflow engine and look at some examples of scheduling jobs using the Oozie workflow.
Oozie is a scheduler to manage Hadoop jobs, with the ability to make decisions on the conditions or states of previous jobs or the presence of certain files.
Make sure that you have completed the recipe of Hadoop cluster configuration with the edge node configured. HDFS and YARN must be configured and healthy before starting this recipe.
Connect to the edge node
edge1.cyrus.com
and switch to thehadoop
user.Download the Oozie source package, untar it, and build it:
$ tar –xzvfoozie-4.1.0.tar.gz $ cd oozie-4.1.0
Edit the file
pom.xml
to make it suitable for the Java version and the Hadoop version. Change the fields according to the version of Java used. For Hadoop 2.x, it must be version 2.3.0:<targetJavaVersion>1.8</targetJavaVersion> <hadoop.version>2.3.0</hadoop.version> <hbase...