Pentaho Data Integration 4 Cookbook

Pentaho Data Integration 4 Cookbook

Overview of this book

Pentaho Data Integration (PDI, also called Kettle), one of the data integration tools leaders, is broadly used for all kind of data manipulation such as migrating data between applications or databases, exporting data from databases to flat files, data cleansing, and much more. Do you need quick solutions to the problems you face while using Kettle? Pentaho Data Integration 4 Cookbook explains Kettle features in detail through clear and practical recipes that you can quickly apply to your solutions. The recipes cover a broad range of topics including processing files, working with databases, understanding XML structures, integrating with Pentaho BI Suite, and more. Pentaho Data Integration 4 Cookbook shows you how to take advantage of all the aspects of Kettle through a set of practical recipes organized to find quick solutions to your needs. The initial chapters explain the details about working with databases, files, and XML structures. Then you will see different ways for searching data, executing and reusing jobs and transformations, and manipulating streams. Further, you will learn all the available options for integrating Kettle with other Pentaho tools. Pentaho Data Integration 4 Cookbook has plenty of recipes with easy step-by-step instructions to accomplish specific tasks. There are examples and code that are ready for adaptation to individual needs.

Pentaho Data Integration 4 Cookbook

Credits

About the Authors

About the Reviewers

www.PacktPub.com

Preface

Free Chapter

Working with Databases

Introduction

Connecting to a database

Getting data from a database

Getting data from a database by providing parameters

Getting data from a database by running a query built at runtime

Inserting or updating rows in a table

Inserting new rows where a simple primary key has to be generated

Inserting new rows where the primary key has to be generated based on stored values

Deleting data from a table

Creating or altering a database table from PDI (design time)

Creating or altering a database table from PDI (runtime)

Inserting, deleting, or updating a table depending on a field

Changing the database connection at runtime

Loading a parent-child table

Reading and Writing Files

Introduction

Reading a simple file

Reading several files at the same time

Reading unstructured files

Reading files having one field by row

Reading files with some fields occupying two or more rows

Writing a simple file

Writing an unstructured file

Providing the name of a file (for reading or writing) dynamically

Using the name of a file (or part of it) as a field

Reading an Excel file

Getting the value of specific cells in an Excel file

Writing an Excel file with several sheets

Writing an Excel file with a dynamic number of sheets

Manipulating XML Structures

Introduction

Reading simple XML files

Specifying fields by using XPath notation

Validating well-formed XML files

Validating an XML file against DTD definitions

Validating an XML file against an XSD schema

Generating a simple XML document

Generating complex XML structures

Generating an HTML page using XML and XSL transformations

File Management

Introduction

Copying or moving one or more files

Deleting one or more files

Getting files from a remote server

Putting files on a remote server

Copying or moving a custom list of files

Deleting a custom list of files

Comparing files and folders

Working with ZIP files

Looking for Data

Introduction

Looking for values in a database table

Looking for values in a database (with complex conditions or multiple tables involved)

Looking for values in a database with extreme flexibility

Looking for values in a variety of sources

Looking for values by proximity

Looking for values consuming a web service

Looking for values over an intranet or Internet

Understanding Data Flows

Introduction

Splitting a stream into two or more streams based on a condition

Merging rows of two streams with the same or different structures

Comparing two streams and generating differences

Generating all possible pairs formed from two datasets

Joining two or more streams based on given conditions

Interspersing new rows between existent rows

Executing steps even when your stream is empty

Processing rows differently based on the row number

Executing and Reusing Jobs and Transformations

Introduction

Executing a job or a transformation by setting static arguments and parameters

Executing a job or a transformation from a job by setting arguments and parameters dynamically

Executing a job or a transformation whose name is determined at runtime

Executing part of a job once for every row in a dataset

Executing part of a job several times until a condition is true

Creating a process flow

Moving part of a transformation to a subtransformation

Integrating Kettle and the Pentaho Suite

Introduction

Creating a Pentaho report with data coming from PDI

Configuring the Pentaho BI Server for running PDI jobs and transformations

Executing a PDI transformation as part of a Pentaho process

Executing a PDI job from the Pentaho User Console

Generating files from the PUC with PDI and the CDA plugin

Populating a CDF dashboard with data coming from a PDI transformation

Getting the Most Out of Kettle

Introduction

Sending e-mails with attached files

Generating a custom log file

Programming custom functionality

Generating sample data for testing purposes

Working with Json files

Getting information about transformations and jobs (file-based)

Getting information about transformations and jobs (repository-based)

Data Structures

Book's data structure

Museum's data structure

Outdoor data structure

Steel Wheels structure

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Getting data from a database by running a query built at runtime

When you work with databases, most of the times you start by writing an SQL statement that gets the data you need. However, there are situations in which you don't know that statement exactly. Maybe the name of the columns to query are in a file, or the name of the columns by which you will sort will come as a parameter from outside the transformation, or the name of the main table to query changes depending on the data stored in it (for example sales2010). PDI allows you to have any part of the SQL statement as a variable so you don't need to know the literal SQL statement text at design time.

Assume the following situation: You have a database with data about books and their authors, and you want to generate a file with a list of titles. Whether to retrieve the data ordered by title or by genre is a choice that you want to postpone until the moment you execute the transformation.

Getting ready

You will need a book database with the structure explained in Appendix, Data Structures.

How to do it...

Create a transformation.
The column that will define the order of the rows will be a named parameter. So, define a named parameter named ORDER_COLUMN, and put title as its default value.
Note
Remember that Named Parameters are defined in the Transformation setting window and their role is the same as the role of any Kettle variable. If you prefer, you can skip this step and define a standard variable for this purpose.
Now drag a Table Input step to the canvas. Then create and select the connection to the book's database.

In the SQL frame type the following statement:

SELECT * FROM books ORDER BY ${ORDER_COLUMN}

Check the option Replace variables in script? and close the window.
Use an Output step like for example a Text file output step, to send the results to a file, save the transformation, and run it.
Open the generated file and you will see the books ordered by title.
Now try again. Press F9 to run the transformation one more time.
This time, change the value of the ORDER_COLUMN parameter typing genre as the new value.
Press the Launch button.
Open the generated file. This time you will see the titles ordered by genre.

How it works...

You can use Kettle variables in any part of the SELECT statement inside a Table Input step. When the transformation is initialized, PDI replaces the variables by their values provided that the Replace variables in script? option is checked.

In the recipe, the first time you ran the transformation, Kettle replaced the variable ORDER_COLUMN with the word title and the statement executed was:

SELECT * FROM books ORDER BY title

The second time, the variable was replaced by genre and the executed statement was:

SELECT * FROM books ORDER BY genre

Note

As mentioned in the recipe, any predefined Kettle variable can be used instead of a named parameter.

There's more...

You may use variables not only for the ORDER BY clause, but in any part of the statement: table names, columns, and so on. You could even hold the full statement in a variable. Note however that you need to be cautious when implementing this.

Note

A wrong assumption about the metadata generated by those predefined statements can make your transformation crash.

You can also use the same variable more than once in the same statement. This is an advantage of using variables as an alternative to question marks when you need to execute parameterized SELECT statements.

Pentaho Data Integration 4 Cookbook

Pentaho Data Integration 4 Cookbook

Overview of this book

Related Content you might be interested in

Current Title:

Pentaho Data Integration 4 Cookbook

Getting data from a database by running a query built at runtime

Getting ready

How to do it...

Note

How it works...

Note

There's more...

Note

See also