Java for Data Science

Java for Data Science

By : Richard M. Reese, Jennifer L. Reese

Buy this Book

Java for Data Science

By: Richard M. Reese, Jennifer L. Reese

Buy this Book

Overview of this book

para 1: Get the lowdown on Java and explore big data analytics with Java for Data Science. Packed with examples and data science principles, this book uncovers the techniques & Java tools supporting data science and machine learning. Para 2: The stability and power of Java combines with key data science concepts for effective exploration of data. By working with Java APIs and techniques, this data science book allows you to build applications and use analysis techniques centred on machine learning. Para 3: Java for Data Science gives you the understanding you need to examine the techniques and Java tools supporting big data analytics. These Java-based approaches allow you to tackle data mining and statistical analysis in detail. Deep learning and Java data mining are also featured, so you can explore and analyse data effectively, and build intelligent applications using machine learning. para 4: What?s Inside ? Understand data science principles with Java support ? Discover machine learning and deep learning essentials ? Explore data science problems with Java-based solutions

Java for Data Science

Credits

About the Authors

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Getting Started with Data Science

Problems solved using data science

Understanding the data science problem - solving approach

Acquiring data for an application

The importance and process of cleaning data

Visualizing data to enhance understanding

The use of statistical methods in data science

Machine learning applied to data science

Using neural networks in data science

Deep learning approaches

Performing text analysis

Visual and audio analysis

Improving application performance using parallel techniques

Assembling the pieces

Summary

Data Acquisition

Understanding the data formats used in data science applications

Data acquisition techniques

Summary

Data Cleaning

Handling data formats

The nitty gritty of cleaning text

Cleaning images

Summary

Data Visualization

Understanding plots and graphs

Creating index charts

Creating bar charts

Creating stacked graphs

Creating pie charts

Creating scatter charts

Creating histograms

Creating donut charts

Creating bubble charts

Summary

Statistical Data Analysis Techniques

Working with mean, mode, and median

Standard deviation

Sample size determination

Hypothesis testing

Regression analysis

Summary

Machine Learning

Supervised learning techniques

Unsupervised machine learning

Reinforcement learning

Summary

Neural Networks

Training a neural network

Understanding static neural networks

Understanding dynamic neural networks

Additional network architectures and algorithms

Summary

Deep Learning

Deeplearning4j architecture

Deep learning and regression analysis

Restricted Boltzmann Machines

Deep autoencoders

Convolutional networks

Recurrent Neural Networks

Summary

Text Analysis

Implementing named entity recognition

Classifying text

Understanding tagging and POS

Extracting relationships from sentences

Sentiment analysis

Summary

Visual and Audio Analysis

Text-to-speech

Understanding speech recognition

Extracting text from an image

Identifying faces

Classifying visual data

Summary

Mathematical and Parallel Techniques for Data Analysis

Implementing basic matrix operations

Using map-reduce

Various mathematical libraries

Using OpenCL

Using Aparapi

Using Java 8 streams

Summary

Bringing It All Together

Defining the purpose and scope of our application

Understanding the application's architecture

Data acquisition using Twitter

Understanding the TweetHandler class

Other optional enhancements

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Visual and audio analysis

In Chapter 10, Visual and Audio Analysis, we demonstrate several Java techniques for processing sounds and images. We begin by demonstrating techniques for sound processing, including speech recognition and text-to-speech APIs. Specifically, we will use the FreeTTS (http://freetts.sourceforge.net/docs/index.php) API to convert text to speech. We also include a demonstration of the CMU Sphinx toolkit for speech recognition.

The Java Speech API (JSAPI) (http://www.oracle.com/technetwork/java/index-140170.html) supports speech technology. This API, created by third-party vendors, supports speech recognition and speech synthesizers. FreeTTS and Festival (http://www.cstr.ed.ac.uk/projects/festival/) are examples of vendors supporting JSAPI.

In the second part of the chapter, we examine image processing techniques such as facial recognition. This demonstration involves identifying faces within an image and is easy to accomplish using OpenCV (http://opencv.org/).

Also, in Chapter 10, Visual and Audio Analysis, we demonstrate how to extract text from images, a process known as OCR. A common data science problem involves extracting and analyzing text embedded in an image. For example, the information contained in license plate, road signs, and directions can be significant.

In the following example, explained in more detail in Chapter 11, Mathematical and Parallel Techniques for Data Analysis accomplishes OCR using Tess4j (http://tess4j.sourceforge.net/) a Java JNA wrapper for Tesseract OCR API. We perform OCR on an image captured from the Wikipedia article on OCR (https://en.wikipedia.org/wiki/Optical_character_recognition#Applications), shown here:

The ITesseract interface provides numerous OCR methods. The doOCR method takes a file and returns a string containing the words found in the file as shown here:

ITesseract instance = new Tesseract();  
try { 
    String result = instance.doOCR(new File("OCRExample.png")); 
    System.out.println(result); 
} catch (TesseractException e) { 
    System.err.println(e.getMessage()); 
}

A part of the output is shown next:

OCR engines nave been developed into many lunds oiobiectorlented OCR applicatlons, sucn as reoeipt OCR, involoe OCR, check OCR, legal billing document OCR

They can be used ior

- Data entry ior business documents, e g check, passport, involoe, bank statement and receipt

- Automatic number plate recognnlon

As you can see, there are numerous errors in this example that need to be addressed. We build upon this example in Chapter 11, Mathematical and Parallel Techniques for Data Analysis, with a discussion of enhancements and considerations to ensure the OCR process is as effective as possible.

We will conclude the chapter with a discussion of NeurophStudio, a neural network Java-based editor, to classify images and perform image recognition. We train a neural network to recognize and classify faces in this section.

Java for Data Science

By : Richard M. Reese, Jennifer L. Reese

Java for Data Science

By: Richard M. Reese, Jennifer L. Reese

Overview of this book

Related Content you might be interested in

Current Title:

Java for Data Science

Natural Language Processing with Java

Machine Learning in Java

Java Data Science Cookbook

Visual and audio analysis