Mastering Java Machine Learning

Mastering Java Machine Learning

By : Uday Kamath, Krishna Choppella

Buy this Book

Mastering Java Machine Learning

By: Uday Kamath, Krishna Choppella

Buy this Book

Overview of this book

Java is one of the main languages used by practicing data scientists; much of the Hadoop ecosystem is Java-based, and it is certainly the language that most production systems in Data Science are written in. If you know Java, Mastering Machine Learning with Java is your next step on the path to becoming an advanced practitioner in Data Science. This book aims to introduce you to an array of advanced techniques in machine learning, including classification, clustering, anomaly detection, stream learning, active learning, semi-supervised learning, probabilistic graph modeling, text mining, deep learning, and big data batch and stream machine learning. Accompanying each chapter are illustrative examples and real-world case studies that show how to apply the newly learned techniques using sound methodologies and the best Java-based tools available today. On completing this book, you will have an understanding of the tools and techniques for building powerful machine learning models to solve data science problems in just about any domain.

Mastering Java Machine Learning

Credits

Foreword

About the Authors

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Machine Learning Review

Machine learning – history and definition

What is not machine learning?

Machine learning – concepts and terminology

Machine learning – types and subtypes

Datasets used in machine learning

Machine learning applications

Practical issues in machine learning

Machine learning – roles and process

Machine learning – tools and datasets

Summary

Practical Approach to Real-World Supervised Learning

Formal description and notation

Data transformation and preprocessing

Feature relevance analysis and dimensionality reduction

Model building

Model assessment, evaluation, and comparisons

Case Study – Horse Colic Classification

Summary

References

Unsupervised Machine Learning Techniques

Issues in common with supervised learning

Issues specific to unsupervised learning

Feature analysis and dimensionality reduction

Clustering

Outlier or anomaly detection

Real-world case study

Summary

References

Semi-Supervised and Active Learning

Semi-supervised learning

Active learning

Case study in active learning

Summary

References

Real-Time Stream Machine Learning

Assumptions and mathematical notations

Basic stream processing and computational techniques

Concept drift and drift detection

Incremental supervised learning

Incremental unsupervised learning using clustering

Unsupervised learning using outlier detection

Case study in stream learning

Summary

References

Probabilistic Graph Modeling

Probability revisited

Graph concepts

Bayesian networks

Markov networks and conditional random fields

Summary

Deep Learning

Multi-layer feed-forward neural network

Limitations of neural networks

Deep learning

Case study

Summary

References

Text Mining and Natural Language Processing

NLP, subfields, and tasks

Issues with mining unstructured data

Text processing components and transformations

Topics in text mining

Tools and usage

Summary

References

Big Data Machine Learning – The Final Frontier

What are the characteristics of Big Data?

Big Data Machine Learning

Batch Big Data Machine Learning

Case study

Linear Algebra

Vector

Matrix

Probability

Axioms of probability

Bayes' theorem

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Datasets used in machine learning

To learn from data, we must be able to understand and manage data in all forms. Data originates from many different sources, and consequently, datasets may differ widely in structure or have little or no structure at all. In this section, we present a high-level classification of datasets with commonly occurring examples.

Based on their structure, or the lack thereof, datasets may be classified as containing the following:

Structured data: Datasets with structured data are more amenable to being used as input to most machine learning algorithms. The data is in the form of records or rows following a well-known format with features that are either columns in a table or fields delimited by separators or tokens. There is no explicit relationship between the records or instances. The dataset is available chiefly in flat files or relational databases. The records of financial transactions at a bank shown in the following figure are an example of structured data:
Financial card transactional data with labels of fraud
Transaction or market data: This is a special form of structured data where each entry corresponds to a collection of items. Examples of market datasets are the lists of grocery items purchased by different customers or movies viewed by customers, as shown in the following table:
Market dataset for items bought from grocery store
Unstructured data: Unstructured data is normally not available in well-known formats, unlike structured data. Text data, image, and video data are different formats of unstructured data. Usually, a transformation of some form is needed to extract features from these forms of data into a structured dataset so that traditional machine learning algorithms can be applied.
Sample text data, with no discernible structure, hence unstructured. Separating spam from normal messages (ham) is a binary classification problem. Here true positives (spam) and true negatives (ham) are distinguished by their labels, the second token in each instance of data. SMS Spam Collection Dataset (UCI Machine Learning Repository), source: Tiago A. Almeida from the Federal University of Sao Carlos.
Sequential data: Sequential data have an explicit notion of "order" to them. The order can be some relationship between features and a time variable in time series data, or it can be symbols repeating in some form in genomic datasets. Two examples of sequential data are weather data and genomic sequence data. The following figure shows the relationship between time and the sensor level for weather:
Time series from sensor data
Three genomic sequences are taken into consideration to show the repetition of the sequences CGGGT and TTGAAAGTGGTG in all three genomic sequences:
Genomic sequences of DNA as a sequence of symbols.
Graph data: Graph data is characterized by the presence of relationships between entities in the data to form a graph structure. Graph datasets may be in a structured record format or an unstructured format. Typically, the graph relationship has to be mined from the dataset. Claims in the insurance domain can be considered structured records containing relevant claim details with claimants related through addresses, phone numbers, and so on. This can be viewed in a graph structure. Using the World Wide Web as an example, we have web pages available as unstructured data containing links, and graphs of relationships between web pages that can be built using web links, producing some of the most extensively mined graph datasets today:
Insurance claim data, converted into a graph structure showing the relationship between vehicles, drivers, policies, and addresses

Mastering Java Machine Learning

By : Uday Kamath, Krishna Choppella

Mastering Java Machine Learning

By: Uday Kamath, Krishna Choppella

Overview of this book

Related Content you might be interested in

Current Title:

Mastering Java Machine Learning

Machine Learning in Java

Deep Learning with Hadoop

Mastering Machine Learning Algorithms.

Datasets used in machine learning