Mastering Java Machine Learning

Mastering Java Machine Learning

By : Uday Kamath, Krishna Choppella

Buy this Book

Mastering Java Machine Learning

By: Uday Kamath, Krishna Choppella

Buy this Book

Overview of this book

Java is one of the main languages used by practicing data scientists; much of the Hadoop ecosystem is Java-based, and it is certainly the language that most production systems in Data Science are written in. If you know Java, Mastering Machine Learning with Java is your next step on the path to becoming an advanced practitioner in Data Science. This book aims to introduce you to an array of advanced techniques in machine learning, including classification, clustering, anomaly detection, stream learning, active learning, semi-supervised learning, probabilistic graph modeling, text mining, deep learning, and big data batch and stream machine learning. Accompanying each chapter are illustrative examples and real-world case studies that show how to apply the newly learned techniques using sound methodologies and the best Java-based tools available today. On completing this book, you will have an understanding of the tools and techniques for building powerful machine learning models to solve data science problems in just about any domain.

Mastering Java Machine Learning

Credits

Foreword

About the Authors

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Machine Learning Review

Machine learning – history and definition

What is not machine learning?

Machine learning – concepts and terminology

Machine learning – types and subtypes

Datasets used in machine learning

Machine learning applications

Practical issues in machine learning

Machine learning – roles and process

Machine learning – tools and datasets

Summary

Practical Approach to Real-World Supervised Learning

Formal description and notation

Data transformation and preprocessing

Feature relevance analysis and dimensionality reduction

Model building

Model assessment, evaluation, and comparisons

Case Study – Horse Colic Classification

Summary

References

Unsupervised Machine Learning Techniques

Issues in common with supervised learning

Issues specific to unsupervised learning

Feature analysis and dimensionality reduction

Clustering

Outlier or anomaly detection

Real-world case study

Summary

References

Semi-Supervised and Active Learning

Semi-supervised learning

Active learning

Case study in active learning

Summary

References

Real-Time Stream Machine Learning

Assumptions and mathematical notations

Basic stream processing and computational techniques

Concept drift and drift detection

Incremental supervised learning

Incremental unsupervised learning using clustering

Unsupervised learning using outlier detection

Case study in stream learning

Summary

References

Probabilistic Graph Modeling

Probability revisited

Graph concepts

Bayesian networks

Markov networks and conditional random fields

Summary

Deep Learning

Multi-layer feed-forward neural network

Limitations of neural networks

Deep learning

Case study

Summary

References

Text Mining and Natural Language Processing

NLP, subfields, and tasks

Issues with mining unstructured data

Text processing components and transformations

Topics in text mining

Tools and usage

Summary

References

Big Data Machine Learning – The Final Frontier

What are the characteristics of Big Data?

Big Data Machine Learning

Batch Big Data Machine Learning

Case study

Linear Algebra

Vector

Matrix

Probability

Axioms of probability

Bayes' theorem

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Feature analysis and dimensionality reduction

Among the first tools to master are the different feature analysis and dimensionality reduction techniques. As in supervised learning, the need for reducing dimensionality arises from numerous reasons similar to those discussed earlier for feature selection and reduction.

A smaller number of discriminating dimensions makes visualization of data and clusters much easier. In many applications, unsupervised dimensionality reduction techniques are used for compression, which can then be used for transmission or storage of data. This is particularly useful when the larger data has an overhead. Moreover, applying dimensionality reduction techniques can improve the scalability in terms of memory and computation speeds of many algorithms.

Notation

We will use similar notation to what was used in the chapter on supervised learning. The examples are in d dimensions and are represented as vector:

x = (x₁,x₂,…x_d )^T

The entire dataset containing n examples can...

Mastering Java Machine Learning

By : Uday Kamath, Krishna Choppella

Mastering Java Machine Learning

By: Uday Kamath, Krishna Choppella

Overview of this book

Related Content you might be interested in

Current Title:

Mastering Java Machine Learning

Machine Learning in Java

Deep Learning with Hadoop

Mastering Machine Learning Algorithms.

Feature analysis and dimensionality reduction

Notation