ElasticSearch Cookbook

ElasticSearch Cookbook

By : Alberto Paro

Buy this Book

ElasticSearch Cookbook

By: Alberto Paro

Buy this Book

Overview of this book

ElasticSearch is one of the most promising NoSQL technologies available and is built to provide a scalable search solution with built-in support for near real-time search and multi-tenancy. This practical guide is a complete reference for using ElasticSearch and covers 360 degrees of the ElasticSearch ecosystem. We will get started by showing you how to choose the correct transport layer, communicate with the server, and create custom internal actions for boosting tailored needs. Starting with the basics of the ElasticSearch architecture and how to efficiently index, search, and execute analytics on it, you will learn how to extend ElasticSearch by scripting and monitoring its behaviour. Step-by-step, this book will help you to improve your ability to manage data in indexing with more tailored mappings, along with searching and executing analytics with facets. The topics explored in the book also cover how to integrate ElasticSearch with Python and Java applications. This comprehensive guide will allow you to master storing, searching, and analyzing data with ElasticSearch.

ElasticSearch Cookbook

Credits

About the Author

About the Reviewers

www.PacktPub.com

Preface

Free Chapter

Getting Started

Introduction

Understanding node and cluster

Understanding node services

Managing your data

Understanding cluster, replication, and sharding

Communicating with ElasticSearch

Using the HTTP protocol

Using the Native protocol

Using the Thrift protocol

Downloading and Setting Up ElasticSearch

Introduction

Downloading and installing ElasticSearch

Networking setup

Setting up a node

Setting up ElasticSearch for Linux systems (advanced)

Setting up different node types (advanced)

Installing a plugin

Installing a plugin manually

Removing a plugin

Changing logging settings (advanced)

Managing Mapping

Introduction

Using explicit mapping creation

Using dynamic templates in document mapping

Managing nested objects

Managing a child document

Mapping a multifield

Mapping a GeoPoint field

Mapping a GeoShape field

Mapping an IP field

Mapping an attachment field

Adding generic data to mapping

Mapping different analyzers

Standard Operations

Introduction

Creating an index

Deleting an index

Opening/closing an index

Putting a mapping in an index

Checking if an index or type exists

Managing index settings

Speeding up atomic operations (bulk)

Speeding up GET

Search, Queries, and Filters

Executing a scan query

Suggesting a correct query

Counting

Deleting by query

Matching all the documents

Querying/filtering for term

Querying/filtering for terms

Using a prefix query/filter

Using a Boolean query/filter

Using a range query/filter

Using span queries

Using the match query

Using the IDS query/filter

Using the has_child query/filter

Using the top_children query

Using the has_parent query/filter

Using a regexp query/filter

Using exists and missing filters

Using and/or/not filters

Using the geo_bounding_box filter

Using the geo_polygon filter

Using the geo_distance filter

Facets

Introduction

Executing facets

Executing terms facets

Executing range facets

Executing histogram facets

Executing date histogram facets

Executing filter/query facets

Executing statistical facets

Executing term statistical facets

Executing geo distance facets

Scripting

Introduction

Installing additional script plugins

Sorting using script

Computing return fields with scripting

Filtering a search via scripting

Updating with scripting

Rivers

Introduction

Managing a river

Using the CouchDB river

Using the MongoDB river

Using the RabbitMQ river

Using the JDBC river

Using the Twitter river

Cluster and Nodes Monitoring

Introduction

Controlling cluster health via API

Controlling cluster state via API

Getting nodes information via API

Getting node statistic via API

Installing and using BigDesk

Installing and using ElasticSerach-head

Installing and using SemaText SPM

Java Integration

Introduction

Creating an HTTP client

Creating a native client

Managing indices with the native client

Executing a standard search

Executing a facet search

Executing a scroll/scan search

Python Integration

Executing a standard search

Executing a facet search

Plugin Development

Introduction

Creating a site plugin

Creating a simple plugin

Creating a REST plugin

Creating a cluster action

Creating an analyzer plugin

Creating a river plugin

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Managing your data

Unless you are using ElasticSearch as a search engine or a distributed data store, it's important to understand concepts on how ElasticSearch stores and manages your data.

Getting ready

To work with ElasticSearch data, a user must know basic concepts of data management and JSON that is the "lingua franca" for working with ElasticSearch data and services.

How it works...

Our main data container is called index (plural indices) and it can be considered as a database in the traditional SQL world. In an index, the data is grouped in data types called mappings in ElasticSearch. A mapping describes how the records are composed (called fields).

Every record, that must be stored in ElasticSearch, must be a JSON object.

Natively, ElasticSearch is a schema-less datastore. When you put records in it, during insert it processes the records, splits them into fields, and updates the schema to manage the inserted data.

To manage huge volumes of records, ElasticSearch uses the common approach to split an index into many shards so that they can be spread on several nodes. The shard management is transparent in usage—all the common record operations are managed automatically in the ElasticSearch application layer.

Every record is stored in only one shard. The sharding algorithm is based on record ID, so many operations that require loading and changing of records can be achieved without hitting all the shards.

The following schema compares ElasticSearch structure with SQL and MongoDB ones:

ElasticSearch	SQL	MongoDB
Index (Indices)	Database	Database
Shard	Shard	Shard
Mapping/Type	Table	Collection
Field	Field	Field
Record (JSON object)	Record (Tuples)	Record (BSON object)

There's more...

ElasticSearch, internally, has rigid rules about how to execute operations to ensure safe operations on index/mapping/records. In ElasticSearch, the operations are divided as follows:

Cluster operations: At cluster level all write ones are locked, first they are applied to the master node and then to the secondary one. The read operations are typically broadcasted.
Index management operations: These operations follow the cluster pattern.
Record operations: These operations are executed on single documents at shard level.

When a record is saved in ElasticSearch, the destination shard is chosen based on the following factors:

The ID (unique identifier) of the record. If the ID is missing, it is autogenerated by ElasticSearch.
If the routing or parent (covered while learning the parent/child mapping) parameters are defined, the correct shard is chosen by the hash of these parameters.

Splitting an index into shards allows you to store your data in different nodes, because ElasticSearch tries to do shard balancing.

Every shard can contain up to 2^32 records (about 4.2 billion records), so the real limit to shard size is its storage size.

Shards contain your data and during search process all the shards are used to calculate and retrieve results. ElasticSearch performance in big data scales horizontally with the number of shards.

All native records operations (such as index, search, update, and delete) are managed in shards.

The shard management is completely transparent to the user. Only an advanced user tends to change the default shard routing and management to cover their custom scenarios. A common custom scenario is the requirement to put customer data in the same shard to speed up his/her operations (search/index/analytics).

Best practice

It's best practice not to have a too big shard (over 10 GB) to avoid poor performance in indexing due to continuous merge and resizing of index segments.

It's not good to oversize the number of shards to avoid poor search performance due to native distributed search (it works as MapReduce). Having a huge number of empty shards in an index consumes only memory.

ElasticSearch Cookbook

By : Alberto Paro

ElasticSearch Cookbook

By: Alberto Paro

Overview of this book

Related Content you might be interested in

Current Title:

ElasticSearch Cookbook

Managing your data

Getting ready

How it works...

There's more...

Best practice

See also