Learning Apache Cassandra

Book Image

Learning Apache Cassandra

By : Matthew Brown

4 (1)

Book Image

Learning Apache Cassandra

4 (1)

By: Matthew Brown

Overview of this book

Learning Apache Cassandra

Learning Apache Cassandra

Credits

About the Author

About the Author

About the Reviewers

About the Reviewers

www.PacktPub.com

www.PacktPub.com

Preface

Free Chapter

Getting Up and Running with Cassandra

Getting Up and Running with Cassandra

What Cassandra offers, and what it doesn't

Installing Cassandra

Bootstrapping the project

Creating a keyspace

The First Table

The First Table

Creating the users table

Developing a mental model for Cassandra

Organizing Related Data

Organizing Related Data

A table for status updates

Working with status updates

Anatomy of a compound primary key

Beyond two columns

Compound keys represent parent-child relationships

Coupling parents and children using static columns

Refining our mental model

Beyond Key-Value Lookup

Beyond Key-Value Lookup

Looking up rows by partition

Retrieving status updates for a specific time range

Paginating over rows in a partition

Reversing the order of rows

Paginating over multiple partitions

Building an autocomplete function

Establishing Relationships

Establishing Relationships

Modeling follow relationships

Storing follow relationships

Looking up follow relationships

Unfollowing users

Using secondary indexes to avoid denormalization

Denormalizing Data for Maximum Performance

Denormalizing Data for Maximum Performance

A normalized approach

Partial denormalization

Fully denormalizing the home timeline

Write complexity and data integrity

Expanding Your Data Model

Expanding Your Data Model

Viewing a table schema in cqlsh

Adding columns to tables

Deleting columns

Updating the existing rows

Removing a value from a column

Inserts, updates, and upserts

Lightweight transactions have a cost

Collections, Tuples, and User-defined Types

Collections, Tuples, and User-defined Types

The problem with concurrent updates

Collection columns and concurrent updates

Using lists for ordered, nonunique values

Using maps to store key-value pairs

Collections in inserts

Collections and secondary indexes

The limitations of collections

Working with tuples

User-defined types

Choosing between tuples and user-defined types

Comparing data structures

Aggregating Time-Series Data

Aggregating Time-Series Data

Recording discrete analytics observations

Recording aggregate analytics observations

Recording analytics observations

How Cassandra Distributes Data

How Cassandra Distributes Data

Data distribution in Cassandra

Data replication in Cassandra

Handling conflicting data

Distributed deletion

Peeking Under the Hood

Peeking Under the Hood

Using cassandra-cli

The structure of a simple primary key table

Compound primary keys in column families

Collection columns in column families

Authentication and Authorization

Authentication and Authorization

Enabling authentication and authorization

Setting up a user

Controlling access

Authorization in action

Security beyond authentication and authorization

Index

Customer Reviews

4 (1)

5 star

0

4 star

100%

3 star

0

2 star

0

1 star

0

Partial denormalization

Our initial approach to home timelines, which used the existing, fully normalized data structure that we've already built, is technically viable but will perform very poorly at scale. If I follow F users and want a page of size P for my home timeline, Cassandra will need to do the following:

Query F partitions for P rows each
Perform an ordered merge of FxP rows in order to retrieve only the most recent P

The most distressing part of this is the fact that both operations grow in complexity proportionally with the number of people I follow. Let's start by trying to fix this.

The basic goal of the home timeline is to show me the most recent status updates that matter to me. Instead of doing all the work to find out what status updates matter to me, based on who I follow, at read time, let's shift some of the work to write time.

I'll create a table that stores references to status updates that I care about. Whenever someone I follow creates a new status update, I'll add...