Mastering Distributed Tracing

Book Image

Mastering Distributed Tracing

By : Yuri Shkuro

Book Image

Mastering Distributed Tracing

By: Yuri Shkuro

Overview of this book

Mastering Distributed Tracing will equip you to operate and enhance your own tracing infrastructure. Through practical exercises and code examples, you will learn how end-to-end tracing can be used as a powerful application performance management and comprehension tool. The rise of Internet-scale companies, like Google and Amazon, ushered in a new era of distributed systems operating on thousands of nodes across multiple data centers. Microservices increased that complexity, often exponentially. It is harder to debug these systems, track down failures, detect bottlenecks, or even simply understand what is going on. Distributed tracing focuses on solving these problems for complex distributed systems. Today, tracing standards have developed and we have much faster systems, making instrumentation less intrusive and data more valuable. Yuri Shkuro, the creator of Jaeger, a popular open-source distributed tracing system, delivers end-to-end coverage of the field in Mastering Distributed Tracing. Review the history and theoretical foundations of tracing; solve the data gathering problem through code instrumentation, with open standards like OpenTracing, W3C Trace Context, and OpenCensus; and discuss the benefits and applications of a distributed tracing infrastructure for understanding, and profiling, complex systems.

Mastering Distributed Tracing

Mastering Distributed Tracing

Contributors

Preface

Other Books You May Enjoy

Other Books You May Enjoy

Leave a review - let other readers know what you think

Leave a review - let other readers know what you think

Free Chapter

Why Distributed Tracing?

Why Distributed Tracing?

Microservices and cloud-native applications

What is observability?

The observability challenge of microservices

Traditional monitoring tools

Distributed tracing

My experience with tracing

Take Tracing for a HotROD Ride

Take Tracing for a HotROD Ride

Meet the HotROD

The architecture

Contextualized logs

Span tags versus logs

Identifying sources of latency

Resource usage attribution

Distributed Tracing Fundamentals

Distributed Tracing Fundamentals

Request correlation

Anatomy of distributed tracing

Preserving causality

Clock skew adjustment

Instrumentation Basics with OpenTracing

Instrumentation Basics with OpenTracing

Exercise 1 – the Hello application

Exercise 2 – the first trace

Exercise 3 – tracing functions and passing context

Exercise 4 – tracing RPC requests

Exercise 5 – using baggage

Exercise 6 – auto-instrumentation

Exercise 7 – extra credit

Instrumentation of Asynchronous Applications

Instrumentation of Asynchronous Applications

The Tracing Talk chat application

Instrumenting with OpenTracing

Instrumenting asynchronous code

Tracing Standards and Ecosystem

Tracing Standards and Ecosystem

Styles of instrumentation

Anatomy of tracing deployment and interoperability

Five shades of tracing

Know your audience

Tracing with Service Meshes

Tracing with Service Meshes

Observability via a service mesh

The Hello application

Distributed tracing with Istio

Using Istio to generate a service graph

Distributed context and routing

All About Sampling

All About Sampling

Head-based consistent sampling

Tail-based consistent sampling

Partial sampling

Turning the Lights On

Turning the Lights On

Tracing as a knowledge base

Performance analysis

Long-term profiling

Distributed Context Propagation

Distributed Context Propagation

Brown Tracing Plane

Chaos engineering

Traffic labeling

Integration with Metrics and Logs

Integration with Metrics and Logs

Three pillars of observability

The Hello application

Integration with metrics

Integration with logs

Gathering Insights with Data Mining

Gathering Insights with Data Mining

Feature extraction

Components of a data mining pipeline

Feature extraction exercise

The Span Count job

Observing trends

Historical analysis

Ad hoc analysis

Implementing Tracing in Large Organizations

Implementing Tracing in Large Organizations

Why is it hard to deploy tracing instrumentation?

Reduce the barrier to adoption

Building the culture

Tracing Quality Metrics

Troubleshooting guide

Don't be on the critical path

Under the Hood of a Distributed Tracing System

Under the Hood of a Distributed Tracing System

Why host your own?

Bet on emerging standards

Architecture and deployment modes

Monitoring and troubleshooting

Afterword

Index

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Components of a data mining pipeline

There are probably many ways of building near real-time data mining for traces. In Canopy, the feature extraction functionality is built directly into the tracing backend, whereas in Jaeger, it can be done via post-processing add-ons, as we will do in this chapter's code exercise. Major components that are required are shown in Figure 12.1:

Tracing backend, or tracing infrastructure in general, collects tracing data from the microservices of the distributed application
Trace completion trigger makes a judgement call that all spans of the trace have been received and it is ready for processing
Feature extractor performs the actual calculations on each trace
An optional Aggregator combines features from individual traces into an even smaller dataset
Storage records the results of the calculations or aggregations

Figure 12.1: High-level architecture of a data mining pipeline

In the following sections, I will go into detail about the responsibilities of each...