NOTE

How to Design Distributed Tracing

What tracing is, why it is needed, Span/Trace/Annotation data, sampling, collection approaches, storage, display, and common tracing components.

Software Architecture & EngineeringCreated Updated 3 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is Tracing

Record the complete process of a request. It is used to see which nodes the request passed through and how much time each node consumed.

2. Why Tracing Is Needed

It can quickly locate which node in the entire call chain has a problem for a request.

3. How to Implement Tracing

3.1. What Data to Report

Tracing

3.1.1. Span

3.1.1.1. What Is a Span
  • A Span is generated every time a module is called. Each Span contains the Span name, Span ID, and parent Span ID.
    • The Span ID tells you the call order.
    • The parent Span ID tells you the hierarchy.
3.1.1.2. Why a Span Is Needed
  • A call chain passes through multiple modules. A Span is used to identify a request to one module.

3.1.2. Trace

3.1.2.1. What Is a Trace
  • A Trace is generated at the beginning of each call. The Trace contains a Trace ID.
    • The Trace ID can be used to track the complete request.
    • Trace and Span have a one-to-many relationship, and Spans have parent-child relationships with one another.
3.1.2.2. Why a Trace Is Needed
  • A call chain passes through multiple modules. A Trace is used to identify the entire call chain.

3.1.3. Annotation

3.1.3.1. What Is an Annotation
  • Custom additional data on each Span, such as timestamps.
    • cs: Client Sent. The timestamp when the client initiates a request.
    • sr: Server Received. The timestamp when the server receives the request. sr-cs = network latency.
    • ss: Server Sent. The timestamp when the server finishes processing the request and returns it to the client. ss-sr = time needed by the server to process the request.
    • cr: Client Received. The timestamp when the client successfully receives the server response. cr-cs = time needed by the client to obtain the response from the server.
3.1.3.2. Why Annotation Is Needed
  • Store extended information for a Span.

3.1.4. Service Information

  • IP, Host, service name, interface name, whether it is callee or caller.

3.2. Data Sampling Rate

  • Dynamically adjust the sampling threshold according to traffic volume. For example, it can be written into remote configuration.

3.3. How to Obtain Data

3.3.1. Log-Based Tracing

  • Output Trace, Span, and other information to application logs. After aggregating logs from all machines, infer the call chain.
    • For example, Spring Cloud Sleuth.
  • Advantage: low invasiveness and high performance.
  • Disadvantage: logs may be lost, resulting in poor accuracy.

3.3.2. Service-Based Tracing

  • Inject a probe into the service by some means. The probe can monitor the service and collect data, then send it to the tracing system through HTTP or RPC calls.
    • If the RPC framework supports Filters, write a callee + caller Filter and extract data from the Header of the RPC request.
    • Examples include Naver Pinpoint and SkyWalking.
  • Advantage: high stability and low invasiveness.
  • Disadvantage: low performance.

3.3.3. Sidecar-Based Tracing

  • A dedicated solution for service mesh.
    • For example, Envoy.
  • Advantage: high stability, high performance, and low invasiveness.
  • Disadvantage: not widely adopted.

3.4. Where Is the Data Stored?

  • MySQL.
  • Tracing - Page 2

3.5. How Is the Data Displayed?

4. Tracing Components

4.1. Google Dapper

4.2. Open Zipkin

4.3. CAT

  • Client
  • Server

4.4. Open Tracing

  • The CNCF Technical Oversight Committee defined an API standard for call chains.

4.5. Open Census

  • A call-chain API standard defined by Google and supported by Microsoft.

4.6. Open Telemetry

  • The CNCF Technical Oversight Committee and Google reconciled and defined a call-chain API standard.

4.7. Spring Cloud Sleuth

4.8. CAT vs. Zipkin vs. Pinpoint

5. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub