NOTE

Designing Timeout and Retry

A historical note on timeout duration, retryable errors, backoff, idempotency, and retry counts.

System DesignCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Are Timeout and Retry?

  • Timeout: service A calls service B. To prevent A from waiting forever for B and holding threads and other resources without releasing them, a timeout must be configured.
  • Retry: service A calls service B. If A’s request to B encounters a network error, a timeout may occur. Because the network problem may be temporary, retrying several times may succeed.

2. How to Design Timeout and Retry

2.1. Timeout

2.1.1. Timeout Duration

  • For a newly launched service, use load testing: when the N-nines latency is 200 ms, take the average latency as the timeout.
  • For an existing service, use monitoring: average response time.

2.2. Retry

2.2.1. Which Errors Can Be Retried?

  • Call timeout, or a retryable error returned by the callee, such as busy, rate limited, under maintenance, insufficient resources, etc.

2.2.2. Retry Strategy

  • Exponential backoff.
  • Reset the timeout.

2.2.3. Prerequisite for Retry

  • The callee is idempotent.

2.2.4. Retry Count

  • Use load testing to calculate the maximum supported concurrency.
  • If the upstream passes a time budget, determine the retry count from that time; otherwise use a default value, for example Linux TCP retries five times.

3. Timeout and Retry Component

Designing a Fault-Tolerance Component.md

4. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub