NOTE
Designing Timeout and Retry
A historical note on timeout duration, retryable errors, backoff, idempotency, and retry counts.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Are Timeout and Retry?
- Timeout: service A calls service B. To prevent A from waiting forever for B and holding threads and other resources without releasing them, a timeout must be configured.
- Retry: service A calls service B. If A’s request to B encounters a network error, a timeout may occur. Because the network problem may be temporary, retrying several times may succeed.
2. How to Design Timeout and Retry
2.1. Timeout
2.1.1. Timeout Duration
- For a newly launched service, use load testing: when the N-nines latency is 200 ms, take the average latency as the timeout.
- For an existing service, use monitoring: average response time.
2.2. Retry
2.2.1. Which Errors Can Be Retried?
- Call timeout, or a retryable error returned by the callee, such as busy, rate limited, under maintenance, insufficient resources, etc.
2.2.2. Retry Strategy
- Exponential backoff.
- Reset the timeout.
2.2.3. Prerequisite for Retry
- The callee is idempotent.
2.2.4. Retry Count
- Use load testing to calculate the maximum supported concurrency.
- If the upstream passes a time budget, determine the retry count from that time; otherwise use a default value, for example Linux TCP retries five times.
3. Timeout and Retry Component
Designing a Fault-Tolerance Component.md
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub