NOTE
Designing a Degradation System
A historical note on service degradation, when to use it, and several degradation approaches.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is Degradation?
- Degradation is a fallback plan, a best-effort measure taken after the system encounters a failure.
- Compared with rate limiting and circuit breaking, which lean toward technical measures, degradation leans toward the business layer.
2. When Is Degradation Needed?
- Use load testing to understand the system QPS. Degradation is needed when actual QPS exceeds system QPS.
3. What Should Be Degraded?
- Consider whether it is a core path or non-core path and the impact after degradation.
4. How to Degrade
4.1. Stop Secondary Features
- Use switches to stop some unimportant services or switch feature states between versions, giving CPU, memory, or data resources to more important features.
- For example, export functions and scheduled tasks.
4.2. Simplify Features
- An aggregation service no longer returns the full data set and only returns part of the data.
4.3. Reduce Consistency
- Use asynchronous processing to simplify the flow.
- Use cache on errors.
- Reduce quality. For example, if the CDN cannot handle the load in a live-streaming scenario, lower the video quality.
5. How to Implement Degradation
5.1. Switch
5.2. Playbook
5.3. Degradation Component
Designing a Fault-Tolerance Component.md
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub