NOTE

Designing a Degradation System

A historical note on service degradation, when to use it, and several degradation approaches.

System DesignCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is Degradation?

  • Degradation is a fallback plan, a best-effort measure taken after the system encounters a failure.
  • Compared with rate limiting and circuit breaking, which lean toward technical measures, degradation leans toward the business layer.

2. When Is Degradation Needed?

  • Use load testing to understand the system QPS. Degradation is needed when actual QPS exceeds system QPS.

3. What Should Be Degraded?

  • Consider whether it is a core path or non-core path and the impact after degradation.

4. How to Degrade

4.1. Stop Secondary Features

  • Use switches to stop some unimportant services or switch feature states between versions, giving CPU, memory, or data resources to more important features.
    • For example, export functions and scheduled tasks.

4.2. Simplify Features

  • An aggregation service no longer returns the full data set and only returns part of the data.

4.3. Reduce Consistency

  • Use asynchronous processing to simplify the flow.
  • Use cache on errors.
  • Reduce quality. For example, if the CDN cannot handle the load in a live-streaming scenario, lower the video quality.

5. How to Implement Degradation

5.1. Switch

5.2. Playbook

5.3. Degradation Component

Designing a Fault-Tolerance Component.md

6. Reference

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub