NOTE

Designing a Highly Available System

A historical note on high availability across design, development, testing, release, and operations.

System DesignCreated Updated 2 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is a Highly Available System?

  • A system with high availability.
    • Availability means that when part of the system encounters a failure, we can anticipate and handle that failure. In essence, this is fault tolerance.
    • High availability:
      • availability = mean time between failures / (mean time between failures + mean time to recovery)
      • Availability is generally expressed in “nines”. 99.99% is common; more nines represent higher availability.

2. How to Design a Highly Available System

2.1. Design Stage

2.1.1. Design Review

  • Compare solutions.

2.2. Development Stage

2.2.1. Cache

2.2.2. Degradation

2.2.3. Rate Limiting

2.2.4. Isolation

2.2.5. Circuit Breaking

2.2.6. Timeout and Retry

2.2.7. Pooling

2.3. Testing Stage

2.3.1. Functional Testing

  • Full-process testing should ideally be fully automated. Many test cases can be accumulated for full-process regression, although the system needs to support this. I have seen too many incidents where problems were not discovered in advance because QA did not have enough capacity to perform full-process regression. The testing principle is therefore to automate as much as possible and cover the full process, leaving precious human resources for steps that can only be tested manually.

2.3.2. Unit Tests

2.3.3. Load Testing

2.3.4. Risk Playbooks

  • Regularly imagine all kinds of possible risk points and prepare playbooks for them so that problems can be solved quickly.
    1. Classify and organize common historical production incidents by scenario, such as a MySQL failure playbook, MQ failure playbook, and order-creation API failure playbook.
    2. Validate the playbooks through a fault-injection platform.

2.3.5. Code Review

2.4. Release Stage

2.4.1. Change Control

  • Notify upstream and downstream staff.

2.4.2. Canary Release and Rollback

2.4.3. Avoid Single Points of Failure

  • Deploy services in multiple Regions, and deploy multiple Zones in each Region.
    • Region: a geographical area, such as North China or South China. Different regions are connected through the Internet, not an intranet.
    • Zone: a physical area within the same Region, for example Shanghai and Hangzhou within the same regional grouping. Different Zones are connected through an intranet, not the Internet.

Cloud Database Redis Regions and Availability Zones OPPO Multi-Active Architecture

2.4.3.1. Dual-Active Across Locations
  • Refers to real-time synchronization, so it can only be implemented across Zones within the same Region.
2.4.3.2. Disaster Recovery Across Locations
  • Refers to non-real-time synchronization, so it can be implemented across Zones or Regions.

2.5. Operations Stage

2.5.1. Monitoring

2.5.2. Chaos Engineering

  • How to Design Chaos Engineering

2.5.3. Retrospective

Production incidents must be reviewed promptly. Retrospectives help discover problems in existing processes and mechanisms, prevent everyone from falling into the same pit again, and continuously improve emergency playbooks.

3. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub