NOTE

Designing a Frequency-Control System

A sanitized historical note preserving atomic Redis counters, batch APIs, hot-key handling, degradation, and performance-testing methodology.

System DesignCreated Updated 3 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is the Requirement?

General frequency-control service refactoring (internal requirement link removed).

2. Why Build It?

The original business system went through several rounds of business integration and evolution. It accumulated problems in architecture, extensibility, reuse, and maintainability, so the frequency-control logic was extracted into a shared capability.

3. Design

3.1. Storage Design

Use Redis. Store a string value in the form business_id_key -> count.

3.2. API Design

enum ErrCode {
  ERR_CODE_OK = 0;
  ERR_CODE_TIMEOUT = 10001;
  ERR_CODE_KEY_NOT_FOUND = 10002;
}

enum TimeUnit {
    TIME_UNIT_SEC = 0;
    TIME_UNIT_DAY = 1;
}

message AddFreqReq {
    uint32 business_id = 1;
    string key = 2;
    uint32 time_span = 3;
    TimeUnit time_unit = 4;
}

message AddFreqRsp {
    int32 code = 1;
    string msg = 2;
    uint32 business_id = 3;
    string key = 4;
    uint32 freq = 5;
    uint32 ttl = 6;
}

message BatchAddFreqReq { repeated AddFreqReq requests = 1; }
message BatchAddFreqRsp { repeated AddFreqRsp results = 1; }
message GetFreqReq { uint32 business_id = 1; string key = 2; }
message GetFreqRsp {
    int32 code = 1;
    string msg = 2;
    uint32 business_id = 3;
    string key = 4;
    uint32 freq = 5;
    uint32 ttl = 6;
}
message BatchGetFreqReq { repeated GetFreqReq requests = 1; }
message BatchGetFreqRsp { repeated GetFreqRsp results = 1; }

service FrequencyControl {
    rpc AddFreq(AddFreqReq) returns (AddFreqRsp);
    rpc BatchAddFreq(BatchAddFreqReq) returns (BatchAddFreqRsp);
    rpc GetFreq(GetFreqReq) returns (GetFreqRsp);
    rpc BatchGetFreq(BatchGetFreqReq) returns (BatchGetFreqRsp);
}

3.3. Architecture Design

3.3.1. Architecture Diagram

Frequency-control architecture

3.3.2. Sequence Diagram

Frequency-control sequence

3.3.3. Batch API Design

  1. Limit request length.
  2. Validate all parameters. If one validation fails, fail the whole batch before partial execution.
  3. Define partial-failure behavior: ignore failed items, fail the whole batch, or provide rollback.

Reuse the single-request req/rsp structures to reduce cognitive load. The batch response can contain an outer code/msg and per-item code/msg.

3.3.4. Custom code/msg Instead of error

Partial failures need to be represented.

3.3.5. Explicit Validation

The whole batch still needs overall validation.

3.3.6. Internal Error Codes + External Error Codes

3.3.7. Idempotency Design

An order/request number can be generated by the caller to make retried writes idempotent. Whether it is necessary depends on how strict the business semantics are.

3.3.8. Degradation Design

Write degradation is decided by the business side. Read network errors may fall back to a short-lived local cache when stale values are acceptable.

3.3.9. Why Wrap Lua

The counter update and TTL initialization need to be atomic.

3.3.10. Redis Hot Keys

  • Detect hot keys through storage-side Top-N monitoring and periodic service-side polling.
  • For hot read keys, nodes may keep a very short-TTL local copy.
  • Strict write counters must not be independently cached across nodes.

4. Performance Testing

For an I/O-heavy service, QPS can be roughly estimated by QPS ≈ concurrency / response_time, but final capacity must come from benchmarking.

First test: debug logging enabled.
Second test: debug logging disabled; performance improves noticeably.
Third test: replace Eval with EvalSha; improvement is limited.

To raise QPS, either reduce latency or increase effective concurrency. Use flame graphs, trace, packet capture, and Redis-side metrics together.

4.1. AddFreq

  1. First test: CPU becomes the bottleneck; flame graphs show GC and Redis-client wrapper overhead.
  2. Adjust the GC strategy and test again; GC share decreases.
  3. Switch Redis clients and test again; client-wrapper overhead decreases.
  4. Try network-layer optimization; improvement is limited.
  5. Benchmark Redis directly and confirm the final bottleneck is Redis CPU.

The public version does not retain internal performance-plugin names, machine specifications, QPS numbers, or production monitoring screenshots.

4.2. BatchAddFreq

  1. Different Redis clients show little throughput difference.
  2. Application CPU, memory, and disk are not saturated, so investigate the network.
  3. Use trace to observe network waiting.
  4. Check latency, bandwidth, and request-packet size.
  5. Use tcpdump / Wireshark to inspect batches and reduce Lua-script transfer size.
  6. Cross-check with a local Redis and redis-benchmark; the main bottleneck is still single-node Redis CPU rather than application CPU or link bandwidth.

One important lesson: do not rely only on averaged CPU in proxy/replication architectures, because an average can hide a saturated node.

4.3. GetFreq

Same as AddFreq.

4.4. BatchGetFreq

Same as BatchAddFreq.

5. Deployment

A common call sequence is:

  1. Read frequency-control state.
  2. Process business logic.
  3. Write frequency-control state.

Step 3 can be asynchronous; the key optimization target is read latency.

Two common Redis multi-region replication models are:

  1. Leader-Leader: low local read/write latency, but conflicting writes are harder to reconcile.
  2. Leader-Follower: writes go to a leader and reads use nearby followers; suitable for read-heavy workloads that tolerate replication delay.

Whether a follower can be read depends on consistency requirements. Strict limits may require a strongly consistent source; protective thresholds that tolerate small error may use a nearby replica.

6. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub