NOTE
Designing a Leaderboard System
A sanitized historical design note preserving sorted sets, idempotent updates, hotspot handling, sharding, and reconciliation.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is the Requirement?
- Requirement ticket
- Prototype
2. Why Build It?
- What problem does it solve?
- Can the requirement be omitted?
- Can the requirement be simplified?
3. Requirement Analysis
Understanding concepts should be combined with concrete system examples.
3.1. What Does the Existing Flow Look Like?
Ask operations or product to demonstrate the flow. Review read/write flows and back-office / user-facing flows.
3.2. What Does the New Flow Look Like?
- Roles in the system and what each role can do
- What the flow looks like when each role performs an action
3.3. QPS Estimation
Use variables to estimate participant scale and hotspot-object online scale. The public version does not retain real production counts or capacity.
4. Design
4.1. Storage Design
Overwrite: zadd 1_2_3 1 member
Increment or decrement: zincrby 1_2_3 5 member
Read by descending score: zrevrange 1_2_3 0 -1 withscores
4.2. API Design
ReadRankReq {
business_id
rank_scope
start
limit
optional timestamp
optional self_member
optional order_type
}
ReadRankRsp {
rank_list[{member_id, score, rank}]
optional self_rank
optional total
}
UpdateRankReq {
business_id
rank_scope
operations[{member_id, score, op_type}]
request_id // idempotency
}
UpdateRankRsp {
results[{member_id, score}]
}
DeleteRankMemberReq {
business_id
rank_scope
member_ids[]
}
4.2.1. Read Leaderboard Details by Rank
- req
- business_id
- rank_id
- subrank_id
- start
- limit
- me
- rsp
- list
- member
- score
- me
- member
- score
- list
4.2.2. Update Leaderboard
- req
- business_id
- rank_id
- subrank_id
- list
- member
- score
- request_id (idempotency)
- op_type (set / increment)
- rsp
- list
- member
- score
- list
4.3. Architecture Design
Consider security, high concurrency, high availability, and maintainability together.
- Split into microservices
- Draw the application architecture diagram
- Draw the sequence diagram
4.4. Code Design
4.4.1. update
pipeline + EvalSha Lua + fallback Eval Lua.
The original implementation encoded a time-based tie-breaker into the score so that, for equal business scores, the member who reached the score earlier ranks first.
function update(rank_key, member, score, request_id, op_type, ttl):
atomically:
if request_id has been processed:
return previous result
mark request_id processed with ttl
if op_type == SET:
ZADD rank_key score member
else if op_type == INCR:
ZINCRBY rank_key score member
if rank_key is newly created:
EXPIRE rank_key ttl
return ZSCORE rank_key member
4.4.2. get
- getByRank
zrevrange rank_key start start+limit-1 withscores
zrange rank_key start start+limit-1 withscores
- getByMember
function getByMember(rank_key, member, order):
score = ZSCORE(rank_key, member)
if score does not exist:
return not_found
rank = order == ASC ? ZRANK(rank_key, member) : ZREVRANK(rank_key, member)
return rank, score
4.5. Idempotency Design
Add a request_id to writes, generated by the caller, and place the idempotency check in the same atomic operation as the leaderboard update.
4.6. Handling High Concurrency
Leaderboard characteristics:
- Only head data needs to be displayed, not the entire leaderboard.
- The workload is read-heavy and write-light.
4.6.1. High-Concurrency Writes
Use a message queue to absorb bursts. Since the workload is usually read-heavy, writes are not normally the first optimization target.
4.6.2. High-Concurrency Reads
The display path reads Top N, the user’s score and rank, then aggregates profile data.
Profile data is suitable for a read-through local cache. Leaderboard scores change frequently, so score caching requires separate evaluation. Benchmark the single-key path first; if Redis is sufficient, avoid over-optimizing.
4.7. Hot Leaderboards
- Detect hot keys
- Preconfigure known hotspots.
- Monitor Top-N keys by access volume.
- Handle hot keys
- Let Redis handle ordinary hotspots.
- For extreme hotspots, dual-write Redis and a local in-memory read copy.
General flow:
- Optionally disable leaderboards for hotspot objects that do not need them.
- For hotspot objects that do need leaderboards, write Redis first and update local cache through replayable events.
- On node startup, load a snapshot from Redis, replay from the event position, and only serve the local copy after catching up.
- Serve Top N from local cache; query personal rank from Redis.
- Query ordinary leaderboards directly from Redis.
4.8. Very Large Leaderboards
- Detect large keys by configuration or member-count monitoring.
- Split large leaderboards.
- Write: hash the member to a suffix and write to a zset shard.
- Read:
- Option 1: query Top N from all shards in parallel and merge in memory.
- Option 2: periodically merge shards into a read-only Top-N view.
- Personal rank may not appear in Top N, so maintain a separate personal score/rank read model.
4.9. Reconciliation
The business side retains score-change records, and the leaderboard side retains processed-update records or a checkpoint.
Periodically compare a recent safe window. If a difference is found, resend the missing update and rely on request_id for idempotent repair.
5. Effort Estimation
- Estimate one API at roughly 0.5–2 days.
6. Development
7. Testing
- Redis: update throughput, Top-N query throughput, personal-rank query throughput
- Message queue: per-partition consumption throughput and replay speed
- Service: local sorted-set queries, Redis queries, hotspot replay time after startup
8. Release
- Release checklist
- Deployment
- Deployment diagram
machine_count = expected_system_QPS / measured_QPS_per_machine
9. Operations
9.1. Redis Global Replication
Refer to the multi-region replication discussion in the frequency-control note.
10. Optimization
11. Summary
- Compare solutions
- Problems encountered and how they were solved
- Design highlights
- Pain points and improvements
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub