NOTE
1.15 Elasticsearch Concurrency Control
English translation of the original VNote ‘Elasticsearch Concurrency Control’, preserving its examples, structure, and historical notes.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is Elasticsearch’s Concurrency-Control Mechanism?
- Multiple requests modifying the same data at the same time can cause concurrency problems (lost updates), so a concurrency-control mechanism is needed.
- There are two kinds of concurrency-control mechanisms: pessimistic locking and optimistic locking.
- Elasticsearch uses an optimistic-locking mechanism: immutable segment files + version numbers.
- Immutable segment files:
- Like Java String, immutability means there is no concurrency problem. Databases solve this through row locks.
- Version numbers:
- Used to solve request-ordering problems.
- Databases work similarly. Refer to Database Optimistic Locking and Pessimistic Locking.md (related note not yet public): first query the data and version number, then compare the version number during update. If they match, update; otherwise return an error to the user and let them retry.
- Before Elasticsearch 6.7,
versionwas used; afterward,seq_no + primary_termis used.
- Immutable segment files:
- Elasticsearch uses an optimistic-locking mechanism: immutable segment files + version numbers.
2. Why Does Elasticsearch Need Concurrency Control?
- On one hand, it solves ordering problems when data is concurrently modified on a single node.
- Concurrent modification in systems such as MySQL has the same problem. Refer to Database Optimistic Locking and Pessimistic Locking.md (related note not yet public).
- On the other hand, it solves ordering problems when the primary shard replicates to replica shards.
- After an Elasticsearch primary shard finishes a CUD operation, it needs to replicate it to replica shards in parallel and asynchronously. Because of the network, replication requests may arrive at replica shards out of order.
- To prevent an old replication request from arriving later than a new replication request and causing an old version of a document to overwrite a new version, a concurrency-control mechanism is needed.
- Why don’t Kafka, MySQL, and Redis have this problem?
- Redis, Kafka, and MySQL use log-based replication. Logs have an order, so this problem does not exist.
- Elasticsearch sends synchronization requests from the primary shard to replica shards in parallel. Network requests cannot guarantee ordering, so this problem exists.
3. How to Use Elasticsearch Concurrency Control
3.1. Elasticsearch 6.7 version Mechanism
- Generate a version.
- A version is automatically generated when creating a Document.
PUT /website/blog/1/_create { "title": "My first blog entry", "text": "Just trying this out..." } GET /website/blog/1 { "_index" : "website", "_type" : "blog", "_id" : "1", "_version" : 1, "found" : true, "_source" : { "title": "My first blog entry", "text": "Just trying this out..." } }
- A version is automatically generated when creating a Document.
- Version-conflict detection.
- Internal version: update only when the version matches.
PUT /website/blog/1?version=1 { "title": "My first blog entry", "text": "Starting to get the hang of this..." } - External version: update only when the version is greater than the original version.
PUT /website/blog/2?version=5&version_type=external { "title": "My first external blog entry", "text": "Starting to get the hang of this..." }
- Internal version: update only when the version matches.
3.2. Elasticsearch 7 seq_no + primary_term Mechanism
4. Concurrency-Conflict Example
4.1. 409 version conflict
update_by_idanddelete_by_iddo not produce this error, whileupdate_by_queryanddelete_by_querymay produce it.- Two
update_by_idoperations are fine.- Disable refresh.
PUT /tb_item/_settings { "index" : { "refresh_interval" : -1 } }- Update by ID.
POST /tb_item/_update/968185 { "doc": { "sellPoint": "test1" } }- Update again by ID.
POST /tb_item/_update/968185 { "doc": { "sellPoint": "test2" } }- Manually refresh.
POST /tb_item/_refresh- Result.
GET /tb_item/_search { "query": { "ids" : { "values" : ["968185"] } } } - First
update_by_id, then no refresh, thendelete_by_queryreports an error.- Disable refresh.
PUT /tb_item/_settings { "index" : { "refresh_interval" : -1 } }- Update by ID.
POST /tb_item/_update/968185 { "doc": { "sellPoint": "test1" } }- Then run
delete_by_query.
POST /tb_item/_delete_by_query { "query": { "match": { "id": "968185" } } }- Result.
{ "took": 1, "timed_out": false, "total": 1, "deleted": 0, "batches": 1, "version_conflicts": 1, "noops": 0, "retries": { "bulk": 0, "search": 0 }, "throttled_millis": 0, "requests_per_second": -1, "throttled_until_millis": 0, "failures": [ { "index": "tb_item", "type": "_doc", "id": "968185", "cause": { "type": "version_conflict_engine_exception", "reason": "[968185]: version conflict, required seqNo [1000], primary term [7]. current document has seqNo [1001] and primary term [7]", "index_uuid": "jOSL4QJNTkWa7p1Z9s-ztQ", "shard": "0", "index": "tb_item" }, "status": 409 } ] }
- Two
4.1.1. Cause Analysis
- First perform the update-by-ID operation.
- Refer to Elasticsearch CRUD Flow.
- At this point, the result is that the document has been written into the memory buffer but has not been refreshed into the filesystem cache. At the same time, it has been written into the translog and fsynced to disk.
- Then perform
delete_by_query.- Refer to Elasticsearch CRUD Flow.
- At this point, the primary node obtains the old document’s version + id and initiates deletion. When deleting from the translog, it finds that the versions do not match, so an error is reported.
4.1.2. Solution
- Force a refresh between the update-by-ID and
delete_by_queryoperations withPOST /tb_item/_refresh, ensuring thatdelete_by_queryobtains the latest version of the data.
5. References
- Elasticsearch delete_by_query 409 version conflict - Elastic Stack / Elasticsearch - Discuss the Elastic Stack
- ElasticSearch6.7使用_delete_by_query产生版本冲突(version conflict)问题_学无止境,永不停歇-CSDN博客
- 乐观并发控制 | Elasticsearch: 权威指南 | Elastic
- 并发控制_百度百科
- elasticsearch学习笔记(十三)——Elasticsearch乐观锁并发控制实战 - SegmentFault 思否
- Optimistic concurrency control | Elasticsearch Guide [7.13] | Elastic
- Relation between “_version”, “_seq_no” and “_primary_term” - Elastic Stack / Elasticsearch - Discuss the Elastic Stack
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub