NOTE

1.10 Using Elasticsearch

English translation of the original VNote ‘Using Elasticsearch’, preserving its structure, examples, code, and historical learning notes.

Elasticsearch / SearchCreated Updated 5 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. Data Types

1.1. String

  • text: will be tokenized.
  • keyword: will not be tokenized. Equivalent to the old not_analyzed.

1.2. Numeric

long, integer, short, byte, double, float

1.3. Date

date: 2015-01-01, 2015-01-01T12:10:30Z, 1420070400001

1.4. Boolean

boolean

2. Using the RESTful API

2.1. Cluster Management

2.1.1. Check Cluster Health

GET /_cat/health?v

epoch      timestamp cluster       status node.total node.data shards pri relo init unassign pending_tasks max_task_wait_time active_shards_percent
1546235661 13:54:21  elasticsearch yellow          1         1      1   1    0    0        1             0                  -                 50.0%
  • Explanation:
    • cluster: cluster name.
    • status
      • green: the primary shard and replica shard of every index are active.
      • yellow: the primary shards of every index are active, but some replica shards are not active and are unavailable.
      • red: not all primary shards of all indices are active; some indices have lost data.
    • node.total: number of master + data nodes.
    • node.data: number of data nodes.
    • unassign: number of unassigned shards.
    • active_shards_percent: percentage of available shards.

2.1.2. View Which Indices Exist in the Cluster

GET /_cat/indices?v

health status index   uuid                   pri rep docs.count docs.deleted store.size pri.store.size
yellow open   .kibana id1SV_oGSjyGosKxeJApww   1   1          1            0      3.1kb          3.1kb

2.2. Index Operations

2.2.1. Create an Index

PUT /test_index?pretty

2.2.2. Delete an Index

// Delete one
DELETE /my_index
// Delete multiple
DELETE /index_one,index_two
// Delete by wildcard
DELETE /index_*
// Delete all
DELETE /_all

2.3. Mapping Operations

2.3.1. Create a Mapping

PUT /my_index
{
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 0
  },
  "mappings": {
    "my_type": {
      "properties": {
        "my_field": {
          "type": "text"
        }
      }
    }
  }
}

2.3.2. Customize the Dynamic-Mapping Strategy

There are three choices:

  • true: when an unknown field is encountered, perform dynamic mapping.
  • false: when an unknown field is encountered, ignore it.
  • strict: when an unknown field is encountered, report an error.
# Global strict strategy; change to true for address
PUT /my_index
{
    "mappings": {
        "my_type": {
            "dynamic": "strict",
            "properties": {
                "title": {
                    "type": "text"
                },
                "address": {
                    "type": "object",
                    "dynamic": "true"
                }
            }
        }
    }
}

2.3.3. View Mapping Information

GET /my_index/_mapping/my_type

2.3.4. Analyzer

# Specify an analyzer for a new field
PUT /my_index/_mapping/my_type
{
  "properties": {
    "content": {
      "type": "text",
      "analyzer": "my_analyzer"
    }
  }
}

# Test the analyzer
GET /my_index/_analyze
{
  "text": "tom&jerry are a friend in the house, <a>, HAHA!!",
  "analyzer": "my_analyzer"
}

2.4. Document Operations

2.4.1. Create a Document

PUT /ecommerce/product/1
{
    "name" : "gaolujie yagao",
    "desc" :  "gaoxiao meibai",
    "price" :  30,
    "producer" :      "gaolujie producer",
    "tags": [ "meibai", "fangzhu" ]
}

{
  "_index": "ecommerce",
  "_type": "product",
  "_id": "1",
  "_version": 1,// Version number; can be used to implement optimistic locking
  "result": "updated",
  "_shards": {
    "total": 2,// Total number of primary + replica shards
    "successful": 1,// Number of shards written successfully
    "failed": 0// Number of failed shards
  },
  "created": false
}

Syntax: /index/type/id. Newer versions have removed type. ES automatically creates the index; we do not need to create it in advance. When ES creates a document, if this ID already exists, the old one is deleted and a new document is created with the same ID and version + 1. If you want creation to report an error when the document already exists, use PUT /index/type/id/_create.

2.4.1.1. Data Routing

The data of an index is divided among multiple shards. When creating a document, it is necessary to decide which shard this document should be routed to.

  • Routing algorithm
shard = hash(routing value) % number_of_primary_shards
  • Specify a routing value. _id is used by default. Another field can be specified with put /index/type/id?routing=user_id.

2.4.2. Query Documents

2.4.2.1. Query String

Append parameters directly to the URL.

2.4.2.1.1. Query by ID
GET /ecommerce/product/1
2.4.2.1.2. Query All
GET /ecommerce/product/_search

{
  "took": 2,// How many milliseconds it took
  "timed_out": false,// Whether it timed out; here it did not
  "_shards": {// Shard results reached by the request
    "total": 5,// There are 5 shards in total; the request reaches all primary shards (replica shards can also be used)
    "successful": 5,// Number of shards that responded successfully
    "failed": 0
  },
  "hits": {
    "total": 2,// Number of query results
    "max_score": 1,// The higher the score, the better the match
    "hits": [// Detailed document data matching the search
      {
        "_index": "ecommerce",
        "_type": "product",
        "_id": "2",
        "_score": 1,
        "_source": {
          "name": "jiajieshi yagao",
          "desc": "youxiao fangzhu",
          "price": 25,
          "producer": "jiajieshi producer",
          "tags": [
            "fangzhu"
          ]
        }
      },
      {
        "_index": "ecommerce",
        "_type": "product",
        "_id": "3",
        "_score": 1,
        "_source": {
          "name": "zhonghua yagao",
          "desc": "caoben zhiwu",
          "price": 40,
          "producer": "zhonghua producer",
          "tags": [
            "qingxin"
          ]
        }
      }
    ]
  }
}
2.4.2.2. Query DSL

Pass parameters using a RequestBody.

2.4.2.3. match_all

Query all.

  • select * from product
GET /ecommerce/product/_search
{
  "query": {
    "match_all": {}
  }
}
2.4.2.4. match

Tokenize first, then query using or or and.

  • select * from product where name like '%yagao%'
GET /ecommerce/product/_search
{
  "query": {
    "match": {
      "name": "yagao"
    }
  }
}

GET /ecommerce/product/_search
{
  "query": {
    "bool": {
      "must":     { "match": { "name": "yagao" }}
    }
  }
}
  • select * from product where name like '%yagao%' or name like '%maojin%'
GET /ecommerce/product/_search
{
  "query": {
    "match": {
      "name": "yagao maojin"
    }
  }
}
  • select * from article where title like '%yagao%' and title like '%maojin%'
GET /forum/article/_search
{
    "query": {
        "match": {
            "title": {
          		"query": "java elasticsearch",
          		"operator": "and"
   	        }
        }
    }
}

# Match at least 3 terms
GET /forum/article/_search
{
    "query": {
        "match": {
            "title": {
          		"query": "java elasticsearch spark hadoop",
          		"minimum_should_match": "75%"
   	        }
        }
    }
}

match -> term + should

{
    "match": { "title": "java elasticsearch"}
}
// Converted to
{
  "bool": {
    "should": [
      { "term": { "title": "java" }},
      { "term": { "title": "elasticsearch"}}
    ]
  }
}
{
    "match": {
        "title": {
            "query":    "java elasticsearch",
            "operator": "and"
        }
    }
}
// Converted to
{
  "bool": {
    "must": [
      { "term": { "title": "java" }},
      { "term": { "title": "elasticsearch"   }}
    ]
  }
}
{
    "match": {
        "title": {
            "query":                "java elasticsearch hadoop spark",
            "minimum_should_match": "75%"
        }
    }
}
// Converted to
{
  "bool": {
    "should": [
      { "term": { "title": "java" }},
      { "term": { "title": "elasticsearch"   }},
      { "term": { "title": "hadoop" }},
      { "term": { "title": "spark" }}
    ],
    "minimum_should_match": 3
  }
}
2.4.2.5. term

Perform an exact field query, generally used for keyword, date, and integer types. If the field queried by a match query is not_analyzed, it is equivalent to a term query.

  • select * from product where name = 'yagao'
GET /ecommerce/product/_search
{
  "query": {
    "term": {
      "name": "yagao"
    }
  }
}
  • select * from product where name = 'yagao maojin'
GET /ecommerce/product/_search
{
  "query": {
    "term": {
      "name": "yagao maojin"
    }
  }
}
2.4.2.6. match_phrase

Do not split the phrase; match the entire phrase.

  • select * from product where name like '%yagao producer%'
GET /ecommerce/product/_search
{
    "query" : {
        "match_phrase" : {
            "producer" : "yagao producer"
        }
    }
}

{
  "took": 10,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 1,
    "max_score": 0.70293105,
    "hits": [
      {
        "_index": "ecommerce",
        "_type": "product",
        "_id": "4",
        "_score": 0.70293105,
        "_source": {
          "name": "special yagao",
          "desc": "special meibai",
          "price": 50,
          "producer": "special yagao producer",
          "tags": [
            "meibai"
          ]
        }
      }
    ]
  }
}

If it is match, it is actually or like: select * from product where name like '%yagao%' or name like '%producer%'

GET /ecommerce/product/_search
{
    "query" : {
        "match" : {
            "producer" : "yagao producer"
        }
    }
}
2.4.2.7. bool

Combine multiple query conditions.

GET /tb_item/_doc/_search
{
"query": {
    "bool": {# Overall must and must_not are connected with and; should is optional and only improves relevance
        "must": { "match": {"title": "电视"}},# Conditions inside must use and
        "must_not": { "term": {"id": "927779"}},# Conditions inside must_not use and-not
        "should": { "match": {"sellPoint": "好评"}}# Conditions inside should use or
        }
    }
}
2.4.2.8. fuzzy

During search, the entered search text may contain misspellings. Fuzzy-search technology automatically corrects misspelled search text and then tries to match data in the index.

GET /my_index/my_type/_search
{
  "query": {
    "fuzzy": {
      "text": {
        "value": "surprize",
        "fuzziness": 2// Maximum number of letters that can be corrected before matching the data; default is 2
      }
    }
  }
}
GET /my_index/my_type/_search
{
  "query": {
    "match": {
      "text": {
        "query": "SURPIZE ME",
        "fuzziness": "AUTO",
        "operator": "and"
      }
    }
  }
}
2.4.2.9. _source

Fields returned by the query.

  • select name, price from product
GET /ecommerce/product/_search
{
  "query": {
    "match_all": {}
  },
  "_source": ["name","price"]
}
2.4.2.10. range

Range query.

  • select * from article where view_cnt between 30 and 60
GET /forum/article/_search
{
  "query": {
    "constant_score": {
      "filter": {
        "range": {
          "view_cnt": {
            "gte": 30,
            "lte": 60
          }
        }
      }
    }
  }
}
2.4.2.11. sort

Sort.

  • select * from product where name like '%yagao%' order by price desc
GET /ecommerce/product/_search
{
  "query": {
    "match": {
      "name": "yagao"
    },
  "sort": [
    {
      "price": {
        "order": "desc"
      }
    }
  ]
  }
}
2.4.2.12. limit

Paginated query.

  • select * from product limit 0, 1
GET /ecommerce/product/_search
{
  "query": {
    "match_all": {}
  },
  "from": 1,
  "size": 1
}
2.4.2.13. filter

The only difference from query is that it does not participate in relevance scoring.

  • select * from product where name like '%yagao%' and price >= 25
GET /ecommerce/product/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "match": {
            "name": "yagao"
          }
        }
      ],
      "filter": {
        "range": {
          "price": {
            "gte": 25
          }
        }
      }
    }
  }
}
2.4.2.14. highlight

Highlight search results. You can configure the highlight HTML tags, force a particular highlighter, and highlight multiple fields.

GET /ecommerce/product/_search
{
    "query" : {
        "match_phrase" : {
            "producer" : "yagao producer"
        }
    },
    "highlight": {
      "fields": {
        "producer": {}
      }
    }
}

{
  "took": 61,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 1,
    "max_score": 0.70293105,
    "hits": [
      {
        "_index": "ecommerce",
        "_type": "product",
        "_id": "4",
        "_score": 0.70293105,
        "_source": {
          "name": "special yagao",
          "desc": "special meibai",
          "price": 50,
          "producer": "special yagao producer",
          "tags": [
            "meibai"
          ]
        },
        "highlight": {
          "producer": [
            "special <em>yagao</em> <em>producer</em>"
          ]
        }
      }
    ]
  }
}

The principle is to save a snapshot of the current data. The disadvantage is that pages can only be traversed downward one page at a time, similar to scrolling down a Weibo feed. You cannot jump freely between pages; doing so would be even less efficient.

2.4.2.15.1. Rebuild an Index With Zero Downtime

Prerequisite: the index used by the client is an alias. Create a new index and create the text field as a string type. Use the scroll API to query the data in batches. Use the bulk API to insert it into the new index in batches. Remove the old index from the alias and associate the new index with the alias.

2.4.2.16. Weight Control

You can use boost to assign weights to search conditions and quantify their importance. By default, all search conditions have the same weight, which is 1.

GET /forum/article/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "match": {
            "title": "blog"
          }
        }
      ],
      "should": [
        {
          "match": {
            "title": {
              "query": "java"
            }
          }
        },
        {
          "match": {
            "title": {
              "query": "hadoop"
            }
          }
        },
        {
          "match": {
            "title": {
              "query": "elasticsearch"
            }
          }
        },
        {
          "match": {
            "title": {
              "query": "spark",
              "boost": 5
            }
          }
        }
      ]
    }
  }
}

2.4.3. Aggregate Documents

2.4.3.1. Grouping

select tags, count(*) from product group by tags

First, fielddata=false by default on a text field used for grouping, so it needs to be set to true.

PUT /ecommerce/product/_mapping
{
  "properties": {
    "tags":{
      "type": "text",
      "fielddata": true
    }
  }
}

Then aggregation can be used.

GET /ecommerce/product/_search
{
  "aggs": {
    "group_by_tags": {
      "terms": {
        "field": "tags"
      }
    }
  }
}

{
  "took": 146,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 4,
    "max_score": 0,
    "hits": []
  },
  "aggregations": {
    "group_by_tags": {
      "doc_count_error_upper_bound": 0,
      "sum_other_doc_count": 0,
      "buckets": [
        {
          "key": "fangzhu",
          "doc_count": 2
        },
        {
          "key": "meibai",
          "doc_count": 2
        },
        {
          "key": "qingxin",
          "doc_count": 1
        }
      ]
    }
  }
}
2.4.3.2. Group After Querying

select tags, count(*) from product where name like '%yagao%' group by tags

GET /ecommerce/product/_search
{
  "query": {
    "match": {
      "name": "yagao"
    }
  },
  "aggs": {
    "group_by_tags": {
      "terms": {
        "field": "tags"
      }
    }
  },
  "size": 0
}
2.4.3.3. Calculate an Average After Grouping

select avg(price) from product where tags like '%tags%' group by tags

GET /ecommerce/product/_search
{
  "aggs": {
    "group_by_tags": {
      "terms": {
        "field": "tags"
      },
      "aggs": {
        "avg_by_price": {
          "avg": {
            "field": "price"
          }
        }
      }
    }
  },
  "size": 0
}

{
  "took": 3,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 4,
    "max_score": 0,
    "hits": []
  },
  "aggregations": {
    "group_by_tags": {
      "doc_count_error_upper_bound": 0,
      "sum_other_doc_count": 0,
      "buckets": [
        {
          "key": "fangzhu",
          "doc_count": 2,
          "avg_by_price": {
            "value": 27.5
          }
        },
        {
          "key": "meibai",
          "doc_count": 2,
          "avg_by_price": {
            "value": 40
          }
        },
        {
          "key": "qingxin",
          "doc_count": 1,
          "avg_by_price": {
            "value": 40
          }
        }
      ]
    }
  }
}
2.4.3.4. Sort After Grouping

select avg(price) from product where tags like '%tags%' group by tags order by avg(price) desc

GET /ecommerce/product/_search
{
  "aggs": {
    "group_by_tags": {
      "terms": {
        "field": "tags",
        "order": {
          "avg_by_price": "desc"
        }
      },
      "aggs": {
        "avg_by_price": {
          "avg": {
            "field": "price"
          }
        }
      }
    }
  },
  "size": 0
}

2.4.4. Modify a Document

2.4.4.1. Full Update
PUT /ecommerce/product/1
{
    "name" : "jiaqiangban gaolujie yagao",
    "desc" :  "gaoxiao meibai",
    "price" :  30,
    "producer" :      "gaolujie producer",
    "tags": [ "meibai", "fangzhu" ]
}
2.4.4.2. Partial Update
POST /ecommerce/product/1/_update
{
  "doc": {
    "name": "jiaqiangban gaolujie yagao"
  }
}
2.4.4.2.1. Built-In Optimistic-Locking Concurrency Control
POST /test_index/test_type/11/_update?retry_on_conflict=2
{
  "doc": {
    "num" : 2
  }
}

retry_on_conflict is the number of retries when a concurrency conflict occurs:

  • Get the document data and latest version.
  • Update by comparing this version number with the server’s version number.
  • Retry if it fails.

2.4.5. Bulk Operations

2.4.5.1. Bulk Query

When querying one record at a time, for example querying 100 records, 100 network requests need to be sent, which has substantial overhead. With a bulk query, querying 100 records requires only 1 network request, reducing network-request overhead by 100 times.

# Different indices
GET /_mget
{
   "docs" : [
      {
         "_index" : "test_index",
         "_type" :  "test_type",
         "_id" :    10
      },
      {
         "_index" : "test_index",
         "_type" :  "test_type",
         "_id" :    11
      }
   ]
}

# The same index
GET /test_index/_mget
{
   "docs" : [
      {
         "_type" :  "test_type",
         "_id" :    10
      },
      {
         "_type" :  "test_type",
         "_id" :    11
      }
   ]
}
2.4.5.2. Bulk Create/Delete/Update
  • delete: delete one document; only one JSON line is needed.
  • create: PUT /index/type/id/_create; force creation / report an error if it already exists.
  • index: ordinary PUT operation; can create a document or fully replace a document.
  • update: perform a partial-update operation.
POST /_bulk
{ "delete": { "_index": "test_index", "_type": "test_type", "_id": "3" }}
{ "create": { "_index": "test_index", "_type": "test_type", "_id": "12" }}
{ "test_field":    "test12" }
{ "index":  { "_index": "test_index", "_type": "test_type", "_id": "2" }}
{ "test_field":    "replaced test2" }
{ "update": { "_index": "test_index", "_type": "test_type", "_id": "1", "_retry_on_conflict" : 3} }
{ "doc" : {"test_field2" : "bulk test1"} }
2.4.5.2.1. Best bulk size

A bulk request is loaded into memory. If it is too large, performance will decrease, so the best bulk size needs to be found through repeated testing. Generally, start with 1000-5000 records and gradually increase. In terms of byte size, preferably keep it between 5-15 MB.

2.4.6. Delete a Document

DELETE /ecommerce/product/1

2.4.7. Document Metadata

2.4.7.1. _index

Represents which index a document is stored in.

2.4.7.2. _type

An index is usually divided into multiple types.

2.4.7.3. _id

Represents the unique identifier of a document. Together with index and type, it can uniquely identify and locate a document.

  • Manually generated ID:
PUT /test_index/test_type/1
{
  "test_content": "test test"
}
  • Automatically generated ID:
POST /test_index/test_type
{
  "test_content": "test test"
}
- Length is 20 characters.
- URL-safe: a Base64-encoded ID can be passed in a URL.
- GUID method: collisions cannot occur when generated in parallel in a distributed system.
2.4.7.4. _source

The field returned by default when querying.

# 2. Add
PUT /test_index/test_type/1
{
  "test_content": "test test",
  "test_content2": "test test2"
}

# Query
GET /test_index/test_type/1
{
  "_index": "test_index",
  "_type": "test_type",
  "_id": "1",
  "_version": 2,
  "found": true,
  "_source": {// The default _source is what we added
    "test_content": "test test",
    "test_content2": "test test2"
  }
}

The returned _source can be customized.

GET /test_index/test_type/_search
{
  "query": {
    "match": {
      "_id": "1"
    }
  },
  "_source": ["test_content","test_content2"]
}
2.4.7.5. _all

Package all fields together as an _all field and build an index for it. When no field is specified for a search, the _all field is used for the search.

  • This field is enabled by default and can be disabled manually.
PUT /my_index/_mapping/my_type3
{
  "_all": {"enabled": false}
}
  • include_in_all can be configured at the field level to control whether the field’s value is included in the _all field.
PUT /my_index/_mapping/my_type4
{
  "properties": {
    "my_field": {
      "type": "text",
      "include_in_all": false
    }
  }
}
2.4.7.6. _version

The value is 0 at creation time and automatically increases by 1 when modified or deleted.

2.4.7.6.1. Implement Optimistic Locking to Resolve Concurrent-Update Conflicts
  1. First add a record. At this point version = 1.
PUT /test_index/test_type/7
{
  "test_field": "test test"
}
  1. Update the data with version = 1. Client 1 updates successfully.
PUT /test_index/test_type/7?version=1
{
  "test_field": "test client 1"
}
  1. Update the data again with version = 1.
PUT /test_index/test_type/7?version=1
{
  "test_field": "test client 2"
}

3. Elasticsearch and MySQL Synchronization

4. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub