NOTE
1.6 Elasticsearch Analyzer
English translation of the original VNote ‘Elasticsearch Analyzer’, preserving its structure and examples.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is an Analyzer?
Split a sentence into individual words and perform normalization on each word.
1.1. Recall
Increase the number of results that can be found during search.
2. Components of an Analyzer
2.1. character filter
Preprocess a piece of text before tokenization. Common examples include filtering HTML tags (<span>hello<span> –> hello) and converting & to and (I&you –> I and you).
2.2. tokenizer
Tokenization: hello you and me –> hello, you, and, me.
2.3. token filter
lowercase, stop word, synonym: dogs –> dog, liked –> like, Tom –> tom, a/the/an –> removed, mother –> mom, small –> little.
3. Install IK Analyzer
- Download https://github.com/medcl/elasticsearch-analysis-ik/releases/download/v7.4.2/elasticsearch-analysis-ik-7.4.2.zip
- Extract it under
elasticsearch-5.2.0\plugin\ik. - Restart ES.
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub