elasticsearch实战应用
Elasticsearch 是一个分布式、基于文档的搜索引擎,具有高效的全文搜索、分析和分布式存储能力,广泛应用于日志分析、全文检索、推荐系统、实时监控等场景。以下是 Elasticsearch 在实际应用中的一些典型场景和技术实现。
1. 全文搜索与过滤
基本全文搜索
Elasticsearch 的核心功能之一是全文搜索。它可以对大规模的文本数据进行高效检索,并支持模糊搜索、分词匹配、多字段查询等。
-
简单查询:
match查询{ "query": { "match": { "content": "Elasticsearch best practices" } } }该查询会对
content字段进行分词,并对包含关键词 “Elasticsearch” 和 “best practices” 的文档进行匹配。 -
多字段查询:
multi_match
如果你想在多个字段中同时进行搜索,可以使用multi_match查询:{ "query": { "multi_match": { "query": "Elasticsearch guide", "fields": ["title", "content"] } } }
布尔查询(Bool Query)
当需要进行复杂的查询时,可以通过 bool 查询将多个条件组合起来,支持 must(必须匹配)、should(可选匹配)和 must_not(排除匹配)。
- 组合查询
{ "query": { "bool": { "must": [ { "match": { "title": "Elasticsearch" } } ], "should": [ { "match": { "tags": "search" } } ], "must_not": [ { "match": { "status": "deprecated" } } ] } } }
过滤器(Filter)
在一些场景中,除了搜索关键词,还可能需要对结果进行过滤,如基于日期、数值范围等。filter 不会影响文档的相关性得分,通常与 bool 查询配合使用。
- 日期范围过滤
{ "query": { "bool": { "must": { "match": { "content": "Elasticsearch" } }, "filter": { "range": { "publish_date": { "gte": "2022-01-01", "lte": "2023-01-01" } } } } } }
2. 日志分析与监控
Elasticsearch 通常与 Logstash、Kibana(Elastic Stack)一起使用,用于构建日志分析平台。日志分析主要包括实时收集、存储、查询与可视化。
日志收集与处理
日志通过 Logstash 或 Filebeat 等工具采集到 Elasticsearch 中,常见的做法是将日志按照时间进行索引划分,以便于管理和优化查询性能。
- 时间分段索引:创建按日期划分的索引,如每天生成一个新索引。
{ "index_patterns": ["logs-*"], "settings": { "number_of_shards": 3, "number_of_replicas": 1 }, "mappings": { "properties": { "timestamp": { "type": "date" }, "log_level": { "type": "keyword" }, "message": { "type": "text" } } } }
日志搜索和分析
通过 Kibana 提供的查询界面,用户可以快速检索特定时间段内的日志,并通过数据聚合功能进行分析。
- 按时间段检索错误日志
{ "query": { "bool": { "must": { "match": { "log_level": "ERROR" } }, "filter": { "range": { "timestamp": { "gte": "now-1d/d", "lt": "now/d" } } } } } }
3. 数据聚合与分析
Elasticsearch 支持强大的聚合功能,允许你对数据进行统计、分组、求和、取平均值等复杂的分析操作,适合处理各种报表和数据分析需求。
基本聚合
使用 aggregation 可以对数据进行统计分析。以下是按 status 字段进行分组,并计算每个状态的文档数量。
- 按字段分组计数
{ "size": 0, "aggs": { "status_count": { "terms": { "field": "status" } } } }
嵌套聚合
你可以进行嵌套聚合,计算某个字段在不同维度下的数值统计。
- 按状态分组,并在每个状态下按日期分组
{ "size": 0, "aggs": { "status_count": { "terms": { "field": "status" }, "aggs": { "date_count": { "date_histogram": { "field": "publish_date", "interval": "month" } } } } } }
基数聚合
基数聚合(cardinality aggregation)用于计算某个字段的唯一值的数量。例如,统计独立用户的访问量:
- 独立用户统计
{ "size": 0, "aggs": { "unique_users": { "cardinality": { "field": "user_id" } } } }
4. 自动完成与建议功能
Elasticsearch 提供了几种方式来实现自动完成(Autocomplete)和搜索建议(Search Suggestions),在构建搜索引擎时非常实用。
基于 completion 字段的自动完成
使用 completion 字段类型实现快速的自动完成功能。
-
创建索引时定义自动完成字段
{ "mappings": { "properties": { "title": { "type": "text" }, "suggest": { "type": "completion" } } } } -
查询建议
{ "suggest": { "title_suggestion": { "prefix": "elas", "completion": { "field": "suggest" } } } }
基于 edge_ngram 的自动完成
另一种自动完成的方式是使用 edge_ngram 分词器。这种方式更加灵活,可以根据部分输入进行匹配。
-
索引设置:定义
edge_ngram分词{ "settings": { "analysis": { "filter": { "edge_ngram_filter": { "type": "edge_ngram", "min_gram": 1, "max_gram": 20 } }, "analyzer": { "autocomplete": { "type": "custom", "tokenizer": "standard", "filter": ["lowercase", "edge_ngram_filter"] } } } }, "mappings": { "properties": { "title": { "type": "text", "analyzer": "autocomplete", "search_analyzer": "standard" } } } } -
查询自动完成结果
{ "query": { "match": { "title": { "query": "elas", "analyzer": "standard" } } } }
5. 实时数据监控与告警
Elasticsearch 可以用来做实时数据监控,结合 Watcher 或其他监控工具,可以在特定条件下触发告警。
- 定义监控条件和告警
例如,当系统错误日志达到一定数量时发送通知:{ "trigger": { "schedule": { "interval": "10s" } }, "input": { "search": { "request": { "indices": ["logs-*"], "body": { "query": { "bool": { "filter": { "term": { "log_level": "ERROR" } } } } } } } }, "condition": { "compare": { "ctx.payload.hits.total": { "gt": 100 } } }, "actions": { "email_admin": { "email": { "to": "admin@example.com", "subject": "Error Logs Exceeded Threshold", "body": "More than 100 error logs found in the last 10 seconds." } } } }
总结
通过
更多推荐
所有评论(0)