Elasticsearch 是一个分布式、基于文档的搜索引擎,具有高效的全文搜索、分析和分布式存储能力,广泛应用于日志分析、全文检索、推荐系统、实时监控等场景。以下是 Elasticsearch 在实际应用中的一些典型场景和技术实现。

1. 全文搜索与过滤

基本全文搜索

Elasticsearch 的核心功能之一是全文搜索。它可以对大规模的文本数据进行高效检索,并支持模糊搜索、分词匹配、多字段查询等。

  • 简单查询:match 查询

    {
      "query": {
        "match": {
          "content": "Elasticsearch best practices"
        }
      }
    }
    

    该查询会对 content 字段进行分词,并对包含关键词 “Elasticsearch” 和 “best practices” 的文档进行匹配。

  • 多字段查询:multi_match
    如果你想在多个字段中同时进行搜索,可以使用 multi_match 查询:

    {
      "query": {
        "multi_match": {
          "query": "Elasticsearch guide",
          "fields": ["title", "content"]
        }
      }
    }
    
布尔查询(Bool Query)

当需要进行复杂的查询时,可以通过 bool 查询将多个条件组合起来,支持 must(必须匹配)、should(可选匹配)和 must_not(排除匹配)。

  • 组合查询
    {
      "query": {
        "bool": {
          "must": [
            { "match": { "title": "Elasticsearch" } }
          ],
          "should": [
            { "match": { "tags": "search" } }
          ],
          "must_not": [
            { "match": { "status": "deprecated" } }
          ]
        }
      }
    }
    
过滤器(Filter)

在一些场景中,除了搜索关键词,还可能需要对结果进行过滤,如基于日期、数值范围等。filter 不会影响文档的相关性得分,通常与 bool 查询配合使用。

  • 日期范围过滤
    {
      "query": {
        "bool": {
          "must": { "match": { "content": "Elasticsearch" } },
          "filter": {
            "range": {
              "publish_date": {
                "gte": "2022-01-01",
                "lte": "2023-01-01"
              }
            }
          }
        }
      }
    }
    

2. 日志分析与监控

Elasticsearch 通常与 Logstash、Kibana(Elastic Stack)一起使用,用于构建日志分析平台。日志分析主要包括实时收集、存储、查询与可视化。

日志收集与处理

日志通过 Logstash 或 Filebeat 等工具采集到 Elasticsearch 中,常见的做法是将日志按照时间进行索引划分,以便于管理和优化查询性能。

  • 时间分段索引:创建按日期划分的索引,如每天生成一个新索引。
    {
      "index_patterns": ["logs-*"],
      "settings": {
        "number_of_shards": 3,
        "number_of_replicas": 1
      },
      "mappings": {
        "properties": {
          "timestamp": {
            "type": "date"
          },
          "log_level": {
            "type": "keyword"
          },
          "message": {
            "type": "text"
          }
        }
      }
    }
    
日志搜索和分析

通过 Kibana 提供的查询界面,用户可以快速检索特定时间段内的日志,并通过数据聚合功能进行分析。

  • 按时间段检索错误日志
    {
      "query": {
        "bool": {
          "must": { "match": { "log_level": "ERROR" } },
          "filter": {
            "range": {
              "timestamp": {
                "gte": "now-1d/d",
                "lt": "now/d"
              }
            }
          }
        }
      }
    }
    

3. 数据聚合与分析

Elasticsearch 支持强大的聚合功能,允许你对数据进行统计、分组、求和、取平均值等复杂的分析操作,适合处理各种报表和数据分析需求。

基本聚合

使用 aggregation 可以对数据进行统计分析。以下是按 status 字段进行分组,并计算每个状态的文档数量。

  • 按字段分组计数
    {
      "size": 0,
      "aggs": {
        "status_count": {
          "terms": {
            "field": "status"
          }
        }
      }
    }
    
嵌套聚合

你可以进行嵌套聚合,计算某个字段在不同维度下的数值统计。

  • 按状态分组,并在每个状态下按日期分组
    {
      "size": 0,
      "aggs": {
        "status_count": {
          "terms": {
            "field": "status"
          },
          "aggs": {
            "date_count": {
              "date_histogram": {
                "field": "publish_date",
                "interval": "month"
              }
            }
          }
        }
      }
    }
    
基数聚合

基数聚合(cardinality aggregation)用于计算某个字段的唯一值的数量。例如,统计独立用户的访问量:

  • 独立用户统计
    {
      "size": 0,
      "aggs": {
        "unique_users": {
          "cardinality": {
            "field": "user_id"
          }
        }
      }
    }
    

4. 自动完成与建议功能

Elasticsearch 提供了几种方式来实现自动完成(Autocomplete)和搜索建议(Search Suggestions),在构建搜索引擎时非常实用。

基于 completion 字段的自动完成

使用 completion 字段类型实现快速的自动完成功能。

  • 创建索引时定义自动完成字段

    {
      "mappings": {
        "properties": {
          "title": {
            "type": "text"
          },
          "suggest": {
            "type": "completion"
          }
        }
      }
    }
    
  • 查询建议

    {
      "suggest": {
        "title_suggestion": {
          "prefix": "elas",
          "completion": {
            "field": "suggest"
          }
        }
      }
    }
    
基于 edge_ngram 的自动完成

另一种自动完成的方式是使用 edge_ngram 分词器。这种方式更加灵活,可以根据部分输入进行匹配。

  • 索引设置:定义 edge_ngram 分词

    {
      "settings": {
        "analysis": {
          "filter": {
            "edge_ngram_filter": {
              "type": "edge_ngram",
              "min_gram": 1,
              "max_gram": 20
            }
          },
          "analyzer": {
            "autocomplete": {
              "type": "custom",
              "tokenizer": "standard",
              "filter": ["lowercase", "edge_ngram_filter"]
            }
          }
        }
      },
      "mappings": {
        "properties": {
          "title": {
            "type": "text",
            "analyzer": "autocomplete",
            "search_analyzer": "standard"
          }
        }
      }
    }
    
  • 查询自动完成结果

    {
      "query": {
        "match": {
          "title": {
            "query": "elas",
            "analyzer": "standard"
          }
        }
      }
    }
    

5. 实时数据监控与告警

Elasticsearch 可以用来做实时数据监控,结合 Watcher 或其他监控工具,可以在特定条件下触发告警。

  • 定义监控条件和告警
    例如,当系统错误日志达到一定数量时发送通知:
    {
      "trigger": {
        "schedule": { "interval": "10s" }
      },
      "input": {
        "search": {
          "request": {
            "indices": ["logs-*"],
            "body": {
              "query": {
                "bool": {
                  "filter": { "term": { "log_level": "ERROR" } }
                }
              }
            }
          }
        }
      },
      "condition": {
        "compare": {
          "ctx.payload.hits.total": { "gt": 100 }
        }
      },
      "actions": {
        "email_admin": {
          "email": {
            "to": "admin@example.com",
            "subject": "Error Logs Exceeded Threshold",
            "body": "More than 100 error logs found in the last 10 seconds."
          }
        }
      }
    }
    

总结

通过

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐