在开发多智能体系统时,我们常常面临这样的挑战:如何让智能体记住用户偏好、历史对话和关键知识,从而提供更个性化、更准确的响应?AutoGen 框架中的 Memory(记忆)机制和 RAG(检索增强生成)模式为这个问题提供了完美解决方案。今天,我们就来深入探讨如何在 AutoGen 中实现智能体的记忆管理和知识检索,让智能体真正拥有 "长期记忆" 的能力。

一、Memory 协议:智能体的记忆核心

记忆系统的基本概念

Memory 协议是 AutoGen 中智能体记忆功能的基础,它定义了一套标准接口,让智能体能够存储、检索和管理上下文信息。在实际应用中,记忆系统就像智能体的 "大脑存储区",负责保存用户偏好、历史对话、知识库等关键信息。

Memory 协议包含以下核心方法:

  • add:向记忆存储中添加新条目
  • query:检索与查询相关的信息
  • update_context:将检索到的信息注入智能体的上下文
  • clear:清除记忆存储中的所有条目
  • close:释放记忆存储使用的资源

这些方法构成了智能体记忆功能的基础,无论是简单的列表记忆还是复杂的向量数据库记忆,都需要实现这些接口。

ListMemory:最简单的记忆实现

作为入门示例,AutoGen 提供了基于列表的记忆实现ListMemory,它按时间顺序存储记忆条目,并在检索时返回所有或部分历史记录。下面是一个使用 ListMemory 维护用户偏好的示例:

python

from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_core.memory import ListMemory, MemoryContent, MemoryMimeType
from autogen_ext.models.openai import OpenAIChatCompletionClient

# 初始化列表记忆
user_memory = ListMemory()

# 添加用户偏好到记忆
await user_memory.add(MemoryContent(content="The weather should be in metric units", mime_type=MemoryMimeType.TEXT))
await user_memory.add(MemoryContent(content="Meal recipe must be vegan", mime_type=MemoryMimeType.TEXT))

# 定义天气查询工具
async def get_weather(city: str, units: str = "imperial") -> str:
    if units == "imperial":
        return f"The weather in {city} is 73 °F and Sunny."
    elif units == "metric":
        return f"The weather in {city} is 23 °C and Sunny."
    else:
        return f"Sorry, I don't know the weather in {city}."

# 创建智能体并关联记忆
assistant_agent = AssistantAgent(
    name="assistant_agent",
    model_client=OpenAIChatCompletionClient(model="gpt-4o-2024-08-06"),
    tools=[get_weather],
    memory=[user_memory],
)

# 运行智能体查询天气
stream = assistant_agent.run_stream(task="What is the weather in New York?")
await Console(stream)

在这个示例中,智能体在响应用户的天气查询时,会自动从 ListMemory 中检索用户偏好的单位设置(公制单位),并在调用天气工具时使用该偏好,最终返回以摄氏度为单位的天气信息。

二、自定义记忆存储:从列表到向量数据库

向量数据库记忆的优势

ListMemory 虽然简单,但在处理大规模知识库时存在明显不足。这时,基于向量数据库的记忆实现就成为更好的选择。向量数据库通过将文本转换为高维向量,利用余弦相似度等算法检索语义相关的内容,相比传统列表记忆具有以下优势:

  • 语义检索:基于内容语义而非精确匹配进行检索
  • 高效查询:适合处理海量文档和知识库
  • 动态扩展:轻松添加新的记忆条目而不影响查询效率

ChromaDBVectorMemory 实战

AutoGen 通过autogen_ext扩展包提供了基于 ChromaDB 的向量记忆实现,下面是如何使用它构建更强大的记忆系统:

python

import os
from pathlib import Path
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_core.memory import MemoryContent, MemoryMimeType
from autogen_ext.memory.chromadb import ChromaDBVectorMemory, PersistentChromaDBVectorMemoryConfig
from autogen_ext.models.openai import OpenAIChatCompletionClient

# 配置ChromaDB记忆存储
chroma_user_memory = ChromaDBVectorMemory(
    config=PersistentChromaDBVectorMemoryConfig(
        collection_name="preferences",
        persistence_path=os.path.join(str(Path.home()), ".chromadb_autogen"),
        k=2,       # 返回前2个最相关结果
        score_threshold=0.4  # 最小相似度分数
    )
)

# 添加用户偏好到向量记忆
await chroma_user_memory.add(
    MemoryContent(
        content="The weather should be in metric units",
        mime_type=MemoryMimeType.TEXT,
        metadata={"category": "preferences", "type": "units"}
    )
)
await chroma_user_memory.add(
    MemoryContent(
        content="Meal recipe must be vegan",
        mime_type=MemoryMimeType.TEXT,
        metadata={"category": "preferences", "type": "dietary"}
    )
)

# 创建关联向量记忆的智能体
model_client = OpenAIChatCompletionClient(model="gpt-4o")
assistant_agent = AssistantAgent(
    name="assistant_agent",
    model_client=model_client,
    tools=[get_weather],
    memory=[chroma_user_memory],
)

# 运行查询
stream = assistant_agent.run_stream(task="What is the weather in New York?")
await Console(stream)

# 清理资源
await model_client.close()
await chroma_user_memory.close()

与 ListMemory 相比,ChromaDBVectorMemory 在处理复杂查询时表现更优。它会根据语义相似度检索最相关的记忆条目,并将其注入智能体的上下文,使响应更加精准。

三、RAG 模式:检索增强生成的完整实现

RAG 的核心流程

RAG(Retrieval Augmented Generation)是当前 AI 系统中广泛应用的技术模式,它将生成过程分为两个关键阶段:

  1. 索引阶段:将文档分块、转换为向量并存储到数据库
  2. 检索阶段:在对话时检索相关文档块并注入生成上下文

这种模式解决了大语言模型 "健忘" 的问题,让智能体能够基于最新、最准确的信息生成响应,尤其适合知识库问答、文档分析等场景。

构建完整的 RAG 智能体

下面我们通过一个完整案例,展示如何构建一个基于 AutoGen 的 RAG 智能体系统:

python

import re
from typing import List
import aiofiles
import aiohttp
from autogen_core.memory import Memory, MemoryContent, MemoryMimeType

# 文档索引器:负责加载、分块和存储文档
class SimpleDocumentIndexer:
    """基础文档索引器,用于AutoGen记忆系统"""
    def __init__(self, memory: Memory, chunk_size: int = 1500) -> None:
        self.memory = memory
        self.chunk_size = chunk_size
    
    async def _fetch_content(self, source: str) -> str:
        """从URL或文件获取内容"""
        if source.startswith(("http://", "https://")):
            async with aiohttp.ClientSession() as session:
                async with session.get(source) as response:
                    return await response.text()
        else:
            async with aiofiles.open(source, "r", encoding="utf-8") as f:
                return await f.read()
    
    def _strip_html(self, text: str) -> str:
        """移除HTML标签并规范化空白"""
        text = re.sub(r"<[^>]*>", " ", text)
        text = re.sub(r"\s+", " ", text)
        return text.strip()
    
    def _split_text(self, text: str) -> List[str]:
        """将文本分割为固定大小的块"""
        chunks: list[str] = []
        for i in range(0, len(text), self.chunk_size):
            chunk = text[i : i + self.chunk_size]
            chunks.append(chunk.strip())
        return chunks
    
    async def index_documents(self, sources: List[str]) -> int:
        """将文档索引到记忆中"""
        total_chunks = 0
        for source in sources:
            try:
                content = await self._fetch_content(source)
                # 如果是HTML内容则移除标签
                if "<" in content and ">" in content:
                    content = self._strip_html(content)
                chunks = self._split_text(content)
                for i, chunk in enumerate(chunks):
                    await self.memory.add(
                        MemoryContent(
                            content=chunk, 
                            mime_type=MemoryMimeType.TEXT,
                            metadata={"source": source, "chunk_index": i}
                        )
                    )
                total_chunks += len(chunks)
            except Exception as e:
                print(f"Error indexing {source}: {str(e)}")
        return total_chunks

# 初始化向量记忆并构建RAG智能体
import os
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_ext.memory.chromadb import ChromaDBVectorMemory, PersistentChromaDBVectorMemoryConfig
from autogen_ext.models.openai import OpenAIChatCompletionClient

# 配置向量记忆
rag_memory = ChromaDBVectorMemory(
    config=PersistentChromaDBVectorMemoryConfig(
        collection_name="autogen_docs",
        persistence_path=os.path.join(str(Path.home()), ".chromadb_autogen"),
        k=3,           # 返回前3个相关结果
        score_threshold=0.4
    )
)

# 清空现有记忆(如果有)
await rag_memory.clear()

# 索引AutoGen文档
async def index_autogen_docs():
    indexer = SimpleDocumentIndexer(memory=rag_memory)
    sources = [
        "https://raw.githubusercontent.com/microsoft/autogen/main/README.md",
        "https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/tutorial/agents.html",
        "https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/tutorial/teams.html",
        "https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/tutorial/termination.html",
    ]
    chunks = await indexer.index_documents(sources)
    print(f"Indexed {chunks} chunks from {len(sources)} AutoGen documents")

# 执行索引
await index_autogen_docs()

# 创建RAG智能体
rag_assistant = AssistantAgent(
    name="rag_assistant",
    model_client=OpenAIChatCompletionClient(model="gpt-4o"),
    memory=[rag_memory]
)

# 提问关于AgentChat的问题
stream = rag_assistant.run_stream(task="What is AgentChat?")
await Console(stream)

# 清理资源
await rag_memory.close()

当我们向这个 RAG 智能体提问 "What is AgentChat?" 时,它会执行以下流程:

  1. 将问题转换为向量
  2. 在向量记忆中检索与 "AgentChat" 最相关的文档块
  3. 将检索到的文档块注入 LLM 的上下文
  4. 基于上下文和问题生成回答

实际运行结果如下:

plaintext

---------- user ----------
What is AgentChat?

---------- rag_assistant ----------
AgentChat is part of the AutoGen framework...(基于检索到的文档生成的详细回答)

四、RAG 系统的优化与最佳实践

影响 RAG 性能的关键因素

构建高效的 RAG 系统需要关注以下关键环节:

  1. 文档分块策略:

    • 块大小:100-1500 字为宜,过大会导致上下文冗余,过小会丢失语义连贯
    • 分块方式:基于语义段落分块优于简单字符分割
  2. 向量嵌入质量:

    • 选择适合任务的嵌入模型(如 text-embedding-ada-002)
    • 定期更新嵌入模型以适应数据变化
  3. 检索参数调整:

    • k值:控制返回的相关文档数量,默认 3-5 为宜
    • 相似度阈值:过滤低相关度文档,避免噪声干扰
  4. 元数据过滤:

    • 添加文档类型、来源等元数据
    • 在查询时结合元数据过滤,缩小检索范围

生产环境优化建议

对于实际生产系统,我们建议采取以下优化措施:

python

# 生产环境优化示例
rag_memory = ChromaDBVectorMemory(
    config=PersistentChromaDBVectorMemoryConfig(
        # 其他配置...
        embedding_model="text-embedding-ada-002",  # 使用专业嵌入模型
        chunk_size=1000,                          # 优化分块大小
        metadata_filters={"category": "technical"}, # 元数据过滤
        # 动态调整检索参数
        adaptive_retrieval=True,
        retrieval_adapter={
            "strategy": "hybrid",
            "params": {"semantic_weight": 0.7, "keyword_weight": 0.3}
        }
    )
)
  • 混合检索策略:结合语义检索和关键词检索,提高准确率
  • 自适应参数:根据历史查询效果动态调整k值和相似度阈值
  • 文档更新机制:设置定时任务,自动更新过期文档索引
  • 缓存策略:对高频查询结果进行缓存,减少向量数据库压力

五、总结

通过 Memory 协议和 RAG 模式,我们为智能体赋予了真正的 "记忆能力",使其能够基于历史信息和外部知识库生成更准确、更相关的响应。从简单的 ListMemory 到复杂的向量数据库记忆,AutoGen 提供了完整的记忆解决方案,满足不同场景的需求。

如果本文对你有帮助,别忘了点赞收藏,关注我,一起探索更高效的开发方式~

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐