Spring AI + Ollama 本地模型实战:5分钟搞定MCP调用(附完整代码)
Spring AI与Ollama本地模型融合实战:构建企业级MCP应用架构
最近在和一些技术团队交流时,发现一个挺有意思的现象:很多开发者已经成功在本地部署了DeepSeek或Qwen3这样的优秀大模型,也能通过简单的API调用实现基础对话功能,但当他们想要把这些AI能力真正集成到现有业务系统中时,却遇到了瓶颈。特别是想要实现类似MCP(模型上下文协议)这样的高级功能时,往往需要从零开始搭建复杂的架构。
这让我想起了几年前微服务刚兴起时的场景——大家都能写单个服务,但如何让这些服务协同工作、如何管理服务间的通信,才是真正的挑战。现在AI领域也面临类似的局面:模型部署只是第一步,如何让AI模型与现有系统无缝集成、如何让AI理解业务上下文并调用正确的工具,这才是决定AI应用成败的关键。
今天我想分享的,正是基于Spring AI和Ollama的本地模型MCP集成方案。这个方案最大的优势在于,它不需要你重新发明轮子,而是利用现有的成熟框架,在5-10分钟内就能搭建起一个生产可用的MCP架构。无论你是想为内部系统添加智能助手,还是构建需要调用外部API的AI应用,这套方案都能提供清晰的路径。
1. 环境准备与架构设计
在开始编码之前,我们需要先理解整个架构的设计思路。传统的AI集成往往采用"一问一答"的模式,用户提问,模型回答,这种模式在处理复杂业务逻辑时显得力不从心。而MCP的核心思想是让AI模型能够理解上下文,并根据上下文动态调用合适的工具或服务。
1.1 技术栈选型考量
为什么选择Spring AI + Ollama这个组合?这里有几个关键考量点:
Spring AI的优势:
- 企业级支持:作为Spring生态的一部分,Spring AI天然支持Spring Boot的各种特性,如依赖注入、配置管理、监控等
- 标准化接口:提供了统一的AI模型接口,可以轻松切换不同的模型提供商
- 工具调用框架:内置了完善的工具调用机制,这正是MCP实现的基础
Ollama的价值:
- 本地化部署:数据不出本地,满足企业对数据安全性的严格要求
- 模型管理:统一的模型管理界面,支持多种开源模型
- 性能优化:针对本地运行进行了专门的优化,响应速度快
提示:如果你的应用对延迟要求极高,可以考虑将Ollama部署在与应用服务器同一物理机或同一数据中心内,这能显著减少网络延迟。
1.2 基础环境配置
首先确保你的开发环境已经准备好以下组件:
# 检查Java版本(需要JDK 17或更高)
java -version
# 安装Ollama(如果尚未安装)
# 访问Ollama官网获取适合你操作系统的安装包
# 启动Ollama服务
ollama serve
# 拉取需要的模型(这里以Qwen3:8b为例)
ollama pull qwen3:8b
环境配置完成后,我们可以通过简单的命令测试模型是否正常工作:
# 测试模型响应
curl http://localhost:11434/api/generate -d '{
"model": "qwen3:8b",
"prompt": "你好,请简单介绍一下自己",
"stream": false
}'
如果看到正常的JSON响应,说明Ollama服务已经就绪。接下来我们需要创建一个Spring Boot项目来集成这些能力。
2. Spring AI项目初始化与配置
创建一个新的Spring Boot项目是第一步,但更重要的是如何正确配置项目依赖和属性,确保整个架构能够顺畅运行。
2.1 项目依赖管理
在pom.xml中,我们需要添加Spring AI的相关依赖。这里的关键是版本管理,确保所有组件兼容:
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0
http://maven.apache.org/xsd/maven-4.0.0.xsd">
<parent>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-parent</artifactId>
<version>3.2.0</version>
</parent>
<dependencies>
<!-- Spring Web基础依赖 -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
</dependency>
<!-- Spring AI核心依赖 -->
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-ollama-spring-boot-starter</artifactId>
</dependency>
<!-- 工具调用支持 -->
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-tool</artifactId>
</dependency>
<!-- 测试依赖 -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-test</artifactId>
<scope>test</scope>
</dependency>
</dependencies>
<!-- Spring AI BOM管理 -->
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-bom</artifactId>
<version>1.0.0-M3</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
</project>
这个依赖配置有几个需要注意的地方:
- 版本对齐:Spring AI的版本需要与Spring Boot版本匹配,否则可能出现兼容性问题
- Starter选择:我们使用
spring-ai-ollama-spring-boot-starter而不是通用的starter,这样能获得更好的本地模型集成体验 - 工具依赖:
spring-ai-tool是MCP功能的核心,它提供了工具调用的基础框架
2.2 应用配置详解
在application.yml(或application.properties)中,我们需要配置Ollama连接和模型参数:
spring:
ai:
ollama:
# Ollama服务地址,默认运行在11434端口
base-url: http://localhost:11434
# 聊天模型配置
chat:
model: qwen3:8b
options:
temperature: 0.7
top-p: 0.9
max-tokens: 2048
# 嵌入模型配置(如果需要向量化功能)
embedding:
model: nomic-embed-text
enabled: true
# 工具调用相关配置
tool:
enabled: true
# 工具调用超时设置(毫秒)
call-timeout: 30000
# 应用服务配置
server:
port: 8080
servlet:
context-path: /api
# 日志配置,便于调试
logging:
level:
org.springframework.ai: DEBUG
com.example.mcp: INFO
配置中的几个关键参数说明:
| 参数 | 说明 | 推荐值 |
|---|---|---|
| temperature | 控制输出的随机性 | 0.7-0.9 |
| top-p | 核采样参数,影响多样性 | 0.8-0.95 |
| max-tokens | 单次响应的最大token数 | 根据需求调整 |
| call-timeout | 工具调用超时时间 | 30000ms |
注意:temperature值不宜设置过高,特别是在生产环境中,过高的随机性可能导致输出不稳定。对于需要确定性的业务场景,建议设置在0.3-0.5之间。
3. MCP核心实现:工具定义与注册
MCP的核心在于让AI模型能够理解和调用我们定义的工具。这不仅仅是技术实现,更是一种设计思维的转变——从"AI回答问题"到"AI使用工具解决问题"。
3.1 工具接口设计原则
在设计工具时,我总结了几条实践经验:
- 单一职责:每个工具只做一件事,并且做好
- 明确输入输出:参数和返回值类型要清晰明确
- 错误处理:工具内部要有完善的错误处理机制
- 文档完整:为每个工具提供清晰的说明文档
让我们从一个实际的天气查询工具开始。这个工具虽然简单,但包含了工具设计的核心要素:
import org.springframework.ai.tool.annotation.Tool;
import org.springframework.stereotype.Component;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
@Component
public class WeatherService {
private static final String WEATHER_API_URL = "https://api.weather.example.com";
private final HttpClient httpClient;
private final ObjectMapper objectMapper;
public WeatherService() {
this.httpClient = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(10))
.build();
this.objectMapper = new ObjectMapper();
}
@Tool(
name = "getWeather",
description = "根据城市名称获取当前天气信息,包括温度、湿度、天气状况等",
inputSchema = @Tool.InputSchema(
properties = {
@Tool.Property(name = "city",
type = "string",
description = "城市名称,如'北京'、'上海'")
}
)
)
public WeatherInfo getWeather(String city) {
try {
// 构建API请求
String apiUrl = String.format("%s/current?city=%s",
WEATHER_API_URL,
java.net.URLEncoder.encode(city, "UTF-8"));
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create(apiUrl))
.header("Accept", "application/json")
.timeout(Duration.ofSeconds(5))
.GET()
.build();
// 发送请求并解析响应
HttpResponse<String> response = httpClient.send(
request,
HttpResponse.BodyHandlers.ofString()
);
if (response.statusCode() == 200) {
return objectMapper.readValue(response.body(), WeatherInfo.class);
} else {
throw new RuntimeException("天气API调用失败: " + response.statusCode());
}
} catch (Exception e) {
// 返回一个包含错误信息的天气对象
return WeatherInfo.error(city, e.getMessage());
}
}
// 天气信息数据类
public static class WeatherInfo {
private String city;
private double temperature;
private int humidity;
private String condition;
private String lastUpdated;
private String error;
// 构造方法、getter/setter省略
public static WeatherInfo error(String city, String message) {
WeatherInfo info = new WeatherInfo();
info.setCity(city);
info.setError("获取天气信息失败: " + message);
return info;
}
@Override
public String toString() {
if (error != null) {
return String.format("%s的天气信息暂时无法获取: %s", city, error);
}
return String.format("%s当前天气: %.1f°C, 湿度%d%%, %s。更新时间: %s",
city, temperature, humidity, condition, lastUpdated);
}
}
}
这个工具实现有几个值得注意的设计点:
- 完善的错误处理:即使在API调用失败时,也返回结构化的错误信息
- 清晰的文档:通过注解提供了完整的工具描述和参数说明
- 合理的超时设置:避免因外部服务问题导致整个系统阻塞
3.2 复杂工具链设计
在实际业务中,我们往往需要多个工具协同工作。比如一个电商客服场景,可能需要查询订单、检查库存、计算运费等多个工具。这时候,工具链的设计就变得尤为重要。
@Component
public class ECommerceTools {
@Tool(
name = "searchProducts",
description = "根据关键词搜索商品,支持按价格、销量等排序",
inputSchema = @Tool.InputSchema(
properties = {
@Tool.Property(name = "keyword", type = "string"),
@Tool.Property(name = "category", type = "string", required = false),
@Tool.Property(name = "minPrice", type = "number", required = false),
@Tool.Property(name = "maxPrice", type = "number", required = false),
@Tool.Property(name = "sortBy", type = "string", required = false)
}
)
)
public List<Product> searchProducts(String keyword,
String category,
Double minPrice,
Double maxPrice,
String sortBy) {
// 实现商品搜索逻辑
return productService.search(keyword, category, minPrice, maxPrice, sortBy);
}
@Tool(
name = "checkInventory",
description = "检查商品库存状态",
inputSchema = @Tool.InputSchema(
properties = {
@Tool.Property(name = "productId", type = "string"),
@Tool.Property(name = "warehouseId", type = "string", required = false)
}
)
)
public InventoryStatus checkInventory(String productId, String warehouseId) {
// 实现库存检查逻辑
return inventoryService.getStatus(productId, warehouseId);
}
@Tool(
name = "calculateShipping",
description = "计算商品运费和预计送达时间",
inputSchema = @Tool.InputSchema(
properties = {
@Tool.Property(name = "productIds", type = "array"),
@Tool.Property(name = "destination", type = "string"),
@Tool.Property(name = "shippingMethod", type = "string", required = false)
}
)
)
public ShippingQuote calculateShipping(List<String> productIds,
String destination,
String shippingMethod) {
// 实现运费计算逻辑
return shippingService.calculate(productIds, destination, shippingMethod);
}
}
工具链设计的关键在于工具间的信息传递和执行顺序管理。Spring AI提供了ToolCall机制,可以自动处理工具间的依赖关系。
4. 模型调用与上下文管理
有了工具定义,下一步就是让AI模型学会使用这些工具。这不仅仅是技术实现,更涉及到提示工程和上下文管理的艺术。
4.1 智能提示词设计
提示词的质量直接决定了模型能否正确理解用户意图并调用合适的工具。我发现在实际项目中,一个结构良好的系统提示词能显著提升工具调用的准确率。
@Service
public class AIChatService {
private final ChatClient chatClient;
private final ToolCallbackProvider toolCallbackProvider;
public AIChatService(ChatClient chatClient,
ToolCallbackProvider toolCallbackProvider) {
this.chatClient = chatClient;
this.toolCallbackProvider = toolCallbackProvider;
}
public String chatWithTools(String userMessage, String conversationId) {
// 构建系统提示词
String systemPrompt = """
你是一个智能助手,可以调用各种工具来帮助用户解决问题。
可用工具:
1. getWeather - 查询城市天气
2. searchProducts - 搜索商品
3. checkInventory - 检查库存
4. calculateShipping - 计算运费
使用规则:
1. 仔细分析用户问题,判断是否需要调用工具
2. 如果需要调用工具,请明确说明要调用哪个工具以及参数
3. 工具调用结果会以结构化数据返回,请用自然语言解释给用户
4. 如果用户问题涉及多个步骤,可以按顺序调用多个工具
5. 如果工具调用失败,向用户友好地解释并建议替代方案
当前对话ID:%s
请保持对话的连贯性,记住之前的对话内容。
""".formatted(conversationId);
// 构建消息列表
List<Message> messages = new ArrayList<>();
messages.add(new SystemMessage(systemPrompt));
// 这里可以添加历史消息,维护对话上下文
messages.add(new UserMessage(userMessage));
// 调用模型,启用工具调用
ChatResponse response = chatClient.prompt()
.messages(messages)
.tools(toolCallbackProvider.getTools())
.call();
return response.getResult().getOutput().getContent();
}
}
这个提示词设计有几个关键点:
- 明确工具列表:让模型知道有哪些工具可用
- 定义使用规则:告诉模型如何正确使用工具
- 维护对话上下文:通过conversationId保持对话连贯性
- 错误处理指导:告诉模型在工具失败时该如何应对
4.2 上下文管理策略
在实际对话中,用户可能会连续提问,这就需要我们维护对话的上下文。Spring AI提供了多种上下文管理策略:
@Component
public class ConversationManager {
private final Map<String, List<Message>> conversationStore =
new ConcurrentHashMap<>();
private static final int MAX_HISTORY = 10; // 最大历史消息数
/**
* 添加消息到对话历史
*/
public void addMessage(String conversationId, Message message) {
List<Message> history = conversationStore
.computeIfAbsent(conversationId, k -> new ArrayList<>());
history.add(message);
// 限制历史消息数量,避免token超限
if (history.size() > MAX_HISTORY * 2) { // 乘以2因为包含用户和助手消息
history = history.subList(history.size() - MAX_HISTORY * 2, history.size());
conversationStore.put(conversationId, history);
}
}
/**
* 获取对话历史
*/
public List<Message> getHistory(String conversationId) {
return conversationStore.getOrDefault(conversationId, new ArrayList<>());
}
/**
* 清理过期的对话
*/
@Scheduled(fixedDelay = 3600000) // 每小时清理一次
public void cleanupOldConversations() {
// 实现基于时间的清理逻辑
}
}
上下文管理需要考虑的几个实际问题:
- Token限制:模型有token限制,历史消息不能无限保存
- 内存管理:大量对话历史会占用内存,需要定期清理
- 相关性过滤:不是所有历史消息都有用,可能需要智能过滤
4.3 工具调用优化
当模型决定调用工具时,我们需要确保调用过程高效可靠。这里分享一些优化经验:
@Service
public class OptimizedToolService {
private final Map<String, ToolExecutor> toolExecutors;
private final CircuitBreaker circuitBreaker;
public OptimizedToolService(List<ToolExecutor> executors) {
// 初始化工具执行器映射
this.toolExecutors = executors.stream()
.collect(Collectors.toMap(
e -> e.getToolName(),
e -> e
));
// 配置熔断器
this.circuitBreaker = CircuitBreaker.ofDefaults("toolCall");
}
public ToolResponse executeTool(ToolCall toolCall) {
return circuitBreaker.executeSupplier(() -> {
ToolExecutor executor = toolExecutors.get(toolCall.getName());
if (executor == null) {
return ToolResponse.error("工具不存在: " + toolCall.getName());
}
long startTime = System.currentTimeMillis();
try {
Object result = executor.execute(toolCall.getArguments());
long duration = System.currentTimeMillis() - startTime;
// 记录性能指标
recordMetrics(toolCall.getName(), duration, true);
return ToolResponse.success(result);
} catch (Exception e) {
long duration = System.currentTimeMillis() - startTime;
recordMetrics(toolCall.getName(), duration, false);
log.error("工具调用失败: {}", toolCall.getName(), e);
return ToolResponse.error("工具执行失败: " + e.getMessage());
}
});
}
private void recordMetrics(String toolName, long duration, boolean success) {
// 记录到监控系统
Metrics.counter("tool_calls_total",
"tool", toolName,
"success", String.valueOf(success))
.increment();
Metrics.timer("tool_call_duration", "tool", toolName)
.record(duration, TimeUnit.MILLISECONDS);
}
}
这个优化实现包含了几个重要特性:
- 熔断保护:防止某个工具故障影响整个系统
- 性能监控:记录每个工具的执行时间和成功率
- 错误隔离:工具失败不影响其他功能
5. 高级特性与生产部署
当基础功能实现后,我们需要考虑如何让这个系统在生产环境中稳定运行。这里涉及到性能优化、监控告警、安全防护等多个方面。
5.1 性能优化策略
本地模型调用虽然避免了网络延迟,但仍然有性能优化的空间。以下是一些实测有效的优化方法:
批量处理优化:
@Service
public class BatchProcessingService {
private final ChatClient chatClient;
private final ExecutorService executorService;
public BatchProcessingService() {
// 根据CPU核心数配置线程池
int coreCount = Runtime.getRuntime().availableProcessors();
this.executorService = Executors.newFixedThreadPool(coreCount * 2);
}
/**
* 批量处理用户请求
*/
public List<String> batchProcess(List<UserRequest> requests) {
List<CompletableFuture<String>> futures = requests.stream()
.map(request -> CompletableFuture.supplyAsync(
() -> processSingleRequest(request),
executorService
))
.collect(Collectors.toList());
// 等待所有请求完成
CompletableFuture.allOf(futures.toArray(new CompletableFuture[0]))
.join();
return futures.stream()
.map(CompletableFuture::join)
.collect(Collectors.toList());
}
private String processSingleRequest(UserRequest request) {
// 单个请求处理逻辑
return chatClient.prompt()
.user(request.getMessage())
.call()
.getResult()
.getOutput()
.getContent();
}
@PreDestroy
public void shutdown() {
executorService.shutdown();
}
}
缓存策略实现:
@Component
@CacheConfig(cacheNames = "aiResponses")
public class ResponseCacheService {
@Cacheable(key = "#userMessage.hashCode() + '-' + #contextHash")
public String getCachedResponse(String userMessage, String contextHash) {
// 如果缓存命中,直接返回
// 否则返回null,触发实际处理
return null;
}
@CachePut(key = "#userMessage.hashCode() + '-' + #contextHash")
public String cacheResponse(String userMessage,
String contextHash,
String response) {
return response;
}
/**
* 计算上下文哈希,用于缓存键
*/
public String computeContextHash(List<Message> conversationHistory) {
// 基于历史消息计算哈希值
String historyString = conversationHistory.stream()
.map(m -> m.getContent())
.collect(Collectors.joining("||"));
return DigestUtils.md5DigestAsHex(historyString.getBytes());
}
}
5.2 监控与告警体系
生产环境必须要有完善的监控。以下是一些关键的监控指标:
| 监控类别 | 具体指标 | 告警阈值 | 处理建议 |
|---|---|---|---|
| 服务可用性 | 服务健康检查 | 连续3次失败 | 检查Ollama服务状态 |
| 性能指标 | 平均响应时间 | > 5秒 | 优化提示词或升级硬件 |
| 性能指标 | P95响应时间 | > 10秒 | 检查是否有慢查询 |
| 业务指标 | 工具调用成功率 | < 95% | 检查工具服务状态 |
| 业务指标 | 用户满意度 | < 4.0/5.0 | 优化提示词或工具设计 |
| 资源使用 | 内存使用率 | > 80% | 增加内存或优化模型 |
| 资源使用 | GPU使用率 | > 90% | 考虑模型量化或分流 |
实现监控的代码示例:
@RestController
@RequestMapping("/metrics")
public class MetricsController {
private final MeterRegistry meterRegistry;
@GetMapping("/health")
public ResponseEntity<Map<String, Object>> healthCheck() {
Map<String, Object> health = new HashMap<>();
// 检查Ollama连接
boolean ollamaHealthy = checkOllamaHealth();
health.put("ollama", ollamaHealthy ? "UP" : "DOWN");
// 检查模型加载状态
boolean modelLoaded = checkModelStatus();
health.put("model", modelLoaded ? "LOADED" : "UNLOADED");
// 检查工具服务
boolean toolsHealthy = checkToolsHealth();
health.put("tools", toolsHealthy ? "HEALTHY" : "UNHEALTHY");
HttpStatus status = ollamaHealthy && modelLoaded && toolsHealthy
? HttpStatus.OK : HttpStatus.SERVICE_UNAVAILABLE;
return ResponseEntity.status(status).body(health);
}
@GetMapping("/performance")
public Map<String, Object> performanceMetrics() {
Map<String, Object> metrics = new HashMap<>();
// 获取响应时间指标
Timer responseTimer = meterRegistry.find("ai.response.time").timer();
if (responseTimer != null) {
metrics.put("avgResponseTime", responseTimer.mean());
metrics.put("p95ResponseTime", responseTimer.percentile(0.95));
}
// 获取工具调用指标
Counter successCounter = meterRegistry.find("tool.calls.success").counter();
Counter totalCounter = meterRegistry.find("tool.calls.total").counter();
if (totalCounter != null && successCounter != null) {
double successRate = totalCounter.count() > 0
? successCounter.count() / totalCounter.count()
: 1.0;
metrics.put("toolSuccessRate", successRate);
}
return metrics;
}
}
5.3 安全防护措施
AI系统的安全同样重要,特别是当它能够调用外部工具时:
@Configuration
public class SecurityConfig {
@Bean
public ToolCallValidator toolCallValidator() {
return new ToolCallValidator() {
@Override
public ValidationResult validate(ToolCall toolCall,
UserContext userContext) {
// 检查工具调用权限
if (!hasPermission(userContext, toolCall.getName())) {
return ValidationResult.rejected(
"用户没有调用工具" + toolCall.getName() + "的权限"
);
}
// 检查参数安全性
if (containsSensitiveData(toolCall.getArguments())) {
return ValidationResult.rejected(
"工具调用包含敏感数据"
);
}
// 检查调用频率
if (exceedsRateLimit(userContext)) {
return ValidationResult.rejected(
"调用频率过高,请稍后再试"
);
}
return ValidationResult.approved();
}
private boolean hasPermission(UserContext context, String toolName) {
// 实现权限检查逻辑
return true;
}
private boolean containsSensitiveData(Map<String, Object> args) {
// 检查是否包含敏感信息
return false;
}
private boolean exceedsRateLimit(UserContext context) {
// 实现频率限制逻辑
return false;
}
};
}
@Bean
public InputSanitizer inputSanitizer() {
return new InputSanitizer() {
@Override
public String sanitize(String input) {
// 移除潜在的恶意内容
String sanitized = input;
// 移除HTML标签
sanitized = sanitized.replaceAll("<[^>]*>", "");
// 移除可能的脚本注入
sanitized = sanitized.replaceAll("(?i)javascript:", "");
sanitized = sanitized.replaceAll("(?i)onload=", "");
sanitized = sanitized.replaceAll("(?i)onerror=", "");
// 限制输入长度
if (sanitized.length() > 10000) {
sanitized = sanitized.substring(0, 10000);
}
return sanitized;
}
};
}
}
安全防护需要多层次的策略:
- 输入验证:所有用户输入都要经过严格的验证和清洗
- 权限控制:确保用户只能调用有权限的工具
- 频率限制:防止滥用和DDoS攻击
- 输出过滤:对AI的输出也要进行安全检查
5.4 部署与扩展考虑
当应用准备好投入生产时,部署架构也需要精心设计:
# docker-compose.prod.yml
version: '3.8'
services:
ollama:
image: ollama/ollama:latest
container_name: ollama-service
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
networks:
- ai-network
spring-ai-app:
build: .
container_name: spring-ai-app
ports:
- "8080:8080"
environment:
- SPRING_AI_OLLAMA_BASE_URL=http://ollama:11434
- SPRING_PROFILES_ACTIVE=prod
depends_on:
- ollama
networks:
- ai-network
deploy:
replicas: 3
restart_policy:
condition: on-failure
nginx:
image: nginx:alpine
container_name: nginx-lb
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf
depends_on:
- spring-ai-app
networks:
- ai-network
volumes:
ollama_data:
networks:
ai-network:
driver: bridge
这个部署配置考虑了:
- 服务发现:通过Docker网络让服务间可以互相访问
- 负载均衡:使用Nginx分发请求到多个Spring AI实例
- GPU支持:为Ollama容器配置GPU加速
- 数据持久化:确保模型数据不会丢失
- 高可用:Spring AI应用有多个副本
在实际部署中,我还发现几个值得注意的点:
- 模型预热:服务启动时自动加载常用模型,避免第一次调用时的延迟
- 连接池管理:合理配置HTTP连接池,避免连接泄露
- 日志聚合:使用ELK或类似方案集中管理日志
- 配置中心:将配置外置,便于不同环境的管理
整个架构搭建完成后,你会发现原本复杂的AI集成变得清晰可控。Spring AI提供了标准的接口和最佳实践,Ollama让本地模型运行变得简单,而MCP模式则让AI真正成为了业务系统的一部分。
我在实际项目中采用这套架构后,最大的感受是开发效率的提升。以前需要花费数天甚至数周才能完成的AI功能集成,现在可以在几小时内完成原型开发。更重要的是,这套架构具有良好的扩展性——当需要添加新的AI能力时,只需要定义新的工具并更新提示词,而不需要修改核心架构。
当然,每个项目都有其特殊性,这套方案也需要根据具体需求进行调整。比如对于实时性要求极高的场景,可能需要考虑模型量化或硬件加速;对于数据敏感性强的场景,可能需要加强安全审计和访问控制。但无论如何,Spring AI + Ollama + MCP这个技术栈,为本地AI应用开发提供了一个坚实而灵活的基础。
更多推荐
所有评论(0)