OpenMontage HeyGen 照片数字人(Talking Photos)实战:从肖像上传、Avatar Group 创建到 Avatar IV 与 AI 生成头像

【免费下载链接】OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio. 【免费下载链接】OpenMontage 项目地址: https://gitcode.com/GitHub_Trending/op/OpenMontage

本文基于 OpenMontage 仓库中 avatar-video 技能下的参考文档 photo-avatars.md,系统讲解如何用 HeyGen API 把一张静态照片变成会说话的数字人:覆盖「上传图片 → 创建照片头像组 → 轮询处理状态 → 生成视频」的完整链路、Avatar IV 一步直出方案、AI 文本生成照片头像的参数约束,以及头像组的管理接口与照片质量要求。读完后,你可以在 OpenMontage 的头像视频流水线中独立完成照片数字人从资产创建到视频产出的全部操作。

照片数字人在 OpenMontage 中的位置

OpenMontage 把 HeyGen 集成在 avatar-video 技能中(见 SKILL.md),该技能面向「精确控制」场景:自己选头像、写精确脚本、配置每个场景。其 frontmatter 声明了运行前提——需要 HEYGEN_API_KEY 环境变量,且所有请求都要求 X-Api-Key 请求头:

curl -X GET "https://api.heygen.com/v2/avatars" \
  -H "X-Api-Key: $HEYGEN_API_KEY"

照片数字人(Photo Avatars / Talking Photos)是 avatar-video 技能「Advanced Features」分支之一:SKILL.md 的 Quick Reference 表中将 “Create avatar from photo” 指向本文档。它与两个基础文档协同:

  • assets.md:资产上传端点 POST https://upload.heygen.com/v1/asset 的完整字段说明,是照片数字人流程第 1 步的基础;
  • video-generation.md:/v2/video/generate 端点的 character 字段定义,其中 character.type 只允许 "avatar" 或 "talking_photo",当 type 为 talking_photo 时 talking_photo_id 必填——这正是照片数字人产物的消费入口。

在仓库的工具层,tools/video/heygen_video.py 中的 HeyGenVideo 工具封装了 HeyGen 云端视频生成(install_instructions 明确提示设置 HEYGEN_API_KEY,并配置了 wan_video、hunyuan_video 等回退工具);tools/avatar/talking_head.py 中的 TalkingHead 工具则提供本地化的「照片转说话头像」(photo_to_video 能力)。本文档讲解的 API 链路是技能层直接调用 HeyGen 端点的完整参考实现。

流程一:从上传照片到生成视频(四步链路)

官方工作流为:Upload Image → Create Avatar Group → Use in Video。创建出来的照片头像 id(与 group_id 相同)就是视频生成时的 talking_photo_id。

Step 1:上传图片,获取 image_key

把肖像照以原始二进制 POST 到资产上传端点,Content-Type 必须与文件 MIME 类型一致:

curl -X POST "https://upload.heygen.com/v1/asset" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: image/jpeg" \
  --data-binary '@./portrait.jpg'

响应示例:

{
  "code": 100,
  "data": {
    "id": "741299e941764988b432ed3a6757878f",
    "name": "741299e941764988b432ed3a6757878f",
    "file_type": "image",
    "url": "https://resource2.heygen.ai/image/.../original.jpg",
    "image_key": "image/741299e941764988b432ed3a6757878f/original.jpg"
  }
}

注意:必须保存 image_key 字段(而不是 id)。image_key 是创建照片头像时使用的 S3 路径。结合 assets.md 的完整字段说明,上传响应中 code: 100 表示成功,data.image_key 为「仅图片类型」才有值的字段(string | null),用于创建上传类照片头像;资产上限为 10MB,且闲置资产可能被服务端清理。

Step 2:创建照片头像组(Avatar Group)

用 image_key 调 POST https://api.heygen.com/v2/photo_avatar/avatar_group/create,服务端会处理该图片并产出一个可复用的照片头像:

curl -X POST "https://api.heygen.com/v2/photo_avatar/avatar_group/create" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image_key": "image/741299e941764988b432ed3a6757878f/original.jpg",
    "name": "My Photo Avatar"
  }'
字段类型必填说明
image_keystring✓上传响应中的 S3 图片 key
namestring✓头像显示名称
generation_idstring 若使用 AI 生成的照片(见下文「AI 照片头像」一节),需携带生成任务 ID

响应示例:

{
  "error": null,
  "data": {
    "id": "045c260bc0364727b2cbe50442c3a5bf",
    "image_url": "https://files2.heygen.ai/...",
    "created_at": 1771798135.777256,
    "name": "My Photo Avatar",
    "status": "pending",
    "group_id": "045c260bc0364727b2cbe50442c3a5bf",
    "is_motion": false,
    "business_type": "uploaded"
  }
}

返回的 id(与 group_id 相同)即视频生成所需的 talking_photo_id。

Step 3:轮询等待处理完成

照片头像初始状态为 status: "pending",通常数秒内转为 "completed"。状态查询端点:

Endpoint: GET https://api.heygen.com/v2/photo_avatar/{id}

curl "https://api.heygen.com/v2/photo_avatar/045c260bc0364727b2cbe50442c3a5bf" \
  -H "X-Api-Key: $HEYGEN_API_KEY"

在 status 变为 "completed" 之前不要发起视频生成。

Step 4:用于视频生成

拿到 talking_photo_id 后,构造 /v2/video/generate 的请求体,character.type 设为 "talking_photo":

const videoConfig = {
  video_inputs: [
    {
      character: {
        type: "talking_photo",
        talking_photo_id: "045c260bc0364727b2cbe50442c3a5bf",
      },
      voice: {
        type: "text",
        input_text: "Hello! This is my photo avatar speaking.",
        voice_id: "1bd001e7e50f421d891986aad5158bc8",
      },
    },
  ],
  dimension: { width: 1920, height: 1080 },
};

对照 video-generation.md 的字段表,talking_photo 类型下还可用可选字段 avatar_style("normal" / "closeUp" / "circle")、scale(缩放系数)和 offset({x, y} 位置偏移)来调整人物在画面中的呈现;voice 支持 text / audio / silence 三种输入类型,text 时 voice_id 与 input_text 必填。

完整工作流参考实现(TypeScript / Python)

TypeScript 版

import fs from "fs";
import path from "path";

interface AssetUploadResponse {
  code: number;
  data: {
    id: string;
    image_key: string;
    url: string;
  };
}

interface PhotoAvatarResponse {
  error: string | null;
  data: {
    id: string;
    group_id: string;
    image_url: string;
    name: string;
    status: string;
    is_motion: boolean;
    business_type: string;
  };
}

async function createPhotoAvatar(
  imagePath: string,
  name: string
): Promise<string> {
  // 1. 上传图片
  const resolvedPath = path.resolve(imagePath);
  const fileBuffer = fs.readFileSync(resolvedPath);
  const uploadResponse = await fetch("https://upload.heygen.com/v1/asset", {
    method: "POST",
    headers: {
      "X-Api-Key": process.env.HEYGEN_API_KEY!,
      "Content-Type": "image/jpeg",
    },
    body: fileBuffer,
  });

  const uploadJson: AssetUploadResponse = await uploadResponse.json();
  if (uploadJson.code !== 100) {
    throw new Error("Upload failed");
  }

  const imageKey = uploadJson.data.image_key;

  // 2. 创建头像组
  const createResponse = await fetch(
    "https://api.heygen.com/v2/photo_avatar/avatar_group/create",
    {
      method: "POST",
      headers: {
        "X-Api-Key": process.env.HEYGEN_API_KEY!,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({ image_key: imageKey, name }),
    }
  );

  const createJson: PhotoAvatarResponse = await createResponse.json();
  if (createJson.error) {
    throw new Error(createJson.error);
  }

  const photoAvatarId = createJson.data.id;

  // 3. 等待处理完成
  await waitForPhotoAvatar(photoAvatarId);

  return photoAvatarId;
}

async function waitForPhotoAvatar(id: string): Promise<void> {
  for (let i = 0; i < 30; i++) {
    const response = await fetch(
      `https://api.heygen.com/v2/photo_avatar/${id}`,
      { headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! } }
    );

    const json: PhotoAvatarResponse = await response.json();

    if (json.data.status === "completed") return;
    if (json.data.status === "failed") {
      throw new Error("Photo avatar processing failed");
    }

    await new Promise((r) => setTimeout(r, 2000));
  }

  throw new Error("Photo avatar processing timed out");
}

async function createVideoFromPhoto(
  photoPath: string,
  script: string,
  voiceId: string
): Promise<string> {
  // 1. 创建照片头像
  const talkingPhotoId = await createPhotoAvatar(photoPath, "Video Avatar");

  // 2. 生成视频
  const response = await fetch("https://api.heygen.com/v2/video/generate", {
    method: "POST",
    headers: {
      "X-Api-Key": process.env.HEYGEN_API_KEY!,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      video_inputs: [
        {
          character: {
            type: "talking_photo",
            talking_photo_id: talkingPhotoId,
          },
          voice: {
            type: "text",
            input_text: script,
            voice_id: voiceId,
          },
        },
      ],
      dimension: { width: 1920, height: 1080 },
    }),
  });

  const { data } = await response.json();
  return data.video_id;
}

轮询策略从源码结构看是「最多 30 次、每次间隔 2 秒」,即约 60 秒超时窗口,pending 之外的终态只有 completed 与 failed。

Python 版

import requests
import os
import time

def create_photo_avatar(image_path: str, name: str) -> str:
    api_key = os.environ["HEYGEN_API_KEY"]

    # 1. 上传图片
    with open(image_path, "rb") as f:
        upload_resp = requests.post(
            "https://upload.heygen.com/v1/asset",
            headers={
                "X-Api-Key": api_key,
                "Content-Type": "image/jpeg",
            },
            data=f,
        )

    upload_data = upload_resp.json()
    if upload_data.get("code") != 100:
        raise Exception("Upload failed")

    image_key = upload_data["data"]["image_key"]

    # 2. 创建头像组
    create_resp = requests.post(
        "https://api.heygen.com/v2/photo_avatar/avatar_group/create",
        headers={
            "X-Api-Key": api_key,
            "Content-Type": "application/json",
        },
        json={"image_key": image_key, "name": name},
    )

    create_data = create_resp.json()
    if create_data.get("error"):
        raise Exception(create_data["error"])

    photo_avatar_id = create_data["data"]["id"]

    # 3. 等待处理完成
    for _ in range(30):
        status_resp = requests.get(
            f"https://api.heygen.com/v2/photo_avatar/{photo_avatar_id}",
            headers={"X-Api-Key": api_key},
        )
        status = status_resp.json()["data"]["status"]
        if status == "completed":
            return photo_avatar_id
        if status == "failed":
            raise Exception("Photo avatar processing failed")
        time.sleep(2)

    raise Exception("Photo avatar processing timed out")

Avatar IV:跳过头像组、直出视频

Avatar IV 是 HeyGen 最新的照片数字人技术,直接以「上传的 image_key + 脚本 + 音色」生成视频,绕过 avatar group 创建步骤,适合不需要跨视频复用同一头像的一次性生产。

Endpoint: POST https://api.heygen.com/v2/video/av4/generate

curl -X POST "https://api.heygen.com/v2/video/av4/generate" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image_key": "image/741299e941764988b432ed3a6757878f/original.jpg",
    "script": "Hello! This is Avatar IV with enhanced quality.",
    "voice_id": "1bd001e7e50f421d891986aad5158bc8",
    "video_orientation": "landscape",
    "video_title": "My Avatar IV Video"
  }'
字段类型必填说明
image_keystring✓资产上传得到的 S3 图片 key
scriptstring✓数字人要说出的文本
voice_idstring✓使用的音色
video_orientationstring "portrait" / "landscape" / "square"
video_titlestring 视频标题
fitstring "cover" 或 "contain"
custom_motion_promptstring 动作/表情描述
enhance_custom_motion_promptboolean是否用 AI 增强动作提示词

TypeScript 封装:

interface AvatarIVRequest {
  image_key: string;
  script: string;
  voice_id: string;
  video_orientation?: "portrait" | "landscape" | "square";
  video_title?: string;
  fit?: "cover" | "contain";
  custom_motion_prompt?: string;
  enhance_custom_motion_prompt?: boolean;
}

interface AvatarIVResponse {
  error: null | string;
  data: {
    video_id: string;
  };
}

async function generateAvatarIVVideo(
  config: AvatarIVRequest
): Promise<string> {
  const response = await fetch(
    "https://api.heygen.com/v2/video/av4/generate",
    {
      method: "POST",
      headers: {
        "X-Api-Key": process.env.HEYGEN_API_KEY!,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(config),
    }
  );

  const json: AvatarIVResponse = await response.json();

  if (json.error) {
    throw new Error(json.error);
  }

  return json.data.video_id;
}

画幅方向与 Fit 取值

Orientation尺寸适用场景
portrait720x1280TikTok、Stories
landscape1280x720YouTube、Web
square720x720Instagram Feed
Fit行为
cover铺满画面,可能裁掉边缘
contain完整放入画面,可能露出背景

自定义动作提示词

custom_motion_prompt 用于控制数字人的动作与表情,可配合 enhance_custom_motion_prompt: true 让 AI 扩写提示词:

const videoId = await generateAvatarIVVideo({
  image_key: "image/.../original.jpg",
  script: "Let me tell you about our product.",
  voice_id: "1bd001e7e50f421d891986aad5158bc8",
  custom_motion_prompt: "nodding head and smiling",
  enhance_custom_motion_prompt: true,
});

两种方案怎么选:需要多个视频复用同一张照片形象(并叠加训练)时走 Avatar Group 流程;单条视频、追求最新质量时直接用 Avatar IV。

流程二:AI 生成照片头像(无需本地照片)

当手头没有合适的肖像照时,可以用文本描述直接生成合成照片,再喂给数字人管线。

Endpoint: POST https://api.heygen.com/v2/photo_avatar/photo/generate

重要:以下 8 个字段全部必填。 API 会拒绝缺失任何字段的请求。当用户只给出「生成一个职业男性形象」这类模糊需求时,需要追问或为缺失字段选定合理默认值。

必填字段

字段类型允许值
namestring生成头像的名称
ageenum"Young Adult" / "Early Middle Age" / "Late Middle Age" / "Senior" / "Unspecified"
genderenum"Woman" / "Man" / "Unspecified"
ethnicityenum"White" / "Black" / "Asian American" / "East Asian" / "South East Asian" / "South Asian" / "Middle Eastern" / "Pacific" / "Hispanic" / "Unspecified"
orientationenum"square" / "horizontal" / "vertical"
poseenum"half_body" / "close_up" / "full_body"
styleenum"Realistic" / "Pixar" / "Cinematic" / "Vintage" / "Noir" / "Cyberpunk" / "Unspecified"
appearancestring外观文本提示词(服装、氛围、灯光等),最长 1000 字符

curl 示例:

curl -X POST "https://api.heygen.com/v2/photo_avatar/photo/generate" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Sarah Product Demo",
    "age": "Young Adult",
    "gender": "Woman",
    "ethnicity": "White",
    "orientation": "horizontal",
    "pose": "half_body",
    "style": "Realistic",
    "appearance": "Professional woman with a friendly smile, wearing a navy blue blazer over a white blouse, soft studio lighting, clean neutral background"
  }'

响应只返回生成任务 ID:

{
  "error": null,
  "data": {
    "generation_id": "6a7f7f2795de4599bec7cf1e06babe30"
  }
}

查询生成状态

Endpoint: GET https://api.heygen.com/v2/photo_avatar/generation/{generation_id}

成功后会返回多张候选图(URL 与 image_key 成对给出):

{
  "error": null,
  "data": {
    "id": "6a7f7f2795de4599bec7cf1e06babe30",
    "status": "success",
    "image_url_list": [
      "https://resource2.heygen.ai/photo_generation/.../image1.jpg",
      "https://resource2.heygen.ai/photo_generation/.../image2.jpg",
      "https://resource2.heygen.ai/photo_generation/.../image3.jpg",
      "https://resource2.heygen.ai/photo_generation/.../image4.jpg"
    ],
    "image_key_list": [
      "photo_generation/.../image1.jpg",
      "photo_generation/.../image2.jpg",
      "photo_generation/.../image3.jpg",
      "photo_generation/.../image4.jpg"
    ]
  }
}

状态机为 pending / processing / success / failed。TypeScript 参考实现:

interface GeneratePhotoAvatarRequest {
  name: string;
  age: "Young Adult" | "Early Middle Age" | "Late Middle Age" | "Senior" | "Unspecified";
  gender: "Woman" | "Man" | "Unspecified";
  ethnicity: "White" | "Black" | "Asian American" | "East Asian" | "South East Asian" | "South Asian" | "Middle Eastern" | "Pacific" | "Hispanic" | "Unspecified";
  orientation: "square" | "horizontal" | "vertical";
  pose: "half_body" | "close_up" | "full_body";
  style: "Realistic" | "Pixar" | "Cinematic" | "Vintage" | "Noir" | "Cyberpunk" | "Unspecified";
  appearance: string;
}

interface GeneratePhotoAvatarResponse {
  error: string | null;
  data: {
    generation_id: string;
  };
}

interface PhotoGenerationStatus {
  error: string | null;
  data: {
    id: string;
    status: "pending" | "processing" | "success" | "failed";
    msg: string | null;
    image_url_list?: string[];
    image_key_list?: string[];
  };
}

async function generatePhotoAvatar(
  config: GeneratePhotoAvatarRequest
): Promise<string> {
  const response = await fetch(
    "https://api.heygen.com/v2/photo_avatar/photo/generate",
    {
      method: "POST",
      headers: {
        "X-Api-Key": process.env.HEYGEN_API_KEY!,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(config),
    }
  );

  const json: GeneratePhotoAvatarResponse = await response.json();

  if (json.error) {
    throw new Error(`Photo avatar generation failed: ${json.error}`);
  }

  return json.data.generation_id;
}

async function waitForPhotoGeneration(
  generationId: string
): Promise<string[]> {
  for (let i = 0; i < 60; i++) {
    const response = await fetch(
      `https://api.heygen.com/v2/photo_avatar/generation/${generationId}`,
      { headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! } }
    );

    const json: PhotoGenerationStatus = await response.json();

    if (json.error) throw new Error(json.error);

    if (json.data.status === "success") {
      return json.data.image_key_list!;
    }

    if (json.data.status === "failed") {
      throw new Error(json.data.msg ?? "Photo generation failed");
    }

    await new Promise((r) => setTimeout(r, 5000));
  }

  throw new Error("Photo generation timed out");
}

从实现看,照片生成的轮询比头像处理更保守:60 次、间隔 5 秒(约 5 分钟窗口),因为文生图任务耗时更长。

AI 照片 → 头像组 → 视频的组合链路

选用某张 AI 生成图后,把它当作 image_key 交给 avatar group 创建接口,并额外携带 generation_id:

// 1. 生成 AI 照片
const generationId = await generatePhotoAvatar({
  name: "Product Demo Host",
  age: "Young Adult",
  gender: "Woman",
  ethnicity: "Unspecified",
  orientation: "horizontal",
  pose: "half_body",
  style: "Realistic",
  appearance: "Professional woman, navy blazer, friendly smile, soft lighting",
});

// 2. 等待生成完成并取第一张结果
const imageKeys = await waitForPhotoGeneration(generationId);
const selectedImageKey = imageKeys[0];

// 3. 用 AI 照片创建头像组(注意带上 generation_id)
const createResponse = await fetch(
  "https://api.heygen.com/v2/photo_avatar/avatar_group/create",
  {
    method: "POST",
    headers: {
      "X-Api-Key": process.env.HEYGEN_API_KEY!,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      image_key: selectedImageKey,
      name: "Product Demo Host",
      generation_id: generationId,
    }),
  }
);

const { data } = await createResponse.json();
const talkingPhotoId = data.id;

// 4. 状态变为 "completed" 后生成视频
const videoId = await generateVideo({
  video_inputs: [{
    character: {
      type: "talking_photo",
      talking_photo_id: talkingPhotoId,
    },
    voice: {
      type: "text",
      input_text: "Welcome to our product demo!",
      voice_id: "1bd001e7e50f421d891986aad5158bc8",
    },
  }],
  dimension: { width: 1920, height: 1080 },
});

生成前检查清单(Pre-Generation Checklist)

调用 AI 生成接口前,确保 8 个字段全部有值:

#字段追问话术 / 建议默认值
1name这个头像叫什么名字?
2ageYoung Adult / Early Middle Age / Late Middle Age / Senior?
3genderWoman / Man?
4ethnicity选哪个族裔?(见上方枚举值)
5orientationhorizontal(横版)/ vertical(竖版)/ square?
6posehalf_body(推荐)/ close_up / full_body?
7styleRealistic(推荐)/ Cinematic / 其他?
8appearance描述服装、表情、灯光、背景

若用户只给出模糊需求(如「生成一个看起来专业的男性」),应追问缺失字段或采用合理默认值(例如 "Early Middle Age"、"Realistic" 风格、"half_body" 姿态、"horizontal" 方向)。

Appearance 提示词写法

appearance 是文本提示词,越具体越好:

好的示例:

  • "Professional woman with shoulder-length brown hair, wearing a light blue button-down shirt, warm friendly smile, soft studio lighting, clean white background"
  • "Young man with short black hair, casual tech startup style, wearing a dark hoodie, confident expression, modern office background with plants"

避免:

  • 含糊描述,如 "a nice person"
  • 相互冲突的属性
  • 指定具体真实人物

头像组的管理接口

列出已有 Talking Photos

查询账号下全部照片数字人:

Endpoint: GET https://api.heygen.com/v1/talking_photo.list

curl "https://api.heygen.com/v1/talking_photo.list" \
  -H "X-Api-Key: $HEYGEN_API_KEY"

响应:

{
  "code": 100,
  "data": [
    {
      "id": "ef0ed70f72c6497793e5e36e434d2aea",
      "image_url": "https://files2.heygen.ai/talking_photo/.../image.WEBP",
      "circle_image": ""
    }
  ]
}

列表中每个 id 都可以直接作为视频生成时的 talking_photo_id。

向已有头像组追加照片

Endpoint: POST https://api.heygen.com/v2/photo_avatar/avatar_group/add

async function addPhotosToGroup(
  groupId: string,
  imageKeys: string[],
  name: string
): Promise<void> {
  const response = await fetch(
    "https://api.heygen.com/v2/photo_avatar/avatar_group/add",
    {
      method: "POST",
      headers: {
        "X-Api-Key": process.env.HEYGEN_API_KEY!,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
        group_id: groupId,
        image_keys: imageKeys,
        name,
      }),
    }
  );

  const json = await response.json();
  if (json.error) {
    throw new Error(json.error);
  }
}

训练头像组

对头像组进行训练可提升动画质量:

Endpoint: POST https://api.heygen.com/v2/photo_avatar/train

curl -X POST "https://api.heygen.com/v2/photo_avatar/train" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"group_id": "045c260bc0364727b2cbe50442c3a5bf"}'

查询训练状态:

Endpoint: GET https://api.heygen.com/v2/photo_avatar/train/status/{group_id}

查询与删除

查询详情: GET https://api.heygen.com/v2/photo_avatar/{id}

async function getPhotoAvatar(id: string): Promise<PhotoAvatarResponse> {
  const response = await fetch(
    `https://api.heygen.com/v2/photo_avatar/${id}`,
    { headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! } }
  );
  return response.json();
}

删除单个照片头像: DELETE https://api.heygen.com/v2/photo_avatar/{id}

async function deletePhotoAvatar(id: string): Promise<void> {
  const response = await fetch(
    `https://api.heygen.com/v2/photo_avatar/${id}`,
    {
      method: "DELETE",
      headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! },
    }
  );

  if (!response.ok) {
    throw new Error("Failed to delete photo avatar");
  }
}

删除整个头像组: DELETE https://api.heygen.com/v2/photo_avatar_group/{group_id}

async function deletePhotoAvatarGroup(groupId: string): Promise<void> {
  const response = await fetch(
    `https://api.heygen.com/v2/photo_avatar_group/${groupId}`,
    {
      method: "DELETE",
      headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! },
    }
  );

  if (!response.ok) {
    throw new Error("Failed to delete photo avatar group");
  }
}

API 端点速查表

端点方法说明
upload.heygen.com/v1/assetPOST上传图片(返回 image_key)
/v2/photo_avatar/avatar_group/createPOST由 image_key 创建照片头像组
/v2/photo_avatar/avatar_group/addPOST向已有组追加照片
/v2/photo_avatar/trainPOST训练头像组
/v2/photo_avatar/train/status/{group_id}GET查询训练状态
/v2/photo_avatar/{id}GET查询照片头像详情/状态
/v2/photo_avatar/{id}DELETE删除照片头像
/v2/photo_avatar_group/{id}DELETE删除头像组
/v2/photo_avatar/photo/generatePOST文本生成 AI 照片
/v2/photo_avatar/generation/{id}GET查询 AI 生成状态
/v2/video/av4/generatePOSTAvatar IV:由 image_key 直出视频
/v1/talking_photo.listGET列出账号下全部 talking photos
/v2/video/generatePOST用 talking_photo_id 生成视频

(除首行为独立域名外,其余路径均基于 https://api.heygen.com。)

照片要求、最佳实践与限制

技术要求

项要求
格式JPEG、PNG
分辨率最低 512x512px
文件大小10MB 以下
面部可见度清晰、正面

质量指引

  1. 光线——面部光照均匀、自然
  2. 表情——中性或轻微微笑
  3. 背景——简洁、无杂乱元素
  4. 面部位置——居中、不被裁切
  5. 清晰度——锐利、对焦准确
  6. 角度——正面或轻微侧角

最佳实践

  1. 使用高质量照片——输入质量直接决定输出质量
  2. 优先正脸肖像——最适合动画化
  3. 保持中性表情——动画效果更自然
  4. 追求最佳画质时用 Avatar IV——最新一代技术
  5. 训练头像组——可提升动画质量
  6. 复用 talking_photo_id——头像创建一次后,在多个视频中复用同一 ID

已知限制

  • 照片质量对输出影响显著
  • 侧面大角度照片支持有限
  • 全身照可能无法正常动画化
  • 部分表情可能显得不自然
  • 处理时长随复杂度波动

小结

OpenMontage 的 photo-avatars 参考文档给出的照片数字人能力可以归纳为三条路径:一是「上传照片 → Avatar Group → 轮询 → talking_photo_id 入 video_inputs」的标准四步链路,适合需要复用与训练的正式生产;二是 Avatar IV(/v2/video/av4/generate)的一步直出,牺牲组级复用换取最新画质与更短链路;三是 AI 照片生成(8 字段全必填的 photo/generate)再回流到头像组链路,解决「没有合适肖像」的问题。三条路径共享同一套 HEYGEN_API_KEY 鉴权与 image_key 资产模型,配合 assets.md 的上传规范与 video-generation.md 的场景化请求体,即可在 OpenMontage 的头像视频流水线中完成从静态照片到成片的全链路生产。

【免费下载链接】OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio. 【免费下载链接】OpenMontage 项目地址: https://gitcode.com/GitHub_Trending/op/OpenMontage

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐