即梦 AI(字节跳动)
1. 介绍
1.1. 平台概况
即梦 AI 由字节跳动旗下剪映团队孵化,百度百科标注其运营主体为深圳市脸萌科技有限公司。它是本组平台中"模型—产品—云 API"三层打通最完整的国内样本:底层为字节 Seed 团队的 Seedream(图像)与 Seedance(视频)模型,中层为剪映 / CapCut / 豆包 / 小云雀等生态产品,对外经火山引擎提供企业 API。
在 AI Harness 六层能力模型中,即梦的定位可以概括为:L1 上下文工程投入极重(多图参考 + 视觉信号 + 联网检索 + 思维链),L6 治理层已被监管实测并暴露短板。它是本组唯一被公开记载因未落实生成合成内容标识要求而被查处的平台,这一事实使其成为研究"L6 治理层如何从事后补救转为内建机制"的关键样本。
| 项 | 内容 | 置信度 |
|---|---|---|
| 开发商 | 字节跳动旗下剪映团队;运营主体标注为深圳市脸萌科技有限公司 | 中高 |
| 上线时间 | 2024-03 底以"剪映 Dreamina"内测;2024-05-09 更名"即梦"并全量上线 AI 作图与 AI 视频;2024-07-31 安卓版上架 | 中高 |
| 模型底座 | 图像端 Seedream 系列,视频端 Seedance 系列(同属字节 Seed 团队) | 高 |
| 开放形态 | Web(jimeng.jianying.com)+ iOS / Android App + 小程序;企业侧经火山引擎提供 API | 高 |
| C 端会员价 | 存在两套冲突口径, | 低 |
| B 端 API 价 | 0.2 元/图(3.0 系列)、0.22 元/张(4.0 / 4.6)、0.2 元/次(inpainting)、0.4 元/次(智能超清) | 极高(火山引擎官方) |
1.2. 发展节点
| 时间 | 事件 | 置信度 |
|---|---|---|
| 2024-03 底 | 以"剪映 Dreamina"内测 | 中高 |
| 2024-05-09 | 更名"即梦",全量上线 AI 作图与 AI 视频 | 中高 |
| 2024-07-31 | 安卓版上架 | 中高 |
| 2024-07 | 作为首席 AI 技术支持方参与《三星堆:未来启示录》 | 高 |
| 2026-02 | 接入 Seedance 2.0(图 / 视 / 音 / 文四模态混合输入、15 秒、音画同步、多镜头叙事),同步上线 Seedream 5.0 Lite | 中高 |
| 2026-02-10 | Seedream 5.0 在剪映、CapCut、小云雀正式上线,即梦 AI 灰度测试 | 中高 |
| 2026-04-09 | 推出协作型 AI 叙事创作工具"小章鱼 Octo" | 中高 |
| 2026-04-28 | 即梦 AI 网站因未有效落实人工智能生成合成内容标识规定要求,被网信部门依法查处 | 中高(建议以官方通报复核) |
| 2026-08-05 | 接入 Seedance 2.5 | 中高 |
| 2026-08-08 | 成为第 38 届大众电影百花奖 AIGC 推优单元独家 AIGC 技术合作伙伴 | 中高 |
| 2026-08-26 | 推出影视内容厂牌"即梦片场" | 中 |
1.3. 定价体系
1.3.1. C 端会员(两套冲突口径)
| 口径 | 免费 | 档位一 | 档位二 | 档位三 |
|---|---|---|---|---|
| 口径 A | ¥0 | 基础会员 ¥69/月(1,080 积分/月) | 标准会员 ¥299/月 | 高级会员 ¥499/月(15,000 积分/月) |
| 口径 B | 免费 | ¥79 | ¥239 | ¥649 |
两套口径均来自第三方微信图文截图与自媒体横评,均非官方渠道,本报告并列呈现并统一标注 。选型时应以 App 内实时价格为准。
按口径 A 换算,1,080 积分约对应 4,320 张图(图片约 4 积分/张量级);视频消耗显著更高。该换算为推算值,。
1.3.2. B 端火山引擎 API(官方,最近更新 2026-03-31)
| 能力 | 计费方式 | 价格(元) |
|---|---|---|
| 即梦 AI-文生图 3.0 / 3.1、图生图 3.0 智能参考、AI 营销商品图 3.0 | 按调用次数(单次出图 1 张) | 0.2 元/图 |
| 即梦 AI-图片生成 4.0、素材提取(商品提取 / POD 按需定制) | 按生成张数(单次有概率出多张) | 0.22 元/张 |
| 即梦 AI-图片生成 4.6 | 按生成张数 | 0.22 元/张 |
| 即梦 AI-交互编辑 inpainting | 按调用次数 | 0.2 元/次 |
| 即梦 AI-智能超清 | 按调用次数 | 0.4 元/次 |
| 并发扩充 | 按并发数 | 500 元/日/并发;10,000 元/月/并发 |
| 免费额度 | 体验 200 次,并发 1 | — |
| 欠费策略 | 欠费后 2 小时内可用;24 小时未补缴释放资源 | — |
这组单价是本组中置信度最高的定价数据之一,可直接用于成本测算。对比同组:通义万相 wan2.7-image 为 0.2 元/张,Qwen 系图像编辑 0.14 元/张,FLUX.2 [klein] 4B 为 $0.014 + $0.001/MP。
2. 名词解释
2.1. AI 图像通用术语
| 术语 | 英文 / 缩写 | 释义 | 即梦 / Seedream 的对应实现 |
|---|---|---|---|
| 文生图 | Text-to-Image(T2I) | 仅由文本提示词生成图像 | 支持,中文语义理解处于国内第一梯队 |
| 图生图 | Image-to-Image(I2I) | 以一张或多张图像为条件生成新图像 | 智能参考(图生图 3.0);Seedream 5.0 Edit 支持 1~10 张输入 |
| 局部重绘 | Inpainting | 对指定区域重新生成,区域外保持不变 | 智能画布内局部重绘;API 侧"交互编辑 inpainting" 0.2 元/次 |
| 外扩 | Outpainting | 在画布外扩区域继续生成 | 智能画布"一键扩图" |
| 可控生成 | ControlNet | 以 Canny、Depth、Pose、Mask 等视觉信号控制生成结构 | 原生内建 Canny / Depth / Mask,无需外挂 ControlNet 模型 |
| 参考图 | Reference Image | 作为身份、风格、结构约束输入的图像 | 多图参考,官方口径"十几张",第三方 API 口径 1~10 张 |
| 随机种子 | Seed | 固定后可在同参数下复现同一张图 | 平台层未公开种子锁定与快照端点的官方说明,[待填写] |
| 引导强度 | CFG | 提示词对生成结果的约束强度 | 平台层未公开 CFG 暴露方式,[待填写] |
| 低秩适配 | LoRA | 小参数量微调模块,用于固化人物、风格、服装资产 | C 端未提供用户侧 LoRA 训练与加载 |
| 图像提示适配 | IP-Adapter | 用图像编码器特征注入注意力,实现"以图为提示词" | 未公开使用;功能上由"灵活参考"能力承担 |
| 零样本身份注入 | InstantID | 单张参考图、无需微调即可迁移身份 | 未公开使用;同类能力由 PuLID(同为字节出品)在开源生态承担 |
2.2. 即梦与 Seedream 特有术语
| 术语 | 英文 / 缩写 | 释义 |
|---|---|---|
| 智能画布 | Smart Canvas | 即梦的一站式画布:集成 AI 拼图、局部重绘、一键扩图、图像消除、抠图、多图层编辑,在同一画布内保持风格统一 |
| 积分 | Credits | 即梦 C 端计量单位;图片生成约 4 积分/张量级,视频消耗显著更高 |
| 智能参考 | Smart Reference | 火山引擎侧的图像条件生成能力(图生图 3.0) |
| 素材提取 | Asset Extraction | 从商品图中提取商品主体,含商品提取与 POD 按需定制两个子能力 |
| 数字人分身认证 | Digital Human Identity Verification | 2026-02 起引入的机制:限制真人素材使用,并把"人"固化为可复用资产 |
| 统一生成与编辑架构 | Unified Generation-Editing | Seedream 5.0 把文生图与 SeedEdit 图像编辑整合进同一架构联合训练 |
| 视觉信号控制 | Visual Signal Control | 原生集成 Canny / Depth / Mask,用户可用草图、涂鸦、辅助线引导生成 |
| 上下文推理生成 | In-Context Reasoning | 理解物理与时间约束、3D 空间与复杂语境,在拼图、填字、漫画续画中保持风格一致 |
| 控制笔刷 | Control Brush | Seedream 5.0 新增的精准选择与调整的图像编辑方式 |
| 小章鱼 Octo | Octo | 2026-04-09 推出的协作型 AI 叙事创作工具 |
| 即梦片场 | Jimeng Studio | 2026-08-26 推出的影视内容厂牌 |
2.3. 换装与换脸方向通用术语
说明:即梦 / Seedream 官方在"八大核心能力"中明确把虚拟试穿(virtual try-on)列为"多图参考"的典型场景,但未提供独立的换装产品线与遮罩式换脸能力。
| 术语 | 英文 / 缩写 | 释义 |
|---|---|---|
| 虚拟试穿 | VTON(Virtual Try-On) | 将目标服装"穿"到指定人物图像上并生成视觉可信结果;Seedream 官方点名为多图参考的典型场景 |
| 服装形变 | Garment Warping | 先用 TPS(薄板样条)等几何变换把平铺服装对齐到人体姿态,再送入生成 |
| 服装掩码 | Cloth Mask | 人体解析得到的上衣、下装、外套区域二值图,用于限定重绘范围 |
| 试穿扩散 | Try-on Diffusion | 以扩散模型端到端完成服装与人体融合,不依赖显式形变 |
| 换脸 | Face Swap | 把 A 的脸替换到 B 的面部位置 |
| 人脸重演 | Face Reenactment | 保留身份、迁移表情、口型与头部姿态 |
| 身份保持 | Identity Preservation | 生成结果在多大程度上仍"是那个人" |
3. 功能说明
3.1. 生成与编辑能力
| 能力 | 支持 | 说明 |
|---|---|---|
| 文生图 | 支持 | 中文语义理解为国内第一梯队,无需英文提示词技巧 |
| 图生图 | 支持 | 智能参考;Seedream 5.0 Edit 支持 1~10 张输入 |
| 局部重绘 | 支持 | 智能画布 + API inpainting |
| 扩图 / 消除 / 抠图 | 支持 | 智能画布内建 |
| 多图层编辑 | 支持 | 智能画布内保持风格统一 |
| 视觉信号控制 | 支持 | 原生 Canny / Depth / Mask |
| 联网检索生图 | 支持 | Seedream 5.0 首次支持实时联网检索(RAG)+ CoT 思维链推理 |
| 高级文字渲染 | 支持 | 公式、表格、化学结构、统计图表 |
| 多图输出 | 支持 | 一次操作生成多张图,带全局规划与上下文一致性(分镜、漫画、IP 贴纸包) |
| 4K 输出 | 支持 | 分辨率从 2K 扩展到 4K;自适应长宽比 |
| 视频生成 | 支持 | 文生视频 / 图生视频 / 首尾帧;Seedance 2.0 支持四模态混合输入、15 秒、音画同步、多镜头叙事 |
3.2. 智能画布
智能画布是即梦在 L1 上下文工程上最重要的产品化表达:它把"多张参考图 + 图层 + 局部涂抹区域"组织为一个可视化的上下文工作区。与 Midjourney 的"参数字符串"和 FLUX.2 的"JSON 字段"不同,即梦选择了空间化(画布)的上下文组织方式——用户在画布上摆放什么,模型就看到什么。
优点:直观、零学习成本;缺点:上下文难以序列化与版本化,无法像 ComfyUI 的 JSON 图那样进入 Git 做 diff 与回归。
3.3. 多模态与叙事能力
- 多模态输入:Seedance 2.0 支持图、视、音、文四模态混合输入。
- 叙事编排:故事分镜 + 小章鱼 Octo(协作型叙事创作工具)+ 视频首帧 / 尾帧约束。
- 影视化延伸:2026-08-26 推出"即梦片场"厂牌。
3.4. 商用与素材治理
- 商用授权:会员档位含"生成作品去除品牌水印"权益(截图来源,中置信)。
- 真人素材限制:2026-02 版本起限制真人素材使用,并引入数字人分身认证机制(百度百科口径,具体流程 [待填写])。
4. 平台架构
图 4-1|即梦 AI 平台架构:从 Seed 模型底座到火山引擎云 API
数据来源:基于本文分析绘制的示意图。
4.1. 模型底座
| 模型 | 定位 | 关键特性 |
|---|---|---|
| Seedream 系列 | 图像生成与编辑 | 统一 DiT + 新型高压缩 VAE;5.0 起生成与编辑联合训练 |
| Seedance 系列 | 视频生成 | 2.0 支持四模态混合输入、15 秒、音画同步、多镜头叙事 |
| SeedVLM | 多模态理解 | 微调后用于扩展输入提示词,借助 VLM 的世界知识补全文生图上下文 |
Seedream 5.0 八大核心能力(官方技术页):精准编辑、灵活参考、视觉信号控制、上下文推理生成、多图参考、多图输出、高级文字渲染、自适应比例与 4K。
4.2. 分发与生态
- C 端:Web(jimeng.jianying.com)+ iOS / Android + 小程序,直连剪映生态。
- 生态内嵌:剪映、CapCut、小云雀、豆包。
- B 端:火山引擎 API,采用异步任务模式 + 并发购买(500 元/日/并发,10,000 元/月/并发)。
4.3. 推理优化
官方技术页口径:Seedream 系列采用对抗蒸馏稳定少步推理 + 4/8 bit 混合量化离线平滑 + 投机解码降低延迟;官方称 DiT 图像生成比 Seedream 3.0 快 10 倍以上(厂商自述,未见第三方复现)。
5. Harness 设计
5.1. 六层能力总览
| 层 | 名称 | 即梦 / Seedream 的实现 | 成熟度 | 证据强度 |
|---|---|---|---|---|
| L1 | 上下文工程 | 智能画布(多参考图 + 图层 + 涂抹区)+ 视觉信号 + 联网检索(RAG)+ CoT 推理链 | 强 | 中高 |
| L2 | 工具与执行 | 生成 / 重绘 / 扩图 / 消除 / 抠图 / 超分 / 配音 / 视频,以画布按钮暴露;无开放工具注册 | 中 | 高 |
| L3 | 编排与控制 | 多图层 + 故事分镜 + 小章鱼 Octo 构成轻量编排;视频侧首帧 / 尾帧约束 | 中 | 中高 |
| L4 | 记忆与状态 | 云端素材库 / 创作历史 / 模板库;2026-02 起数字人分身认证 | 中 | 中(需复核) |
| L5 | 评估与观测 | 官方口径"综合评测领先、文生视频与图生视频全球 Elo 第一";未公开具体榜单分数与 Eval Set | 弱 | 低—中 |
| L6 | 治理与安全 | 2026-02 起限制真人素材 + 数字人分身认证;2026-04-28 因未落实《标识办法》被查处 | 弱→整改中 | 中高 |
5.2. L1 上下文工程层
即梦在本组平台中属于 L1 投入最重的阵营,其上下文由五类信号统一编排:
- 文本提示词:中文语义理解为国内第一梯队,无需英文提示词技巧。
- 多图参考:官方口径"十几张"(第三方 API 口径 1~10 张),可提取人物特征、场景风格、物体结构做有机融合。
- 结构信号:原生内建 Canny / Depth / Mask,无需外挂 ControlNet 模型——工程上消除了工具依赖,这是与 Midjourney 的关键差异。
- 外部知识:Seedream 5.0 首次支持实时联网检索(RAG),把外部事实纳入生成上下文。
- 推理链:搭载 CoT 思维链推理,可多步逻辑推理与联网知识整合。
官方明确把虚拟试穿列为多图参考的典型场景——即"保持服装尺度与物理连贯、同时融合人物特征"这类任务,被归入上下文工程而非独立产品能力。
5.3. L2 工具与执行层
- 工具集:生成、重绘、扩图、消除、抠图、超分、配音、视频。
- 暴露形态:画布按钮 + 火山引擎 API(异步任务 + 并发扩充)。
- 无开放工具注册机制:第三方不能注册新工具,不支持 Function Calling 或 MCP。
- 成本可预测性较好:B 端按调用次数 / 生成张数明码标价,可精确测算;C 端按积分计价但存在口径冲突。
5.4. L3 编排与控制层
即梦的编排是轻量编排,介于 Midjourney 的"无编排"与美图设计室 Agent Teams 的"多智能体编排"之间:
| 编排手段 | 能力 | 边界 |
|---|---|---|
| 多图层 | 同一画布内分层编辑并保持风格统一 | 手工操作,不可序列化 |
| 故事分镜 | 分镜级叙事组织 | 面向内容创作,非通用工作流 |
| 小章鱼 Octo | 协作型 AI 叙事创作工具 | 2026-04-09 推出,定位叙事协作 |
| 首帧 / 尾帧约束 | 视频生成的端点控制 | 仅限视频 |
| 模型内生成编辑同构 | 生成→编辑→再生成为同一会话内连续操作,无需切换模型 | 仅限模型层,未上升为流程层 |
缺口:无 DAG、无子智能体派发、无中断与恢复、无工作流工件版本控制。
5.5. L4 记忆与状态层
- 云端素材库 / 创作历史 / 模板库:提供基本的状态持久化。
- 数字人分身认证(2026-02):把"人"固化为可复用资产——这是本组值得注意的 L4 设计,与妙鸭相机的"数字分身"、可灵的"主体创建(Element)"属于同一思路,但即梦将其与素材治理(限制真人素材使用)绑定,是本组中唯一把 L4 与 L6 显式耦合的设计。
- 缺口:未见工程化 Checkpoint、工件版本管理、模型版本锁定端点。智能画布的上下文是空间化的,难以序列化与版本化。
5.6. L5 评估与观测层
- 官方口径称"综合评测领先、文生视频与图生视频全球 Elo 第一";自媒体引述 Elo 1269 / 1351。该数据为特定赛事时点值,应标注为厂商自述或第三方榜单时点值,不建议作为稳定事实引用。
- 未公开官方 Eval Set、Golden Dataset 或回归集。
- 观测口径:C 端为积分消耗,B 端为调用次数与并发数。未见生成前成本预估工具(对比 Leonardo.ai 的 Pricing Calculator 端点)。
5.7. L6 治理与安全层
这是即梦在本组中最具样本价值的一层,因为它已经被监管实测:
事实:2026-04-28,即梦 AI 网站因未有效落实人工智能生成合成内容标识规定要求,被网信部门依法查处。(来源:百度百科"即梦 AI"词条,中高置信,建议以官方通报复核。)
对照《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) 的要求:
| 条款 | 要求 | 即梦侧公开信息 |
|---|---|---|
| 第四条 | 提供下载、复制、导出功能时,应当确保文件中含有满足要求的显式标识 | 未见官方实现说明,[待填写] |
| 第五条 | 应当在文件元数据中添加隐式标识(属性信息、服务提供者名称或编码、内容编号);鼓励添加数字水印 | 未见官方实现说明,[待填写] |
| 第六条 | 传播平台应当核验元数据隐式标识并分三档处理 | 不适用(即梦为服务提供者而非传播平台) |
| 第十条 | 不得恶意删除、篡改、伪造、隐匿标识 | 未见官方说明 |
即梦在 2026-02 已主动限制真人素材使用并引入数字人分身认证,说明其治理意识并不落后;但"标识"这一最基础、最刚性的要求仍出现落实缺口。这印证了本组的核心判断:L6 不能靠政策声明实现,必须靠机制与工程管线实现——标识必须在下载 / 导出链路中自动注入,而不是依赖运营配置。
5.8. 成熟度判断
即梦是"强模型 + 中等 Harness + 治理层已被实测并整改"的形态。其 L1 投入达到本组第一梯队(与 Seedream 模型能力共享),L3 有轻量编排但无工作流工件,L4 有素材库但无版本管理,L5 未工程化,L6 存在已确认的历史缺口。它证明了:模型能力强、生态位置好,都不能替代标识链路的工程实现。
6. 实际案例
6.1. 《三星堆:未来启示录》
- 时间:2024-07。
- 内容:中国首部 AIGC 生成式连续性叙事科幻短剧集在抖音上线,即梦 AI 为首席 AI 技术支持方。
- 技术产出:合作中改进了视频生成能力,包括 24/30/60 fps 补帧、二倍超分、镜头水平与上下移动及方向幅度控制。
- 性质:真实合作案例,但未披露量化效果数据(播放量、转化率等)。
- 置信度:高(百度百科"即梦 AI"词条)。
6.2. 第 38 届大众电影百花奖 AIGC 推优单元
- 时间:2026-08-08。
- 内容:即梦 AI 作为独家 AIGC 技术合作伙伴参与配套活动。
- 性质:品牌合作,无量化效果数据。置信度中高。
6.3. 即梦片场与影视内容厂牌
- 时间:2026-08-26。
- 内容:推出影视内容厂牌"即梦片场",向影视工业化延伸。
- 性质:产品动作,非客户案例。置信度中。
6.4. 监管查处案例(反面)
- 时间:2026-04-28。
- 内容:即梦 AI 网站因未有效落实人工智能生成合成内容标识规定要求被网信部门依法查处。
- 性质:本组唯一可写入文档的真实监管案例,是本组论证"L6 治理层必须以工程机制落地"的直接证据。置信度中高,建议以官方通报复核。
7. 总结
7.1. 优势
- 中文语义理解国内第一梯队,无需英文提示词技巧,团队上手成本最低。
- 视觉信号原生内建(Canny / Depth / Mask),无需外挂 ControlNet,工程依赖更少。
- B 端定价高度透明:火山引擎官方明码标价(0.2 元/图、0.22 元/张、0.2 元/次 inpainting),可直接做成本测算,是本组置信度最高的定价数据之一。
- 生态内嵌深度高:剪映 / CapCut / 小云雀 / 豆包形成天然分发,模型能力可"当日上线即触达亿级用户"。
- 联网检索 + CoT:把外部事实与推理链纳入生成上下文,是 L1 的实质性扩展。
7.2. 局限与适用边界
| 局限 | 影响 |
|---|---|
| C 端定价两套冲突口径 | 采购测算不可靠,须以 App 内实时价为准 |
| 无工作流工件与版本管理 | 智能画布上下文难以序列化,回归验证不可行 |
| 无模型版本锁定端点 | 模型升级后同参数输出漂移 |
| L5 未工程化 | 质量改进依赖榜单时点值与主观判断 |
| L6 已被监管实测并查处 | 企业采购须核实标识链路整改完成情况 |
| 无开放工具注册 | 不能嵌入既有 Agent 工具链 |
适用边界:适合中文内容创作团队、短视频与短剧生产、电商营销图批量生成(走 API);不适合需要强合规审计、可复现回归、跨系统编排的企业级流水线,除非自建标识与版本管理层。
7.3. 选型建议
- 按 API 而非按会员采购:火山引擎定价为官方口径、可测算;C 端会员价两套冲突,不应用于预算编制。
- 自建标识链路:在调用即梦 API 生成后、分发前,由业务侧统一注入显式标识与元数据隐式标识,不依赖平台默认行为——这是从即梦 2026-04-28 查处事件中可直接得出的工程结论。
- 补 L4/L5:用外部系统记录"提示词 + 参考图哈希 + 模型版本 + 参数"四元组,作为最小可复现单元与举证材料。
- 涉真人素材:务必先完成数字人分身认证流程(具体流程 [待填写]),并留存授权证据。
7.4. 合规提示
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号),2025-09-01 施行:第四条显式标识义务、第五条隐式标识义务、第十条禁止恶意删除篡改伪造隐匿标识。
- 《中华人民共和国民法典》第一千零一十九条:不得以利用信息技术手段伪造等方式侵害他人肖像权;第一千零二十条的合理使用情形不包含商业性换脸。
- 《互联网信息服务深度合成管理规定》第十七条:人脸替换、人脸生成等深度合成服务,可能造成公众混淆误认的,应当进行显著标识。
- 北京互联网法院 2026-03 生效判决:举证责任转移——被告主张"AI 偶然撞脸"的须复现创作过程,无法复现则承担举证不能的不利后果。使用即梦等无版本记录工具时,应主动保存生成参数与参考图,以备举证。
信息缺口声明
- C 端会员定价:两套冲突口径(¥69/¥299/¥499 与 ¥79/¥239/¥649),均为第三方来源。以 App 内实时价为准
- 数字人分身认证机制的具体流程:仅百科提及,无细则。[待填写]
- 2026-04-28 查处事件的官方通报原文:现有来源为百度百科词条,。
- 种子锁定、CFG 暴露方式、模型版本锁定端点:未检索到官方说明。[待填写]
- 在中国《标识办法》下的显式与隐式标识实现细节:未见官方公开说明。[待填写]
- 官方带量化效果数据的商家案例:未检索到。《三星堆:未来启示录》与百花奖合作为真实案例但不含效果数据。[无量化数据]
- Seedream 5.0"比 3.0 快 10 倍以上":厂商自述,未见第三方复现。
- Elo 1269 / 1351:自媒体引述的赛事时点值,,不建议作为稳定事实引用。
8. 参考资料
- 火山引擎 · 即梦 AI-图像生成计费说明 — 火山引擎,最近更新 2026-03-31。https://www.volcengine.com/docs/85621/1544714
- 即梦 AI 官网 — 字节跳动。https://jimeng.jianying.com/
- Seedream 5.0 技术页(八大能力 / 统一架构) — 字节 Seed 团队。http://seedream4.org/seedream-4
- 百度百科 · 即梦 AI(含查处事件与产品沿革,二次来源) — https://baike.baidu.com/item/%E5%8D%B3%E6%A2%A6App/67386767
- 百度百科 · Seedream(含发布时间与团队变动,二次来源) — https://baike.baidu.com/item/Seedream/67390954
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — 国家网信办、工业和信息化部、公安部、国家广播电视总局,2025-03-14 发布,2025-09-01 施行。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- 《人工智能生成合成内容标识办法》解读 — 中国政府网 / 新华社,2025-03-16。https://www.gov.cn/zhengce/202503/content_7014281.htm
- 《互联网信息服务深度合成管理规定》第十六条、第十七条 — 国家网信办等,2022。(正文引用条文,无官方链接)
- 《中华人民共和国民法典》第一千零一十八条、第一千零一十九条、第一千零二十条 — 全国人民代表大会,2020。(正文引用条文,无官方链接)
- 经济参考报 ·《技术不是侵权"挡箭牌" 法院这样认定 AI"盗脸"》 — 新华社《经济参考报》,2026-04-17。http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
- 阿里云百炼 · wan2.7-image 模型文档(同价位段横向对比基准) — 阿里云。https://help.aliyun.com/zh/model-studio/wan2-7-image
- Black Forest Labs · FLUX.2 Overview(结构化提示词与固定快照端点对比基准) — BFL,2026。https://docs.bfl.ai/flux_2/flux2_overview
Jimeng AI (ByteDance)
1. Introduction
1.1. Platform Overview
Jimeng AI was incubated by ByteDance's Jianying (CapCut) team, and Baidu Baike lists its operating entity as Shenzhen Lianmeng Technology Co., Ltd. Within this platform group it is the domestic sample where the "model—product—cloud API" three-layer pipeline is the most thoroughly connected: at the base are ByteDance Seed team's Seedream (image) and Seedance (video) models; in the middle are ecosystem products such as Jianying / CapCut / Doubao / Xiaoyunque; and externally it provides enterprise APIs through Volcano Engine.
In the six-layer AI Harness capability model, Jimeng's positioning can be summarized as: heavy investment in L1 context engineering (multi-image reference + visual signals + web retrieval + chain-of-thought), while the L6 governance layer has been tested by regulators and exposed for its shortcomings. It is the only platform in this group publicly recorded as being penalized for failing to implement the requirements on marking AI-generated synthetic content — a fact that makes it a key sample for studying "how the L6 governance layer shifts from after-the-fact remediation to built-in mechanisms."
| Item | Content | Confidence |
|---|---|---|
| Developer | ByteDance's Jianying (CapCut) team; operating entity listed as Shenzhen Lianmeng Technology Co., Ltd. | Medium–high |
| Launch date | Late 2024-03 internal testing as "Jianying Dreamina"; renamed "Jimeng" on 2024-05-09 with full launch of AI image generation and AI video; Android version released 2024-07-31 | Medium–high |
| Model base | Seedream series for image, Seedance series for video (both from ByteDance Seed team) | High |
| Open form | Web (jimeng.jianying.com) + iOS / Android App + mini program; enterprise side provides API via Volcano Engine | High |
| C-side membership price | Two conflicting sets of figures | Low |
| B-side API pricing | 0.2 yuan/image (3.0 series), 0.22 yuan/image (4.0 / 4.6), 0.2 yuan/call (inpainting), 0.4 yuan/call (smart upscaling) | Very high (Volcano Engine official) |
1.2. Development Milestones
| Time | Event | Confidence |
|---|---|---|
| Late 2024-03 | Internal testing as "Jianying Dreamina" | Medium–high |
| 2024-05-09 | Renamed "Jimeng", full launch of AI image generation and AI video | Medium–high |
| 2024-07-31 | Android version released | Medium–high |
| 2024-07 | Participated in Sanxingdui: Revelations of the Future as chief AI technology supporter | High |
| 2026-02 | Integrated Seedance 2.0 (four-modality mixed input of image / video / audio / text, 15 seconds, audio-visual sync, multi-shot narrative); simultaneously launched Seedream 5.0 Lite | Medium–high |
| 2026-02-10 | Seedream 5.0 officially launched in Jianying, CapCut, and Xiaoyunque; gray-scale testing in Jimeng AI | Medium–high |
| 2026-04-09 | Launched the collaborative AI narrative creation tool "Xiaozhangyu Octo" | Medium–high |
| 2026-04-28 | The Jimeng AI website was penalized by cyberspace regulators for failing to effectively implement the requirements on marking AI-generated synthetic content | Medium–high (recommend cross-checking with official notice) |
| 2026-08-05 | Integrated Seedance 2.5 | Medium–high |
| 2026-08-08 | Became the exclusive AIGC technology partner for the AIGC recommendation unit of the 38th Hundred Flowers Awards | Medium–high |
| 2026-08-26 | Launched the film & TV content label "Jimeng Film Set" | Medium |
1.3. Pricing System
1.3.1. C-side Memberships (two conflicting sets of figures)
| Figures | Free | Tier 1 | Tier 2 | Tier 3 |
|---|---|---|---|---|
| Figures A | ¥0 | Basic membership ¥69/month (1,080 credits/month) | Standard membership ¥299/month | Premium membership ¥499/month (15,000 credits/month) |
| Figures B | Free | ¥79 | ¥239 | ¥649 |
Both sets of figures come from third-party WeChat screenshots and self-media reviews, neither from official channels; this report presents them side by side and uniformly marks them [To be verified]. For selection decisions, use the real-time price in the App.
Converted according to Figures A, 1,080 credits roughly correspond to 4,320 images (image generation is on the order of about 4 credits/image); video consumption is significantly higher. This conversion is an estimate.
1.3.2. B-side Volcano Engine API (official, last updated 2026-03-31)
| Capability | Billing method | Price (yuan) |
|---|---|---|
| Jimeng AI - text-to-image 3.0 / 3.1, image-to-image 3.0 smart reference, AI marketing product image 3.0 | Per call (1 image per single call) | 0.2 yuan/image |
| Jimeng AI - image generation 4.0, asset extraction (product extraction / POD on-demand customization) | Per generated image (a single call may yield multiple images) | 0.22 yuan/image |
| Jimeng AI - image generation 4.6 | Per generated image | 0.22 yuan/image |
| Jimeng AI - interactive editing inpainting | Per call | 0.2 yuan/call |
| Jimeng AI - smart upscaling | Per call | 0.4 yuan/call |
| Concurrency expansion | Per concurrency | 500 yuan/day/concurrency; 10,000 yuan/month/concurrency |
| Free quota | 200 trial calls, concurrency 1 | — |
| Arrears policy | Still usable within 2 hours of arrears; resources released if not paid within 24 hours | — |
This set of unit prices is among the highest-confidence pricing data in this group and can be used directly for cost estimation. Comparison within the group: Tongyi Wanxiang wan2.7-image is 0.2 yuan/image, Qwen-series image editing 0.14 yuan/image, FLUX.2 [klein] 4B is $0.014 + $0.001/MP.
2. Glossary
2.1. General AI Image Terminology
| Term | English / Abbreviation | Definition | Corresponding implementation in Jimeng / Seedream |
|---|---|---|---|
| Text-to-Image | Text-to-Image (T2I) | Generate an image solely from a text prompt | Supported; Chinese semantic understanding ranks in the domestic first tier |
| Image-to-Image | Image-to-Image (I2I) | Generate a new image conditioned on one or more images | Smart Reference (image-to-image 3.0); Seedream 5.0 Edit supports 1–10 inputs |
| Inpainting | Inpainting | Regenerate a specified region while leaving the rest unchanged | Local repainting within the Smart Canvas; on the API side "interactive editing inpainting" at 0.2 yuan/call |
| Outpainting | Outpainting | Continue generating in the region extended beyond the canvas | "One-click canvas expansion" in the Smart Canvas |
| Controllable generation | ControlNet | Use visual signals such as Canny, Depth, Pose, and Mask to control generation structure | Canny / Depth / Mask natively built in, no external ControlNet model required |
| Reference image | Reference Image | An image input constraining identity, style, or structure | Multi-image reference; official figures of "a dozen or so", third-party API figures of 1–10 |
| Random seed | Seed | When fixed, reproduces the same image under the same parameters | No official documentation of seed locking and snapshot endpoints at the platform level, [To be filled] |
| Guidance strength | CFG | The strength with which the prompt constrains the generation result | No official documentation of how CFG is exposed at the platform level, [To be filled] |
| Low-rank adaptation | LoRA | A small-parameter fine-tuning module used to fix person, style, and clothing assets | No user-side LoRA training and loading provided on the C side |
| Image prompt adapter | IP-Adapter | Inject image-encoder features into attention to achieve "image as prompt" | Not publicly documented; functionally handled by the "flexible reference" capability |
| Zero-shot identity injection | InstantID | Transfer identity from a single reference image without fine-tuning | Not publicly documented; the equivalent capability is handled in the open-source ecosystem by PuLID (also from ByteDance) |
2.2. Jimeng- and Seedream-Specific Terminology
| Term | English / Abbreviation | Definition |
|---|---|---|
| Smart Canvas | Smart Canvas | Jimeng's one-stop canvas: integrates AI collage, local repainting, one-click canvas expansion, object removal, background cutout, and multi-layer editing, keeping style consistent within a single canvas |
| Credits | Credits | Jimeng's C-side metering unit; image generation is on the order of about 4 credits/image, while video consumption is significantly higher |
| Smart Reference | Smart Reference | An image-conditioned generation capability on the Volcano Engine side (image-to-image 3.0) |
| Asset Extraction | Asset Extraction | Extracts the product subject from a product image, including two sub-capabilities: product extraction and POD on-demand customization |
| Digital Human Identity Verification | Digital Human Identity Verification | A mechanism introduced since 2026-02: restricts the use of real-person assets and consolidates the "person" into a reusable asset |
| Unified Generation-Editing | Unified Generation-Editing | Seedream 5.0 jointly trains text-to-image and SeedEdit image editing within the same architecture |
| Visual Signal Control | Visual Signal Control | Natively integrates Canny / Depth / Mask; users can guide generation with sketches, doodles, and guide lines |
| In-Context Reasoning | In-Context Reasoning | Understands physical and temporal constraints, 3D space, and complex context, keeping style consistent in collage, fill-in-the-blank, and comic continuation |
| Control Brush | Control Brush | A new precise selection-and-adjustment image editing method in Seedream 5.0 |
| Octo | Octo | A collaborative AI narrative creation tool launched on 2026-04-09 |
| Jimeng Studio | Jimeng Studio | A film & TV content label launched on 2026-08-26 |
2.3. General Terminology for Try-On and Face Swap
Note: In its "eight core capabilities", Jimeng / Seedream officially lists virtual try-on as a typical scenario of "multi-image reference", but does not provide a standalone try-on product line or masked face-swap capability.
| Term | English / Abbreviation | Definition |
|---|---|---|
| Virtual Try-On | VTON(Virtual Try-On) | "Wears" the target garment onto a specified person's image and produces a visually credible result; Seedream officially names it a typical scenario of multi-image reference |
| Garment Warping | Garment Warping | First uses geometric transforms such as TPS (thin-plate spline) to align a flattened garment to the body pose, then feeds it into generation |
| Cloth Mask | Cloth Mask | A binary map of the top, bottom, and outerwear regions obtained from human parsing, used to limit the repainting range |
| Try-on Diffusion | Try-on Diffusion | End-to-end fusion of garment and body with a diffusion model, without relying on explicit warping |
| Face Swap | Face Swap | Replaces A's face into B's facial position |
| Face Reenactment | Face Reenactment | Preserves identity while transferring expression, mouth shape, and head pose |
| Identity Preservation | Identity Preservation | The extent to which the generated result is still "that person" |
3. Feature Description
3.1. Generation and Editing Capabilities
| Capability | Support | Notes |
|---|---|---|
| Text-to-image | Supported | Chinese semantic understanding ranks in the domestic first tier; no English prompt tricks needed |
| Image-to-image | Supported | Smart Reference; Seedream 5.0 Edit supports 1–10 inputs |
| Inpainting | Supported | Smart Canvas + API inpainting |
| Canvas expansion / removal / cutout | Supported | Built into the Smart Canvas |
| Multi-layer editing | Supported | Keeps style consistent within the Smart Canvas |
| Visual signal control | Supported | Native Canny / Depth / Mask |
| Web-search image generation | Supported | Seedream 5.0 first supports real-time web retrieval (RAG) + CoT chain-of-thought reasoning |
| Advanced text rendering | Supported | Formulas, tables, chemical structures, statistical charts |
| Multi-image output | Supported | Generate multiple images in one operation, with global planning and context consistency (storyboards, comics, IP sticker packs) |
| 4K output | Supported | Resolution expanded from 2K to 4K; adaptive aspect ratio |
| Video generation | Supported | Text-to-video / image-to-video / first-and-last-frame; Seedance 2.0 supports four-modality mixed input, 15 seconds, audio-visual sync, multi-shot narrative |
3.2. Smart Canvas
The Smart Canvas is Jimeng's most important productized expression of L1 context engineering: it organizes "multiple reference images + layers + locally painted regions" into a visual context workspace. Unlike Midjourney's "parameter string" and FLUX.2's "JSON fields", Jimeng chose a spatial (canvas) way of organizing context — whatever the user places on the canvas is what the model sees.
Advantages: intuitive, zero learning cost; disadvantages: the context is hard to serialize and version, and cannot be entered into Git for diff and regression like ComfyUI's JSON graphs.
3.3. Multimodal and Narrative Capabilities
- Multimodal input: Seedance 2.0 supports mixed input of four modalities: image, video, audio, and text.
- Narrative orchestration: storyboarding + Xiaozhangyu Octo (collaborative narrative creation tool) + first-frame / last-frame constraints for video.
- Film-style extension: launched the "Jimeng Film Set" label on 2026-08-26.
3.4. Commercial and Asset Governance
- Commercial licensing: membership tiers include the "remove brand watermark from generated works" entitlement (from screenshots, medium confidence).
- Real-person asset restriction: since the 2026-02 version, the use of real-person assets has been restrictedlinked, and a digital human identity verification mechanism was introduced (per Baidu Baike; specific flow
[To be filled]).
4. Platform Architecture
图 4-1|即梦 AI 平台架构:从 Seed 模型底座到火山引擎云 API
数据来源:基于本文分析绘制的示意图。
4.1. Model Base
| Model | Positioning | Key features |
|---|---|---|
| Seedream series | Image generation and editing | Unified DiT + new high-compression VAE; joint generation-editing training since 5.0 |
| Seedance series | Video generation | 2.0 supports four-modality mixed input, 15 seconds, audio-visual sync, multi-shot narrative |
| SeedVLM | Multimodal understanding | Fine-tuned to extend input prompts, using the VLM's world knowledge to complete text-to-image context |
Seedream 5.0's eight core capabilities (official technical page): precise editing, flexible reference, visual signal control, in-context reasoning generation, multi-image reference, multi-image output, advanced text rendering, adaptive aspect ratio and 4K.
4.2. Distribution and Ecosystem
- C side: Web (jimeng.jianying.com) + iOS / Android + mini program, directly connected to the Jianying ecosystem.
- Ecosystem embedding: Jianying, CapCut, Xiaoyunque, Doubao.
- B side: Volcano Engine API, using an asynchronous task model + concurrency purchase (500 yuan/day/concurrency, 10,000 yuan/month/concurrency).
4.3. Inference Optimization
Per the official technical page: the Seedream series uses adversarial distillation for stable few-step inference + 4/8-bit mixed-quantization offline smoothing + speculative decoding to reduce latency; the vendor claims DiT image generation is more than 10x faster than Seedream 3.0 (vendor's own claim, no third-party reproduction seen).
5. Harness Design
5.1. Six-Layer Capability Overview
| Layer | Name | Implementation in Jimeng / Seedream | Maturity | Evidence strength |
|---|---|---|---|---|
| L1 | Context engineering | Smart Canvas (multi-reference images + layers + painted regions) + visual signals + web retrieval (RAG) + CoT reasoning chain | Strong | Medium–high |
| L2 | Tools and execution | Generation / repainting / canvas expansion / removal / cutout / upscaling / dubbing / video, exposed via canvas buttons; no open tool registration | Medium | High |
| L3 | Orchestration and control | Multi-layer + storyboard + Xiaozhangyu Octo form lightweight orchestration; first-frame / last-frame constraints on the video side | Medium | Medium–high |
| L4 | Memory and state | Cloud asset library / creation history / template library; digital human identity verification since 2026-02 | Medium | Medium (needs review) |
| L5 | Evaluation and observation | Official claims of "leading in comprehensive benchmarks, No.1 in global Elo for text-to-video and image-to-video"; specific leaderboard scores and Eval Set not disclosed | Weak | Low–medium |
| L6 | Governance and safety | Real-person asset restriction + digital human identity verification since 2026-02; penalized on 2026-04-28 for failing to implement the Marking Measures | Weak → under remediation | Medium–high |
5.2. L1 Context Engineering Layer
Within this platform group, Jimeng belongs to the camp with the heaviest L1 investment; its context is orchestrated uniformly across five categories of signals:
- Text prompts: Chinese semantic understanding ranks in the domestic first tier; no English prompt tricks needed.
- Multi-image reference: official figures of "a dozen or so" (third-party API figures of 1–10), able to extract character features, scene style, and object structure for organic fusion.
- Structural signals: Canny / Depth / Mask natively built in, no external ControlNet model needed — this eliminates tool dependency at the engineering level and is a key difference from Midjourney.
- External knowledge: Seedream 5.0 first supports real-time web retrieval (RAG), bringing external facts into the generation context.
- Reasoning chain: equipped with CoT chain-of-thought reasoning, supporting multi-step logical reasoning and integration of web knowledge.
Officially, virtual try-on is clearly listed as a typical scenario of multi-image reference — tasks such as "maintaining garment scale and physical coherence while fusing character features" are classified under context engineering rather than as a standalone product capability.
5.3. L2 Tools and Execution Layer
- Tool set: generation, repainting, canvas expansion, removal, cutout, upscaling, dubbing, video.
- Exposure form: canvas buttons + Volcano Engine API (asynchronous tasks + concurrency expansion).
- No open tool registration mechanism: third parties cannot register new tools; Function Calling or MCP is not supported.
- Good cost predictability: the B side clearly lists prices per call / per generated image and can be estimated precisely; the C side is priced by credits but has conflicting pricing sets.
5.4. L3 Orchestration and Control Layer
Jimeng's orchestration is lightweight orchestration, sitting between Midjourney's "no orchestration" and Meitu Design Studio's Agent Teams "multi-agent orchestration":
| Orchestration method | Capability | Boundary |
|---|---|---|
| Multi-layer | Layer-based editing within one canvas while keeping style consistent | Manual operation, not serializable |
| Storyboard | Storyboard-level narrative organization | For content creation, not a general-purpose workflow |
| Xiaozhangyu Octo | Collaborative AI narrative creation tool | Launched 2026-04-09, positioned for narrative collaboration |
| First-frame / last-frame constraints | Endpoint control of video generation | Video only |
| In-model generation-editing isomorphism | Generation → editing → regeneration are continuous operations in the same session without switching models | Model layer only, not elevated to the workflow layer |
Gap: no DAG, no sub-agent dispatch, no interruption and resume, no workflow artifact version control.
5.5. L4 Memory and State Layer
- Cloud asset library / creation history / template library: provides basic state persistence.
- Digital human identity verification (2026-02): consolidates the "person" into a reusable asset — a noteworthy L4 design in this group, following the same line of thinking as Miaoya Camera's "digital avatar" and Kling's "subject creation (Element)", but Jimeng ties it to asset governance (restricting real-person asset use), making it the only design in this group that explicitly couples L4 with L6.
- Gap: no engineered Checkpoint, artifact version management, or model-version locking endpoints. The Smart Canvas context is spatial, making it hard to serialize and version.
5.6. L5 Evaluation and Observation Layer
- Officially claims "leading in comprehensive benchmarks, No.1 in global Elo for text-to-video and image-to-video"; self-media cites Elo 1269 / 1351. This figure is a value at a specific event time point, should be labeled as vendor's own claim or a third-party leaderboard time-point value, and is not recommended as a stable fact to cite.
- No official Eval Set, Golden Dataset, or regression set disclosed.
- Observation basis: credits consumed on the C side, call count and concurrency on the B side. No pre-generation cost estimation tool is seen (compare Leonardo.ai's Pricing Calculator endpoint).
5.7. L6 Governance and Safety Layer
This is the layer with the greatest sample value in this group, because it has already been tested by regulators:
Fact: on 2026-04-28, the Jimeng AI website was penalized by cyberspace regulators for failing to effectively implement the requirements on marking AI-generated synthetic content. (Source: Baidu Baike's "Jimeng AI" entry; medium–high confidence; recommend cross-checking with the official notice.)
Against the requirements of the Measures for Marking AI-Generated Synthetic Content (Guoxinban Tongzi [2025] No. 2):
| Clause | Requirement | Jimeng-side public information |
|---|---|---|
| Article 4 | When providing download, copy, or export functions, shall ensure the file contains an explicit label that meets the requirements | No official implementation notes seen, [To be filled] |
| Article 5 | Shall add an implicit label in the file metadata (attribute information, service provider name or code, content number); adding a digital watermark is encouraged | No official implementation notes seen, [To be filled] |
| Article 6 | Dissemination platforms shall verify the implicit label in metadata and handle it in three tiers | Not applicable (Jimeng is a service provider, not a dissemination platform) |
| Article 10 | Must not maliciously delete, alter, forge, or conceal labels | No official notes seen |
Since 2026-02, Jimeng has proactively restricted the use of real-person assets and introduced digital human identity verification, showing that its governance awareness is not behind; yet "marking" — this most basic and most rigid requirement — still shows a gap in implementation. This confirms the group's core conclusion: L6 cannot be achieved by policy statements; it must be achieved through mechanisms and engineering pipelines — the label must be injected automatically in the download / export chain, not left to operational configuration.
5.8. Maturity Assessment
Jimeng takes the form of "strong model + medium Harness + governance layer already tested and remediated". Its L1 investment reaches the group's first tier (shared with Seedream's model capability), L3 has lightweight orchestration but no workflow artifacts, L4 has an asset library but no version management, L5 is not engineered, and L6 has a confirmed historical gap. It demonstrates that: neither strong model capability nor a good ecosystem position can replace the engineering implementation of the labeling chain.
6. Real-World Cases
6.1. Sanxingdui: Revelations of the Future
- Time: 2024-07.
- Content: China's first AIGC-generated sequential-narrative sci-fi short drama series launched on Douyin, with Jimeng AI as chief AI technology supporter.
- Technical output: the collaboration improved video-generation capabilities, including 24/30/60 fps frame interpolation, 2x upscaling, and control over horizontal and vertical camera movement and direction/magnitude.
- Nature: a real collaboration case, but no quantified performance data disclosed (view counts, conversion rates, etc.).
- Confidence: high (Baidu Baike "Jimeng AI" entry).
6.2. 38th Hundred Flowers Awards AIGC Recommendation Unit
- Time: 2026-08-08.
- Content: Jimeng AI participated in the supporting activities as exclusive AIGC technology partner.
- Nature: brand partnership, no quantified performance data. Medium–high confidence.
6.3. Jimeng Film Set and the Film & TV Content Label
- Time: 2026-08-26.
- Content: launched the film & TV content label "Jimeng Film Set", extending toward film industrialization.
- Nature: a product move, not a customer case. Medium confidence.
6.4. Regulatory Penalty Case (negative example)
- Time: 2026-04-28.
- Content: the Jimeng AI website was penalized by cyberspace regulators for failing to effectively implement the requirements on marking AI-generated synthetic content.
- Nature: the only real regulatory case in this group that can be written into the document, and direct evidence for the group's argument that "the L6 governance layer must be implemented through engineering mechanisms". Medium–high confidence; recommend cross-checking with the official notice.
7. Summary
7.1. Strengths
- Chinese semantic understanding ranks in the domestic first tier, requiring no English prompt tricks, so teams face the lowest onboarding cost.
- Visual signals natively built in (Canny / Depth / Mask), no external ControlNet needed, fewer engineering dependencies.
- Highly transparent B-side pricing: officially listed by Volcano Engine (0.2 yuan/image, 0.22 yuan/image, 0.2 yuan/call inpainting), enabling direct cost estimation — among the highest-confidence pricing data in this group.
- Deep ecosystem embedding: Jianying / CapCut / Xiaoyunque / Doubao form natural distribution, so model capabilities can be "released the same day and reach hundreds of millions of users".
- Web retrieval + CoT: brings external facts and reasoning chains into the generation context, a substantive extension of L1.
7.2. Limitations and Applicability Boundaries
| Limitation | Impact |
|---|---|
| Two conflicting C-side pricing sets | Procurement estimation unreliable; must use the real-time price in the App |
| No workflow artifacts and version management | Smart Canvas context hard to serialize; regression verification infeasible |
| No model-version locking endpoint | Same-parameter output drifts after model upgrades |
| L5 not engineered | Quality improvement relies on leaderboard time-point values and subjective judgment |
| L6 already tested and penalized by regulators | Enterprise procurement must verify that the labeling-chain remediation is complete |
| No open tool registration | Cannot be embedded into existing Agent toolchains |
Applicability boundary: suited to Chinese content-creation teams, short-video and short-drama production, and batch generation of e-commerce marketing images (via API); not suited to enterprise-grade pipelines that require strong compliance auditing, reproducible regression, and cross-system orchestration, unless a labeling and version-management layer is built in-house.
7.3. Selection Recommendations
- Procure by API rather than membership: Volcano Engine pricing is the official baseline and can be estimated; the two conflicting C-side membership price sets should not be used for budgeting.
- Build a labeling chain in-house: after calling the Jimeng API to generate, before distribution, the business side should uniformly inject the explicit label and metadata implicit label rather than relying on the platform's default behavior — an engineering conclusion drawn directly from the 2026-04-28 penalty event.
- Supplement L4/L5: use an external system to record the "prompt + reference-image hash + model version + parameters" four-tuple as the minimum reproducible unit and evidentiary material.
- Real-person assets: be sure to complete the digital human identity verification process first (specific flow
[To be filled]), and retain authorization evidence.
7.4. Compliance Notes
- The Measures for Marking AI-Generated Synthetic Content (Guoxinban Tongzi [2025] No. 2), effective 2025-09-01: Article 4 explicit-label obligation, Article 5 implicit-label obligation, Article 10 prohibition on maliciously deleting, altering, forging, or concealing labels.
- Civil Code of the People's Republic of China, Article 1019: must not infringe others' right of portraiture through forgery or other technical means; the fair-use circumstances in Article 1020 do not include commercial face-swapping.
- Provisions on the Administration of Deep Synthesis of Internet Information Services, Article 17: deep-synthesis services such as face replacement and face generation that may cause public confusion or misidentification shall carry a prominent label.
- Effective judgment of Beijing Internet Court in 2026-03: shift of burden of proof — a defendant claiming "an AI coincidentally looks like someone" must reproduce the creation process; failing to reproduce results in the adverse consequence of failed proof. When using tools such as Jimeng that keep no version records, proactively save the generation parameters and reference images for evidence purposes.
Information Gap Statement
- C-side membership pricing: two conflicting sets of figures (¥69/¥299/¥499 and ¥79/¥239/¥649), both from third-party sources.
[To be verified; use the real-time price in the App] - The specific flow of the digital human identity verification mechanism: only mentioned in Baike, no details.
[To be filled] - The original text of the official notice of the 2026-04-28 penalty event: the current source is the Baidu Baike entry.
- Seed locking, CFG exposure, and model-version locking endpoints: no official documentation found.
[To be filled] - Implementation details of explicit and implicit labels under China's Marking Measures: no official public notes seen.
[To be filled] - Official merchant cases with quantified performance data: none found. The Sanxingdui and Hundred Flowers Award collaborations are real cases but contain no performance data.
[No quantified data] - Seedream 5.0 "more than 10x faster than 3.0": vendor's own claim, no third-party reproduction seen.
- Elo 1269 / 1351: event time-point values cited by self-media, not recommended as stable facts to cite.
8. References
- Volcano Engine · Jimeng AI - image generation billing notes — Volcano Engine, last updated 2026-03-31. https://www.volcengine.com/docs/85621/1544714
- Jimeng AI official website — ByteDance. https://jimeng.jianying.com/
- Seedream 5.0 technical page (eight capabilities / unified architecture) — ByteDance Seed team. http://seedream4.org/seedream-4
- Baidu Baike · Jimeng AI (includes penalty event and product history, secondary source) — https://baike.baidu.com/item/%E5%8D%B3%E6%A2%A6App/67386767
- Baidu Baike · Seedream (includes release dates and team changes, secondary source) — https://baike.baidu.com/item/Seedream/67390954
- The Measures for Marking AI-Generated Synthetic Content (Guoxinban Tongzi [2025] No. 2) — Cyberspace Administration of China, Ministry of Industry and Information Technology, Ministry of Public Security, National Radio and Television Administration, published 2025-03-14, effective 2025-09-01. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- Interpretation of the Measures for Marking AI-Generated Synthetic Content — gov.cn / Xinhua News Agency, 2025-03-16. https://www.gov.cn/zhengce/202503/content_7014281.htm
- Provisions on the Administration of Deep Synthesis of Internet Information Services, Articles 16 and 17 — Cyberspace Administration of China et al., 2022. (Clauses cited in the text; no official link)
- Civil Code of the People's Republic of China, Articles 1018, 1019, and 1020 — National People's Congress, 2020. (Clauses cited in the text; no official link)
- Economic Information Daily · "Technology is not a 'shield' for infringement; how courts determined AI 'face theft'" — Xinhua News Agency's Economic Information Daily, 2026-04-17. http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
- Alibaba Cloud Bailian · wan2.7-image model documentation (benchmark for horizontal comparison in the same price band) — Alibaba Cloud. https://help.aliyun.com/zh/model-studio/wan2-7-image
- Black Forest Labs · FLUX.2 Overview (benchmark for structured prompts and fixed-snapshot endpoints) — BFL, 2026. https://docs.bfl.ai/flux_2/flux2_overview