豆包 / Seedance(字节视频模型)AI 漫剧平台研究
1. 介绍
Doubao-Seedance-2.0 是字节跳动 Seed 团队 / 豆包大模型团队研发的视频生成模型,以火山引擎方舟为服务载体。在本组 8 个对象中,Seedance 是治理层最强的样本:它明确禁用真人人脸作为参考素材、要求真人出镜须完成本人形象与声音校验、并对 API 采取分阶段开放策略(背景包括好莱坞制片厂与欧洲娱乐产业团体的版权抵制)。
同时,Seedance 是第一个同时接受四种输入模态的模型(文本、最多 9 张图、最多 3 段视频、最多 3 条音轨),并采用统一的多模态音视频联合生成架构。
Seedance 与即梦 AI 同属字节体系,但形态不同:即梦是面向人的创作平台(SaaS + 渠道),Seedance 是面向开发者的模型服务(MaaS)。本组把两者分开解剖,正是为了区分"平台 Harness"与"模型 Harness"。
1.1. 研发方与模型标识
| 项 | 内容 |
|---|---|
| 研发方 | 字节跳动 Seed 团队 / 豆包大模型团队 |
| 模型名 | Doubao-Seedance-2.0 |
| Model ID | doubao-seedance-2-0-260128(首推) |
| 发布时间 | 2026 年 2 月 |
| 输入类型 | 文本、图片、视频、音频 |
| 输出类型 | 视频 |
| API | /v3/contents/generations |
发布与接入时间线:
| 时间 | 事件 |
|---|---|
| 2026 年 2 月 | Seedance 2.0 发布 |
| 2026 年 2 月 10 日 | Seedance 概念带动传媒板块大涨(读客文化、荣信文化 20% 涨停) |
| 2026 年 2 月 12 日 | 豆包宣布 Seedance 2.0 正式接入豆包 App、电脑端和网页版 |
| 2026 年 4 月 2 日 | 火山引擎启动 Seedance 2.0 面向普通 API 客户开放申请;同日 AI 创新巡展·武汉站宣布 API 面向企业用户开放公测 |
| 2026 年 4 月 7 日 | 美图旗下 AI Agent 产品 RoboNeo 接入 Seedance 2.0 |
| 2026 年 4 月 12 日 | 天娱数科旗下影视级 AI 视频创编平台 CineART 成为首批接入平台之一 |
| 2026 年 4 月 14 日 | 火山引擎正式上线 Seedance 2.0 系列 API 服务 |
| 2026 年 5 月 | Seedance 2.0 参与制作的 8 部 AI 影片在戛纳展映 |
1.2. 定位
Seedance 2.0 的定位是"统一的多模态音视频联合生成架构"——不是多个单模态模型的拼接,而是统一架构下的联合生成。
关键特征:
- 原生 2K 分辨率(2048×1080 或 1080×2048)
- 音视频同步
- 相比 Seedance 1.5 Pro 生成速度提升 30%
- 自带专业运镜、多镜头叙事与文字生成能力
1.3. 定价
| 计费模式 | 价格 | 说明 |
|---|---|---|
| 包含视频输入(视频编辑) | 28 元/百万 tokens | 基于已有素材优化,算力需求更低 |
| 不含视频输入(纯文生视频) | 46 元/百万 tokens | 从零构建画面、时序与物理逻辑,算力消耗呈指数级增长 |
| 15 秒视频 token 消耗 | 308,880 tokens | 另一处写 30.888 万 |
| 单条 15 秒成本(纯生成模式核算) | 约 15 元,折合每秒 1 元 | — |
差异化计费逻辑(21 财经/不慌实验室):视频编辑基于已有素材优化,算力需求更低(28 元);纯视频生成从零构建画面、时序与物理逻辑,算力消耗呈指数级增长(46 元)。这是本组唯一公开了"按算力特性分层定价"逻辑的模型,对成本工程有直接指导意义。
行业预测参照:高盛预测 2030 年全球 AI 视频市场规模将突破 290 亿美元;2025 年近 4 成头部短剧采用 AI 生成技术。
1.4. 开放形态
| 形态 | 说明 |
|---|---|
| 火山引擎方舟 API | /v3/contents/generations;2026 年 4 月 14 日正式上线 Seedance 2.0 系列 API 服务 |
| 豆包 App / 电脑端 / 网页版 | 2026 年 2 月 12 日接入 |
| 即梦 Web 端 / App | 通过即梦平台使用,受即梦的真人素材限制约束 |
| 小云雀 | 短剧 Agent,通过即梦体系使用 |
| 第三方接入 | RoboNeo(美图)、CineART(天娱数科)等 |
API 规格:
| 参数 | 数值 |
|---|---|
| 分辨率 | 480P / 720P / 1080P / 4k |
| 时长 | 4~15s |
| 帧率 | 24fps |
| 并发(非 4K) | 共享 104 |
| 并发(4K) | 独享 1 |
| RPM(非 4K) | 共享 0.6k |
| RPM(4K) | 独享 0.015k |
| 任务类型 | 多模态生视频、视频编辑、视频延长 |
2. 名词解释
| 术语 | 英文 / 缩写 | 释义 |
|---|---|---|
| AI 漫剧 | AI Comic Drama | 介于静态漫画与真人短剧之间的内容形态,以漫画分镜加动态视听语言构成 |
| 动态漫 | Motion Comic | 以静态漫画素材为基础,通过运镜、缩放、局部动效与配音形成的轻微动态视频形态 |
| 分镜 / 分镜脚本 | Storyboard | 将文字剧本转化为画面草图,标注每个镜头的构图、动作、时长 |
| 角色一致性 | Character Consistency | 同一角色在跨镜头、跨集、跨次生成中保持五官、服装、体型、气质稳定的能力 |
| 关键帧 | Keyframe | 定义动画或运镜变化关键状态的帧(起点与终点),对应二维动画中的"原画" |
| 中间帧 / 过渡帧 | In-between / Tween | 关键帧之间通过插值算法自动生成的过渡帧 |
| 首尾帧 | First-Last Frame | 上传首帧与尾帧,由模型补全中间运动轨迹的图生视频控制法 |
| 图生视频 | Image-to-Video(I2V) | 输入一张静态图片,由模型生成数秒动画 |
| 文生视频 | Text-to-Video(T2V) | 输入文本提示词直接生成视频 |
| 口型同步 / 唇形同步 | Lip Sync | 把音频叠加到生成角色上并驱动嘴部动作匹配发音 |
| 镜头语言 | Camera Language | 通过景别、角度、运动、构图与剪辑节奏传递叙事信息的视听表达体系 |
| 多模态音视频联合生成 | Unified Audio-Visual Joint Generation | Seedance 的核心架构特征:视频与音频在同一次推理中联合生成,而非分阶段拼接 |
| 多镜头叙事 | Multi-shot Narrative | 单次生成或单次指令内自动规划并切换多个连贯镜头(全景→中景→特写) |
| 四模态输入 | Four-Modality Input | 文本、图片、视频、音频四种模态同时输入;Seedance 2.0 支持最多 9 图 + 3 视频 + 3 音频 |
| 视频编辑 | Video Editing | 在已有视频素材基础上做修改优化;Seedance 定价 28 元/百万 tokens |
| 视频延长 | Video Extension | 在已有视频基础上续接时长 |
| 数字分身 / 真人校验 | Digital Avatar Verification | 用户需录制本人形象与声音完成校验后才能制作本人 AI 形象出镜 |
| 数字分身认证 | Digital Avatar Authentication | 即梦自 2026 年 2 月引入的认证机制 |
| MaaS | Model as a Service | 模型即服务;Seedance 以火山方舟 API 形态提供 |
| 算法备案 | Algorithm Filing | 《生成式人工智能服务管理暂行办法》要求的备案义务 |
| 显式标识 / 隐式标识 | Explicit / Implicit Label | AI 生成合成内容的两类法定标识:显式为用户可感知提示;隐式嵌入文件元数据 |
| AIGC 元数据字段 | AIGC Metadata Field | 强制性国标 GB 45438—2025 规定的元数据隐式标识字段 |
| tokens | tokens | 模型计费单位;15 秒视频消耗 308,880 tokens |
| RPM | Requests Per Minute | 每分钟请求数;Seedance 非 4K 共享 0.6k,4K 独享 0.015k |
3. 功能说明
3.1. 多模态参考生视频
Seedance 2.0 支持参考图片、视频、音频生视频,是首个同时接受四种输入模态的模型:
| 模态 | 上限 |
|---|---|
| 文本 | 提示词 |
| 图片 | 最多 9 张 |
| 视频 | 最多 3 段 |
| 音频 | 最多 3 条 |
这一容量设计与海螺 H3(最多 12 个参考文件:9 图 + 3 视频 + 3 音频)完全一致,是本组两个最高容量档位之一。
差异在于:Seedance 走扩容路线(直接接受更多模态),海螺走压缩路线(H3-Context-IR 把约 100k token 压缩至约 4k token)。前者简单直接但有硬上限,后者复杂但容量弹性更大。
3.2. 视频编辑与延长
| 任务类型 | 说明 | 计费 |
|---|---|---|
| 多模态生视频 | 从文本/参考生成新视频 | 46 元/百万 tokens(不含视频输入) |
| 视频编辑 | 在已有素材基础上修改优化 | 28 元/百万 tokens(含视频输入) |
| 视频延长 | 在已有视频基础上续接时长 | 含视频输入,按 28 元档计 |
视频编辑与延长的成本优势值得工程关注:28 元 vs 46 元,差 39%。对"先生成关键镜头,再编辑/延长补完"的生产模式,这一差价是重要的成本杠杆。
3.3. 多镜头叙事与运镜
- 多镜头叙事:单次生成内自动规划并切换多个连贯镜头。
- 自带专业运镜:模型内置运镜能力,无需外部指定。
这是把 L3 编排下沉进模型的做法——与即梦同源(同为字节体系),也意味着使用者无法显式编辑镜头序列。
3.4. 首尾帧与文字生成
- 首尾帧生视频:支持首帧 + 尾帧指定,模型补全中间。
- 文字生成能力:模型可生成画面内文字,对漫剧中的标题、招牌、字幕等有直接价值。
4. 平台架构
4.1. 总体架构
图 4-1|Seedance 平台总体架构:从集成生态层到治理层
数据来源:基于本文分析绘制的示意图。
┌──────────────────────────────────────────────────────────────────┐
│ 集成生态层 豆包 App/电脑端/网页版 · 即梦 Web/App · 小云雀 │
│ RoboNeo(美图)· CineART(天娱数科)· 其他第三方 Agent │
├──────────────────────────────────────────────────────────────────┤
│ 服务层 火山引擎方舟 API /v3/contents/generations │
│ 并发:非 4K 共享 104 / 4K 独享 1 │
│ RPM:非 4K 共享 0.6k / 4K 独享 0.015k │
├──────────────────────────────────────────────────────────────────┤
│ 任务层 多模态生视频 · 视频编辑 · 视频延长 │
├──────────────────────────────────────────────────────────────────┤
│ 模型层 Doubao-Seedance-2.0(doubao-seedance-2-0-260128) │
│ 统一多模态音视频联合生成架构 │
│ 原生 2K · 24fps · 4~15s · 多镜头叙事 · 专业运镜 · 文字生成 │
├──────────────────────────────────────────────────────────────────┤
│ 治理层 真人人脸禁用 · 数字分身真人校验 · API 分阶段开放 │
│ · 成本护栏(约 1 元/秒) │
└──────────────────────────────────────────────────────────────────┘ 4.2. 模型层
Doubao-Seedance-2.0 采用统一的多模态音视频联合生成架构——不是"先生成视频再配音"的两阶段方案,而是音视频在同一次推理中联合产出。
关键规格:原生 2K 分辨率(2048×1080 或 1080×2048);API 侧分辨率选项 480P / 720P / 1080P / 4k;时长 4~15s;帧率 24fps;相比 Seedance 1.5 Pro 生成速度提升 30%。
4.3. 服务层
服务层的核心是并发与速率的差异化配置:
| 参数 | 非 4K | 4K |
|---|---|---|
| 并发 | 共享 104 | 独享 1 |
| RPM | 共享 0.6k | 独享 0.015k |
4K 的独享并发为 1、RPM 为 0.015k(即每分钟 15 次),是显著的资源约束。这意味着 4K 生成不适合批量生产,只适合少量关键镜头。
4.4. 集成生态层
Seedance 的集成生态是本组最有代表性的"模型被 Agent 集成"样本:
| 集成方 | 类型 | 时间 |
|---|---|---|
| 豆包 App / 电脑端 / 网页版 | 自有产品 | 2026-02-12 |
| 即梦 Web 端 / App | 同体系平台 | 2026-02 |
| 小云雀 | 短剧 Agent | — |
| RoboNeo(美图旗下 AI Agent 产品) | 第三方 Agent | 2026-04-07 |
| CineART(天娱数科影视级 AI 视频创编平台) | 第三方平台 | 2026-04-12 |
5. Harness 设计
5.1. L1 上下文工程层
Seedance 的 L1 是四模态统一上下文:文本 + 最多 9 图 + 最多 3 视频 + 最多 3 音频。
与前代相比,这是从"文本为主、图片为辅"到"四模态平权"的变化。对 AI 漫剧而言,这意味着一次生成可同时锚定角色(图)、场景(图)、运镜(视频)、音色(音频)与动作(文本)。
判断:与海螺并列本组 L1 最强档,但路线不同——Seedance 扩容,海螺压缩。Seedance 的优势是模态平权且实现简单;海螺的优势是容量弹性大(100k→4k 压缩)。
缺口:
- 无公开的上下文压缩机制——参考素材量超过 9+3+3 上限后,使用者需自行裁剪,且无优先级策略指导。
- 无公开的缓存复用机制——跨镜头重复引用同一角色时是否可复用上下文,未见说明,标
[待填写]。 - 四模态输入的权重分配机制未公开——模型如何平衡图、视频、音频、文本四种条件的优先级,无说明。
5.2. L2 工具与执行层
Seedance 的 L2 是标准的模型 API 形态:
| 工具 | 说明 |
|---|---|
/v3/contents/generations API | 统一生成入口 |
| 任务类型 | 多模态生视频、视频编辑、视频延长 |
| 首尾帧生视频 | 支持 |
| 文字生成 | 模型内置 |
判断:Seedance 的 L2 属"中"档。它提供了完整的模型能力 API,但没有 Agent 原生的工具层——没有 CLI、没有 Skills、没有 Function Calling 的工具描述、没有 MCP 服务器。它采取的是"被第三方 Agent 集成"的路线:RoboNeo、CineART 等平台各自实现工具封装。
这与 PixVerse 的 CLI + Skills(主动适配 Claude Code、Codex、Cursor、OpenClaw)形成对比:Seedance 等别人来接,PixVerse 主动去适配。
缺口:
- 无官方 CLI / Skills / MCP 支持,标
[待填写]。 - 工具粒度粗——只有"生成""编辑""延长"三个任务类型,缺少细粒度工具(Lip Sync、局部重绘、相机控制等)。
- 无沙箱、无并行/串行调度机制(并发由 RPM 限制间接管控)。
5.3. L3 编排与控制层
Seedance 的 L3 是多镜头叙事——模型自带专业运镜与镜头规划,可在单次生成内完成多镜头切换。
判断:与即梦同源,属"模型内隐式规划"路线。优点是零配置、上手快;缺点是不可视化、不可精确编辑、不可版本化、不可导出。
缺口清单:
- 无节点图、无分镜表导出、无时间轴编辑。
- 无公开的条件分支、批处理、循环机制。
- 无中断与恢复机制——长任务失败后需重跑。
- 多镜头叙事的镜头数上限与切换逻辑未见公开说明,标
[待填写]。
对复杂的 100 集连续生产,Seedance 的编排必须由上层 Agent 承载——这正是 CineART、小云雀这类第三方平台存在的价值。
5.4. L4 记忆与状态层
Seedance 的 L4 是本组的短板样本,也是最典型的"模型无状态"案例。
已知能力:角色特征稳定保持;高分辨率下的细节、材质、音色、视效风格、运镜的高精度还原(火山引擎控制台模型描述)。
明确缺口:
- 无资产库:Seedance 作为模型服务,不提供角色/道具/场景的持久化资产载体。对比 Vidu 主体库、PixVerse Character、白日梦角色库,这是本质缺失。
- 无会话概念:API 是无状态的,每次调用独立。跨镜头、跨集的一致性完全依赖调用方在每次请求中重新注入参考。
- 无剧情状态机。
这恰恰印证了本组核心论断:AI 漫剧的跨镜头、跨集一致性本质是跨会话状态保持问题,而模型本身不解决这个问题。Seedance 把这个责任完整地交给了调用方——任何基于 Seedance 构建 AI 漫剧生产系统的团队,必须在上层自建完整的 L4。
反过来说,这也是 Seedance 作为"纯模型层"的合理定位:它提供了强 L1(四模态)与强生成能力,把 L4 留给上层 Harness 去实现。
5.5. L5 评估与观测层
Seedance 的 L5 信号:
| 信号 | 说明 |
|---|---|
| 视频可用率 | 火山引擎描述"大幅提升",无具体数值,标 [待填写] |
| 戛纳展映 | 2026 年 5 月,Seedance 2.0 参与制作的 8 部 AI 影片在戛纳展映——外部质量验证 |
| 生成速度 | 相比 Seedance 1.5 Pro 提升 30% |
判断:Seedance 的 L5 属"强"档,但依据是外部验证(戛纳展映)而非平台内建评估。这与本组其他平台一致——目前没有任何平台提供内建的轨迹追踪、回归集或可用率面板。
可观测性缺口:
- 无官方的可用率统计数值(仅有"大幅提升"的定性描述)。
- 无生成轨迹追踪与结构化日志。
- 无成本分析面板——虽然有明确的 token 计费规则,但缺少用量可视化工具(火山方舟控制台是否提供,未见说明,标
[待填写])。
5.6. L6 治理与安全层
Seedance 的 L6 是本组最强,与即梦并列。
机制一:真人人脸禁用
在即梦 Web 端、小云雀等平台使用 Seedance 2.0 时,系统明确提示暂不支持真人人脸作为参考素材。
机制二:数字分身真人校验
在即梦 App 与豆包 App 中,用户若希望真人形象出镜,须录制本人形象与声音完成真人校验后方可制作数字分身。
这是本组唯一要求"形象 + 声音"双要素校验的机制,比单一形象校验更严格。
机制三:API 分阶段开放
2026 年 3 月底 Seedance 2.0 仍受限于方舟体验中心,API 分阶段放开。背景包括好莱坞制片厂与欧洲娱乐产业团体的版权抵制。2026 年 4 月 2 日启动普通 API 客户开放申请,4 月 14 日正式上线系列 API 服务。
这体现了一种审慎的治理姿态:在版权争议未澄清前,宁可限制开放范围也不贸然全量放开。
机制四:法规遵从
豆包/Doubao 受《生成式人工智能服务管理暂行办法》(2023 年 8 月生效)约束,须完成算法备案、维护核心价值观、对合成媒体加标识。
机制五:成本护栏
约 1 元/秒的定价本身构成成本护栏;4K 独享并发 1、RPM 0.015k 的限流是资源护栏。
缺口:
- 未检索到 Seedance 是否内置符合 GB 45438—2025 的元数据隐式标识的正面说明——虽然法规要求"对合成媒体加标识",但具体实现(显式标识位置/字高/时长、AIGC 元数据字段)未见公开技术文档,标
[待填写]。 - 4K 并发为 1、RPM 0.015k 是硬性产能约束,批量生产需提前规划配额。
- API 开放范围仍可能随版权谈判变化,存在供应不确定性。
5.7. 六层能力矩阵
| 层 | Seedance 的实现 | 成熟度 | 主要缺口 |
|---|---|---|---|
| L1 上下文工程 | 四模态统一上下文(9 图 + 3 视频 + 3 音频 + 文本) | 强 | 无压缩机制;无缓存复用;模态权重不明 |
| L2 工具与执行 | /v3/contents/generations API;生视频/编辑/延长;首尾帧 | 中 | 无 CLI/Skills/MCP;工具粒度粗 |
| L3 编排与控制 | 多镜头叙事(模型内隐式规划);专业运镜 | 强 | 不可编辑、不可导出、不可版本化 |
| L4 记忆与状态 | 角色特征稳定保持(模型侧);无资产库 | 中 | 无资产库、无会话、无状态机——全部需上层自建 |
| L5 评估与观测 | 戛纳展映外部验证;可用率"大幅提升"(无数值) | 强 | 无内建评估、无轨迹追踪 |
| L6 治理与安全 | 真人人脸禁用 + 数字分身真人校验 + API 分阶段开放 + 成本护栏 | 最强 | 隐式标识实现未公开;4K 产能受限 |
6. 实际案例
6.1. 案例一:第三方 Agent 产品接入
- 背景:AI 视频模型要进入专业创作链路,必须能被第三方 Agent 与平台集成。
- 方案:Seedance 2.0 以
/v3/contents/generationsAPI 形式输出,被多方集成——2026 年 4 月 7 日美图旗下 AI Agent 产品 RoboNeo 接入;2026 年 4 月 12 日天娱数科旗下影视级 AI 视频创编平台 CineART 成为首批接入平台之一;同体系的豆包 App/电脑端/网页版、即梦 Web/App、小云雀亦已接入。 - 效果:具体调用量与业务数据未见公开披露,标
[待填写]。 - Harness 解读:这是"模型 Harness"与"平台 Harness"分工的典型样本。Seedance 提供 L1/L2(上下文与执行原语),RoboNeo 与 CineART 在其上自建 L3/L4(编排与状态)。如果你要基于 Seedance 做 AI 漫剧,你扮演的角色就是 CineART——模型不会替你管资产与剧情状态。
6.2. 案例二:戛纳展映的 8 部 AI 影片
- 背景:AI 生成内容的专业认可度,需要权威场景的外部验证。
- 方案:2026 年 5 月,Seedance 2.0 参与制作的 8 部 AI 影片在戛纳展映。
- 效果:这是本组唯一进入国际顶级电影节展映的质量验证事件。具体影片名称、展映单元与评审反馈未见公开披露,标
[待填写]。 - Harness 解读:戛纳展映是 L5 的外部信号而非平台能力。它证明 Seedance 的输出质量达到专业可用水准,但不意味着你的项目也能自动达到该水准——中间隔着 L4(一致性)与 L3(编排)的巨大工程差距。
6.3. 案例三:API 分阶段开放与版权抵制
- 背景:AI 视频模型的训练数据与生成内容涉及大量版权争议。
- 方案:2026 年 3 月底 Seedance 2.0 仍受限于方舟体验中心,API 分阶段放开;2026 年 4 月 2 日启动面向普通 API 客户开放申请,4 月 14 日正式上线系列 API 服务。开放节奏审慎的背景是好莱坞制片厂与欧洲娱乐产业团体的版权抵制。
- 效果:分阶段开放策略使模型能力得以在争议中稳步释放;具体的开放范围与配额规则未见完整披露,标
[待填写]。 - Harness 解读:这是 L6 治理影响 L2 可用性的直接案例——治理决策会限制工具层的可得性。对依赖 Seedance 的生产系统,这意味着存在供应不确定性,应设计多模型兜底策略。
6.4. 案例四:差异化计费的成本工程
- 背景:AI 漫剧的规模化生产对成本极度敏感——行业万播收益已从约 100 元跌至 5~10 元,回本率不足 1.3%。
- 方案:Seedance 采用差异化计费——含视频输入(视频编辑)28 元/百万 tokens,不含视频输入(纯文生视频)46 元/百万 tokens。15 秒视频消耗 308,880 tokens,纯生成模式约 15 元/条,折合每秒 1 元。
- 效果:视频编辑相比纯生成便宜约 39%。对"先生成关键镜头,再编辑/延长补完"的生产模式,这是显著的成本杠杆。具体的节省幅度取决于项目结构,标
[待填写]。 - Harness 解读:这是本组最值得学习的成本工程实践。它提示了一种生产策略:把镜头分为"必须从零生成的关键帧"与"可基于已有素材编辑/延长的补完帧",前者走 46 元档,后者走 28 元档。这种分层在 L3 编排层实现,直接影响 L6 的成本护栏效果。
7. 总结
7.1. 优势
- L6 治理本组最强:真人人脸禁用 + 数字分身真人校验(形象 + 声音双要素)+ API 分阶段开放,是本组最完整、最审慎的治理组合。
- 四模态输入容量最高档之一:9 图 + 3 视频 + 3 音频 + 文本,与海螺并列。
- 统一多模态音视频联合生成架构:非分阶段拼接,是架构层面的统一。
- 原生 2K + 4K 支持:分辨率档位完整(480P / 720P / 1080P / 4k)。
- 外部质量验证最强:8 部 AI 影片戛纳展映。
- 差异化计费逻辑透明:28 元 vs 46 元的定价逻辑有公开的算力解释,便于成本工程。
- 集成生态成熟:被 RoboNeo、CineART 等专业平台集成,验证了模型的可集成性。
- 速度提升明确:相比 Seedance 1.5 Pro 生成速度提升 30%。
7.2. 局限
- L4 完全缺失:无资产库、无会话、无状态机。所有跨镜头、跨集的一致性工作都要调用方自建——这是本组最彻底的"把 L4 交给上层"的设计。
- 无 Agent 原生工具层:无 CLI、无 Skills、无 MCP,只等第三方来接。
- 4K 产能受限:独享并发 1、RPM 0.015k,无法批量生产。
- 编排不可控:多镜头叙事是模型内隐式规划,不可编辑、不可导出、不可版本化。
- 可用率无量化:仅有"大幅提升"的定性描述。
- 成本偏高:约 1 元/秒的纯生成成本,高于海螺(0.8 元/秒 2K)与 Vidu(宣称行业平均 1/3)。
- 供应不确定性:API 开放范围受版权争议影响,可能变化。
- 隐式标识实现未公开:虽受法规约束,但具体技术实现未见文档。
- 时长 4~15 秒:与同组多数平台持平,长镜头需拼接。
7.3. 适用边界
| 适合 | 不适合 |
|---|---|
| 已有自有编排/资产系统的团队(只需强模型层) | 希望平台提供完整一致性方案的团队 |
| 需要四模态复杂参考的镜头 | 需要把生成接进 AI 原生开发环境(无 CLI/Skills) |
| 需要原生音视频联合生成的内容 | 需要 4K 批量生产(并发受限) |
| 需要强合规治理背书的商业交付 | 预算极度敏感的项目(约 1 元/秒) |
| 专业影视/广告的高质量单镜头 | 需要可视化精细编排的复杂剧集 |
| 需要"编辑/延长"降本的生产模式 | 需要稳定长期供应保障的关键业务(开放范围可能变化) |
7.4. 选型建议
- 选 Seedance 的核心理由是模型能力与治理成熟度,不是完整度。若你已经有自己的编排系统与资产库,只需要一个强模型层,Seedance 是优质选择。
- 必须自建完整的 L4:这是使用 Seedance 的前提,不是可选项。你需要:分镜表(结构化 JSON)、角色卡(含参考图与文本锚点)、场景卡、剧情状态机,并在每一次 API 调用中重新注入这些状态。
- 利用差异化计费做成本工程:把镜头分为"必须从零生成的关键帧"(46 元档)与"可基于已有素材编辑/延长的补完帧"(28 元档),在编排层实现分流。
- 4K 只用于关键镜头:4K 独享并发 1、RPM 0.015k,把它留给少量高价值镜头,批量生产走 1080P。
- 设计多模型兜底:由于 API 开放范围可能随版权谈判变化,生产系统应抽象出模型层接口,预留切换到其他模型(如海螺 H3、Vidu Q3)的能力。
- 合规能力需自行验证:虽受《生成式人工智能服务管理暂行办法》约束须"对合成媒体加标识",但具体实现未见公开技术文档。国内平台分发的作品须自行校验是否符合 GB 45438—2025 的显式标识(起始画面边角、字高不低于最短边 5%、持续不少于 2 秒)与 AIGC 元数据隐式标识要求。
- 成本核算按 1 元/秒估算并预留冗余:15 秒 308,880 tokens 约 15 元/条;按项目镜头数 × 返工系数(建议不低于 1.3)计算总预算。
信息缺口声明
- 多模态输入的模态权重机制:模型如何平衡图、视频、音频、文本四种条件的优先级,未见公开说明。
- "角色特征稳定保持"的技术机制与量化指标:火山引擎控制台仅作定性描述,无一致性数值与实现说明。
- 多镜头叙事的规格:单次生成支持的最大镜头数、镜头切换逻辑、是否可指定镜头序列,未见公开说明。
- 上下文缓存复用机制:跨镜头重复引用同一角色时是否可复用上下文以降低 token 消耗,无公开结果。
- Seedance 是否内置符合 GB 45438—2025 的元数据隐式标识:法规要求加标识,但具体技术实现(AIGC 元数据字段写入)未见公开技术文档。
- 可用率的具体数值:仅有"大幅提升"的定性描述,无量化数据。
- API 开放范围与配额规则:分阶段放开的具体标准、各阶段的配额上限、未来是否进一步开放,未见完整披露。
- 戛纳展映 8 部影片的详细信息:影片名称、展映单元、评审反馈未见公开披露。
- 是否提供官方 CLI / Skills / MCP 支持:未检索到相关公开信息。
- 火山方舟控制台的成本分析能力:是否提供用量可视化、预算告警、成本分析面板,未见说明。
- Seedance 2.5 与 2.0 的能力差异:即梦已于 2026 年 8 月 5 日接入 Seedance 2.5,但 2.5 相对 2.0 的参数提升幅度未见官方量化披露。
- 视频编辑任务的具体能力边界:支持哪些编辑操作(换背景、换角色、改运镜、改时长等),未见完整说明。
8. 参考资料
- 火山引擎方舟控制台 Doubao-Seedance-2.0 系列 — https://console.volcengine.com/ark/region:ark+cn-beijing/model/detail?Id=doubao-seedance-2-0
- 21 财经《1 元 1 秒!字节 Seedance2.0 定价出炉》 — https://m.21jingji.com/article/20260305/herald/9065ec418e98b284842d12cab6aa23fb.html
- 百度百科《Seedance 2.0》 — https://baike.baidu.com/item/Seedance%202/67291551
- 百度百科《即梦AI》 — https://baike.baidu.com/item/%E5%8D%B3%E6%A2%A6/64611350
- AI Wiki《Doubao》 — https://aiwiki.ai/wiki/doubao
- 百度百科《AI漫剧》 — https://baike.baidu.com/item/AI%E6%BC%AB%E5%89%A7/68788906
- 百度百科《关键帧动画》 — https://baike.baidu.com/item/%E5%85%B3%E9%94%AE%E5%B8%A7%E5%8A%A8%E7%94%BB/10223838
- 360 百科《关键帧》 — https://baike.so.com/doc/6737995-32354145.html
- 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- 中国政府网(繁体版)《人工智能生成合成内容标识办法》 — https://big5.www.gov.cn/gate/big5/www.gov.cn/zhengce/zhengceku/202503/content_7014286.htm
- 强制性国家标准 GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
- 央视网新闻《人工智能生成合成内容标识、传播、审核等如何发展?》 — https://news.cctv.cn/2025/03/15/ARTI3wMX1ohsE7LsLT4qBVrv250315.shtml
- 新疆生产建设兵团第十二师中级人民法院《AI 合成内容标识新规来了》 — http://btd12szy.btcourt.gov.cn/article/detail/2025/08/id/8955655.shtml
- 今日头条《AI 漫剧的变现逻辑,可以总结为"一基三翼"》 — https://www.toutiao.com/article/7678962433607664163
- KOCPC(英文)DataEye 2026 H1 AI 短剧报告转述 — https://en.kocpc.com.tw/archives/25203
- 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — https://www.thepaper.cn/newsDetail_forward_33993493
- 一品威客《AI 漫剧分镜设计指南》 — https://gonglue.epwk.com/322844.html
Doubao / Seedance (ByteDance Video Model) AI Comic Drama Platform Research
1. Introduction
Doubao-Seedance-2.0 is a video generation model developed by ByteDance's Seed team / Doubao Large Model team, served through Volcano Engine Ark. Among the 8 subjects in this group, Seedance is the sample with the strongest governance layer: it explicitly disallows real human faces as reference materials, requires real-person appearances to complete the person's own image and voice verification, and adopts a phased API opening strategy (a backdrop that includes copyright pushback from Hollywood studios and European entertainment industry groups).
At the same time, Seedance is the first model to simultaneously accept four input modalities (text, up to 9 images, up to 3 video clips, and up to 3 audio tracks), and it adopts a unified multimodal audio-visual joint generation architecture.
Seedance and Jimeng AI both belong to ByteDance, but differ in form: Jimeng is a creation platform for people (SaaS + channels), while Seedance is a model service for developers (MaaS). This group dissects the two separately precisely to distinguish a "platform Harness" from a "model Harness".
1.1. Developer and Model Identification
| Item | Content |
|---|---|
| Developer | ByteDance Seed team / Doubao Large Model team |
| Model name | Doubao-Seedance-2.0 |
| Model ID | doubao-seedance-2-0-260128 (recommended) |
| Release date | February 2026 |
| Input types | Text, image, video, audio |
| Output type | Video |
| API | /v3/contents/generations |
Release and integration timeline:
| Date | Event |
|---|---|
| February 2026 | Seedance 2.0 released |
| Feb 10, 2026 | Seedance concept drove a surge in the media sector (Dooke Culture, Rongxin Culture hit 20% daily limit-up) |
| Feb 12, 2026 | Doubao announced Seedance 2.0 officially integrated into the Doubao app, desktop, and web versions |
| Apr 2, 2026 | Volcano Engine opened Seedance 2.0 applications for general API customers; same day, the AI Innovation Roadshow·Wuhan stop announced open beta of the API for enterprise users |
| Apr 7, 2026 | Meitu's AI Agent product RoboNeo integrated Seedance 2.0 |
| Apr 12, 2026 | Tianyu Shuke's film-grade AI video creation platform CineART became one of the first integration platforms |
| Apr 14, 2026 | Volcano Engine officially launched the Seedance 2.0 series API services |
| May 2026 | 8 AI films produced with Seedance 2.0 screened at Cannes |
1.2. Positioning
Seedance 2.0's positioning is a "unified multimodal audio-visual joint generation architecture" — not a concatenation of multiple unimodal models, but joint generation under a unified architecture.
Key features:
- Native 2K resolution (2048×1080 or 1080×2048)
- Audio-video synchronization
- Generation speed 30% faster than Seedance 1.5 Pro
- Built-in professional camera movement, multi-shot narrative, and text generation capability
1.3. Pricing
| Billing mode | Price | Note |
|---|---|---|
| With video input (video editing) | 28 yuan / million tokens | Optimizes from existing material; lower compute requirement |
| Without video input (pure text-to-video) | 46 yuan / million tokens | Builds images, timing, and physics from scratch; compute consumption grows exponentially |
| Token consumption for a 15-second video | 308,880 tokens | Another source writes 308,880 as 30.888 万 |
| Cost per single 15-second clip (pure generation calculation) | About 15 yuan, or about 1 yuan per second | — |
Differentiated billing logic (21 财经/不慌实验室): video editing optimizes from existing material with a lower compute requirement (28 yuan); pure video generation builds images, timing, and physics from scratch, with exponentially higher compute consumption (46 yuan). This is the only model in this group that has publicly disclosed "tiered pricing by compute characteristics," which has direct implications for cost engineering.
Industry forecast reference: Goldman Sachs predicts the global AI video market will exceed 29 billion USD by 2030; in 2025, nearly 40% of leading short dramas adopted AI generation technology.
1.4. Open Forms
| Form | Note |
|---|---|
| Volcano Engine Ark API | /v3/contents/generations; Seedance 2.0 series API services officially launched Apr 14, 2026 |
| Doubao app / desktop / web | Integrated Feb 12, 2026 |
| Jimeng Web / app | Used through the Jimeng platform, subject to Jimeng's real-person material restrictions |
| Xiaoyunque | Short-drama Agent, used through the Jimeng ecosystem |
| Third-party integrations | RoboNeo (Meitu), CineART (Tianyu Shuke), etc. |
API specifications:
| Parameter | Value |
|---|---|
| Resolution | 480P / 720P / 1080P / 4K |
| Duration | 4~15s |
| Frame rate | 24fps |
| Concurrency (non-4K) | Shared 104 |
| Concurrency (4K) | Dedicated 1 |
| RPM (non-4K) | Shared 0.6k |
| RPM (4K) | Dedicated 0.015k |
| Task types | Multimodal video generation, video editing, video extension |
2. Glossary
| Term | English / Abbreviation | Definition |
|---|---|---|
| AI 漫剧 | AI Comic Drama | A content form between static comics and live-action short dramas, composed of comic storyboards plus dynamic audio-visual language |
| 动态漫 | Motion Comic | A slightly-dynamic video form built on static comic material through camera movement, zoom, localized effects, and dubbing |
| 分镜 / 分镜脚本 | Storyboard | Turns a written script into picture sketches, annotating each shot's composition, action, and duration |
| 角色一致性 | Character Consistency | The ability of the same character to keep facial features, clothing, body shape, and temperament stable across shots, episodes, and generations |
| 关键帧 | Keyframe | A frame defining a key state of animation or camera change (start and end points), corresponding to the "key drawing" in 2D animation |
| 中间帧 / 过渡帧 | In-between / Tween | Transition frames automatically generated between keyframes through interpolation algorithms |
| 首尾帧 | First-Last Frame | An image-to-video control method where the first and last frames are uploaded and the model fills in the intermediate motion trajectory |
| 图生视频 | Image-to-Video (I2V) | Inputting a single static image and having the model generate several seconds of animation |
| 文生视频 | Text-to-Video (T2V) | Directly generating a video from a text prompt |
| 口型同步 / 唇形同步 | Lip Sync | Overlaying audio onto a generated character and driving mouth movement to match the speech |
| 镜头语言 | Camera Language | The audio-visual expression system that conveys narrative information through shot size, angle, movement, composition, and editing rhythm |
| 多模态音视频联合生成 | Unified Audio-Visual Joint Generation | Seedance's core architectural feature: video and audio are generated jointly in a single inference rather than spliced stage by stage |
| 多镜头叙事 | Multi-shot Narrative | Automatically planning and switching between multiple coherent shots (wide → medium → close-up) within a single generation or single instruction |
| 四模态输入 | Four-Modality Input | Simultaneously inputting four modalities — text, image, video, audio; Seedance 2.0 supports up to 9 images + 3 videos + 3 audio tracks |
| 视频编辑 | Video Editing | Modifying and optimizing existing video material; Seedance prices it at 28 yuan / million tokens |
| 视频延长 | Video Extension | Extending the duration on top of an existing video |
| 数字分身 / 真人校验 | Digital Avatar Verification | Users must record their own image and voice to complete verification before their AI persona can appear on screen |
| 数字分身认证 | Digital Avatar Authentication | The authentication mechanism Jimeng introduced in February 2026 |
| MaaS | Model as a Service | Model as a Service; Seedance is provided in the form of the Volcano Ark API |
| 算法备案 | Algorithm Filing | The filing obligation required by the Interim Measures for the Administration of Generative AI Services |
| 显式标识 / 隐式标识 | Explicit / Implicit Label | The two types of statutory labels for AI-generated synthetic content: explicit labels are user-perceptible prompts; implicit labels are embedded in file metadata |
| AIGC 元数据字段 | AIGC Metadata Field | The metadata implicit-label field specified by the mandatory national standard GB 45438—2025 |
| tokens | tokens | The model billing unit; a 15-second video consumes 308,880 tokens |
| RPM | Requests Per Minute | Requests per minute; Seedance shares 0.6k for non-4K and dedicates 0.015k for 4K |
3. Functional Description
3.1. Multimodal Reference Video Generation
Seedance 2.0 supports generating video from reference images, videos, and audio, and is the first model to simultaneously accept four input modalities:
| Modality | Limit |
|---|---|
| Text | Prompt |
| Image | Up to 9 |
| Video | Up to 3 clips |
| Audio | Up to 3 tracks |
This capacity design is exactly consistent with Hailuo H3 (up to 12 reference files: 9 images + 3 videos + 3 audio tracks), placing it among the group's two highest-capacity tiers.
The difference lies in approach: Seedance takes the expansion route (directly accepting more modalities), while Hailuo takes the compression route (H3-Context-IR compresses about 100k tokens to about 4k tokens). The former is simple and direct but has a hard cap; the latter is complex but offers greater capacity flexibility.
3.2. Video Editing and Extension
| Task type | Note | Billing |
|---|---|---|
| Multimodal video generation | Generates new video from text/reference | 46 yuan / million tokens (without video input) |
| Video editing | Modifies and optimizes existing material | 28 yuan / million tokens (with video input) |
| Video extension | Extends duration on top of an existing video | Includes video input, billed at the 28 yuan tier |
The cost advantage of video editing and extension is worth engineering attention: 28 yuan vs 46 yuan, a 39% difference. For a production model of "generate key shots first, then edit/extend to fill in," this price gap is an important cost lever.
3.3. Multi-shot Narrative and Camera Movement
- Multi-shot narrative: automatically plans and switches between multiple coherent shots within a single generation.
- Built-in professional camera movement: the model has built-in camera capability, no external specification needed.
This is an approach of pushing L3 orchestration down into the model — sharing the same origin as Jimeng (both in the ByteDance ecosystem), which also means users cannot explicitly edit the shot sequence.
3.4. First-Last Frames and Text Generation
- First-last frame video generation: supports specifying the first + last frames, with the model filling in the middle.
- Text generation capability: the model can generate in-image text, which is directly valuable for titles, signs, subtitles, etc., in comic dramas.
4. Platform Architecture
4.1. Overall Architecture
图 4-1|Seedance 平台总体架构:从集成生态层到治理层
数据来源:基于本文分析绘制的示意图。
┌──────────────────────────────────────────────────────────────────┐
│ 集成生态层 豆包 App/电脑端/网页版 · 即梦 Web/App · 小云雀 │
│ RoboNeo(美图)· CineART(天娱数科)· 其他第三方 Agent │
├──────────────────────────────────────────────────────────────────┤
│ 服务层 火山引擎方舟 API /v3/contents/generations │
│ 并发:非 4K 共享 104 / 4K 独享 1 │
│ RPM:非 4K 共享 0.6k / 4K 独享 0.015k │
├──────────────────────────────────────────────────────────────────┤
│ 任务层 多模态生视频 · 视频编辑 · 视频延长 │
├──────────────────────────────────────────────────────────────────┤
│ 模型层 Doubao-Seedance-2.0(doubao-seedance-2-0-260128) │
│ 统一多模态音视频联合生成架构 │
│ 原生 2K · 24fps · 4~15s · 多镜头叙事 · 专业运镜 · 文字生成 │
├──────────────────────────────────────────────────────────────────┤
│ 治理层 真人人脸禁用 · 数字分身真人校验 · API 分阶段开放 │
│ · 成本护栏(约 1 元/秒) │
└──────────────────────────────────────────────────────────────────┘ 4.2. Model Layer
Doubao-Seedance-2.0 adopts a unified multimodal audio-visual joint generation architecture — not a two-stage scheme of "generate video first, then add dubbing," but producing audio and video jointly in a single inference.
Key specifications: native 2K resolution (2048×1080 or 1080×2048); API-side resolution options 480P / 720P / 1080P / 4K; duration 4~15s; frame rate 24fps; generation speed 30% faster than Seedance 1.5 Pro.
4.3. Service Layer
The core of the service layer is differentiated concurrency and rate configuration:
| Parameter | Non-4K | 4K |
|---|---|---|
| Concurrency | Shared 104 | Dedicated 1 |
| RPM | Shared 0.6k | Dedicated 0.015k |
The dedicated concurrency for 4K is 1 with an RPM of 0.015k (i.e., 15 requests per minute), a significant resource constraint. This means 4K generation is not suitable for batch production — only for a small number of key shots.
4.4. Integration Ecosystem Layer
Seedance's integration ecosystem is this group's most representative sample of a "model being integrated by Agents":
| Integrator | Type | Date |
|---|---|---|
| Doubao app / desktop / web | Own product | 2026-02-12 |
| Jimeng Web / app | Same-ecosystem platform | 2026-02 |
| Xiaoyunque | Short-drama Agent | — |
| RoboNeo (Meitu's AI Agent product) | Third-party Agent | 2026-04-07 |
| CineART (Tianyu Shuke's film-grade AI video creation platform) | Third-party platform | 2026-04-12 |
5. Harness Design
5.1. L1 Context Engineering Layer
Seedance's L1 is a four-modality unified context: text + up to 9 images + up to 3 videos + up to 3 audio tracks.
Compared with its predecessor, this is a shift from "text-primary, image-secondary" to "four-modality parity." For AI comic dramas, this means a single generation can simultaneously anchor characters (image), scenes (image), camera movement (video), voice (audio), and action (text).
Assessment: it is tied with Hailuo for this group's strongest L1 tier, but with a different approach — Seedance expands, Hailuo compresses. Seedance's advantage is modality parity with simple implementation; Hailuo's advantage is greater capacity flexibility (100k→4k compression).
Gaps:
- No public context compression mechanism — once the reference material exceeds the 9+3+3 limit, users must trim it themselves, with no priority-strategy guidance.
- No public cache-reuse mechanism — whether context can be reused when repeatedly referencing the same character across shots is not documented, marked
[To be filled]. - The weight-allocation mechanism for the four input modalities is not public — how the model balances the priority of image, video, audio, and text conditions is undocumented.
5.2. L2 Tools and Execution Layer
Seedance's L2 is a standard model API form:
| Tool | Note |
|---|---|
/v3/contents/generations API | Unified generation entry point |
| Task types | Multimodal video generation, video editing, video extension |
| First-last frame video generation | Supported |
| Text generation | Built into the model |
Assessment: Seedance's L2 ranks in the "medium" tier. It provides a complete model-capability API, but has no Agent-native tool layer — no CLI, no Skills, no Function Calling tool descriptions, no MCP server. It follows the "integrated by third-party Agents" route: platforms such as RoboNeo and CineART each implement their own tool wrappers.
This contrasts with PixVerse's CLI + Skills (proactively adapting to Claude Code, Codex, Cursor, OpenClaw): Seedance waits for others to integrate it, while PixVerse proactively adapts.
Gaps:
- No official CLI / Skills / MCP support, marked
[To be filled]. - Coarse tool granularity — only the three task types "generate," "edit," and "extend," lacking fine-grained tools (Lip Sync, local redraw, camera control, etc.).
- No sandbox, no parallel/serial scheduling mechanism (concurrency is indirectly governed by RPM limits).
5.3. L3 Orchestration and Control Layer
Seedance's L3 is multi-shot narrative — the model has built-in professional camera movement and shot planning, and can complete multi-shot switching within a single generation.
Assessment: sharing the same origin as Jimeng, it belongs to the "implicit in-model planning" route. Its advantages are zero configuration and fast onboarding; its drawbacks are that it cannot be visualized, precisely edited, versioned, or exported.
Gap checklist:
- No node graph, no storyboard export, no timeline editing.
- No public conditional branching, batch processing, or loop mechanisms.
- No interruption and resume mechanism — long tasks must be rerun after a failure.
- The shot-count limit and switching logic for multi-shot narrative are not publicly documented, marked
[To be filled].
For complex continuous production of 100 episodes, Seedance's orchestration must be carried by an upper-layer Agent — this is precisely the value that third-party platforms like CineART and Xiaoyunque provide.
5.4. L4 Memory and State Layer
Seedance's L4 is this group's shortfall sample, and the most typical case of a "stateless model."
Known capabilities: stable maintenance of character features; high-precision restoration of detail, materials, voice, visual-effects style, and camera movement at high resolution (per Volcano Engine console model description).
Explicit gaps:
- No asset library: as a model service, Seedance does not provide a persistent asset carrier for characters/props/scenes. Compared with Vidu's subject library, PixVerse Character, and the Daydream character library, this is a fundamental absence.
- No session concept: the API is stateless, with each call independent. Cross-shot and cross-episode consistency relies entirely on the caller re-injecting references on every request.
- No plot state machine.
This precisely corroborates this group's core thesis: cross-shot and cross-episode consistency in AI comic dramas is essentially a cross-session state-keeping problem, and the model itself does not solve it. Seedance hands this responsibility entirely to the caller — any team building an AI comic drama production system on Seedance must build a complete L4 at the upper layer itself.
Conversely, this is also Seedance's reasonable positioning as a "pure model layer": it provides a strong L1 (four modalities) and strong generation capability, leaving L4 for an upper-layer Harness to implement.
5.5. L5 Evaluation and Observability Layer
Seedance's L5 signals:
| Signal | Note |
|---|---|
| Video availability rate | Volcano Engine describes it as "greatly improved" with no specific figure, marked [To be filled] |
| Cannes screening | In May 2026, 8 AI films produced with Seedance 2.0 screened at Cannes — external quality verification |
| Generation speed | 30% faster than Seedance 1.5 Pro |
Assessment: Seedance's L5 ranks in the "strong" tier, but it is based on external verification (Cannes screening) rather than platform-built-in evaluation. This is consistent with the other platforms in this group — currently no platform provides a built-in trajectory tracker, regression set, or availability-rate dashboard.
Observability gaps:
- No official availability-rate statistic (only the qualitative "greatly improved" description).
- No generation-trajectory tracking or structured logging.
- No cost-analysis dashboard — although there are explicit token billing rules, usage-visualization tooling is missing (whether the Volcano Ark console provides it is undocumented, marked
[To be filled]).
5.6. L6 Governance and Security Layer
Seedance's L6 is the strongest in this group, tied with Jimeng.
Mechanism one: real-human-face disallowed
When using Seedance 2.0 on platforms such as Jimeng Web and Xiaoyunque, the system explicitly prompts that real human faces are not currently supported as reference material.
Mechanism two: digital avatar real-person verification
In the Jimeng app and Doubao app, if users want a real person's image to appear on screen, they must record their own image and voice to complete real-person verification before creating a digital avatar.
This is the only mechanism in this group requiring the dual-factor "image + voice" verification, stricter than a single image verification.
Mechanism three: phased API opening
As of late March 2026, Seedance 2.0 was still restricted to the Ark experience center, with the API opened in phases. The backdrop includes copyright pushback from Hollywood studios and European entertainment industry groups. Applications for general API customers opened on Apr 2, 2026, and the series of API services officially launched on Apr 14.
This reflects a prudent governance posture: before copyright disputes are clarified, it prefers to limit the scope of access rather than recklessly opening it all at once.
Mechanism four: regulatory compliance
Doubao is bound by the Interim Measures for the Administration of Generative AI Services (effective August 2023) and must complete algorithm filing, uphold core values, and label synthetic media.
Mechanism five: cost guardrail
The roughly 1 yuan/second pricing itself constitutes a cost guardrail; the 4K dedicated concurrency of 1 and RPM of 0.015k rate limiting is a resource guardrail.
Gaps:
- No positive statement was found on whether Seedance has built-in metadata implicit labeling compliant with GB 45438—2025 — although the law requires "labeling synthetic media," the specific implementation (explicit-label position/height/duration, AIGC metadata fields) has no public technical documentation, marked
[To be filled]. - 4K concurrency of 1 and RPM 0.015k are hard capacity constraints; batch production must plan quotas in advance.
- The API's opening scope may still change with copyright negotiations, creating supply uncertainty.
5.7. Six-Layer Capability Matrix
| Layer | Seedance's implementation | Maturity | Main gaps |
|---|---|---|---|
| L1 Context engineering | Four-modality unified context (9 images + 3 videos + 3 audio tracks + text) | Strong | No compression mechanism; no cache reuse; modality weights unclear |
| L2 Tools and execution | /v3/contents/generations API; generate/edit/extend; first-last frames | Medium | No CLI/Skills/MCP; coarse tool granularity |
| L3 Orchestration and control | Multi-shot narrative (implicit in-model planning); professional camera movement | Strong | Not editable, not exportable, not versionable |
| L4 Memory and state | Stable maintenance of character features (model side); no asset library | Medium | No asset library, no session, no state machine — all must be built at the upper layer |
| L5 Evaluation and observability | Cannes external verification; availability "greatly improved" (no figure) | Strong | No built-in evaluation, no trajectory tracking |
| L6 Governance and security | Real-human-face disallowed + digital avatar real-person verification + phased API opening + cost guardrail | Strongest | Implicit-label implementation not public; 4K capacity constrained |
6. Practical Cases
6.1. Case One: Third-Party Agent Product Integration
- Background: for an AI video model to enter the professional creation pipeline, it must be integrable by third-party Agents and platforms.
- Approach: Seedance 2.0 is delivered as a
/v3/contents/generationsAPI and integrated by multiple parties — on Apr 7, 2026, Meitu's AI Agent product RoboNeo integrated it; on Apr 12, 2026, Tianyu Shuke's film-grade AI video creation platform CineART became one of the first integration platforms; the same-ecosystem Doubao app/desktop/web, Jimeng Web/app, and Xiaoyunque have also integrated it. - Result: specific call volumes and business data have not been publicly disclosed, marked
[To be filled]. - Harness interpretation: this is a typical sample of the division of labor between a "model Harness" and a "platform Harness." Seedance provides L1/L2 (context and execution primitives), while RoboNeo and CineART build L3/L4 (orchestration and state) on top of it. If you want to build AI comic dramas on Seedance, you play the role of CineART — the model will not manage assets and plot state for you.
6.2. Case Two: The 8 AI Films Screened at Cannes
- Background: the professional recognition of AI-generated content requires external verification in an authoritative setting.
- Approach: in May 2026, 8 AI films produced with Seedance 2.0 screened at Cannes.
- Result: this is the group's only quality-verification event at an international top-tier film festival. The specific film titles, screening sections, and jury feedback have not been publicly disclosed, marked
[To be filled]. - Harness interpretation: the Cannes screening is an external signal for L5, not a platform capability. It proves that Seedance's output quality reaches a professionally usable standard, but it does not mean your project automatically reaches that standard either — between them lies the large engineering gap of L4 (consistency) and L3 (orchestration).
6.3. Case Three: Phased API Opening and Copyright Pushback
- Background: the training data and generated content of AI video models involve substantial copyright disputes.
- Approach: as of late March 2026, Seedance 2.0 was still restricted to the Ark experience center, with the API opened in phases; applications for general API customers opened on Apr 2, 2026, and the series of API services officially launched on Apr 14. Behind the cautious opening cadence was the copyright pushback from Hollywood studios and European entertainment industry groups.
- Result: the phased opening strategy allowed the model's capabilities to be steadily released amid the disputes; the specific opening scope and quota rules have not been fully disclosed, marked
[To be filled]. - Harness interpretation: this is a direct case of L6 governance affecting L2 availability — governance decisions can restrict tool-layer availability. For production systems relying on Seedance, this means supply uncertainty exists, so a multi-model fallback strategy should be designed.
6.4. Case Four: Cost Engineering via Differentiated Billing
- Background: large-scale production of AI comic dramas is extremely cost-sensitive — industry per-hundred-thousand-plays revenue has fallen from about 100 yuan to 5~10 yuan, with a break-even rate below 1.3%.
- Approach: Seedance uses differentiated billing — with video input (video editing) 28 yuan / million tokens, without video input (pure text-to-video) 46 yuan / million tokens. A 15-second video consumes 308,880 tokens, about 15 yuan per clip in pure generation mode, or about 1 yuan per second.
- Result: video editing is about 39% cheaper than pure generation. For the production model of "generate key shots first, then edit/extend to fill in," this is a significant cost lever. The specific savings depend on project structure, marked
[To be filled]. - Harness interpretation: this is the cost-engineering practice most worth learning in this group. It suggests a production strategy: divide shots into "keyframes that must be generated from scratch" and "completion frames that can be edited/extended from existing material," with the former going through the 46 yuan tier and the latter through the 28 yuan tier. This layering is implemented at the L3 orchestration layer and directly affects the effectiveness of the L6 cost guardrail.
7. Summary
7.1. Strengths
- Strongest L6 governance in the group: real-human-face disallowed + digital avatar real-person verification (image + voice dual factors) + phased API opening, the most complete and cautious governance combination in this group.
- One of the highest four-modality input capacities: 9 images + 3 videos + 3 audio tracks + text, tied with Hailuo.
- Unified multimodal audio-visual joint generation architecture: not stage-by-stage splicing, but unification at the architecture level.
- Native 2K + 4K support: complete resolution tiers (480P / 720P / 1080P / 4K).
- Strongest external quality verification: 8 AI films screened at Cannes.
- Transparent differentiated billing logic: the 28 yuan vs 46 yuan pricing logic has a public compute-based explanation, aiding cost engineering.
- Mature integration ecosystem: integrated by professional platforms such as RoboNeo and CineART, verifying the model's integrability.
- Clear speed improvement: generation speed 30% faster than Seedance 1.5 Pro.
7.2. Limitations
- L4 entirely missing: no asset library, no session, no state machine. All cross-shot and cross-episode consistency work must be built by the caller — this is the group's most thorough "hand L4 to the upper layer" design.
- No Agent-native tool layer: no CLI, no Skills, no MCP — it only waits for third parties to integrate it.
- 4K capacity constrained: dedicated concurrency 1, RPM 0.015k, cannot do batch production.
- Orchestration uncontrollable: multi-shot narrative is implicit in-model planning, not editable, not exportable, not versionable.
- No quantified availability: only the qualitative "greatly improved" description.
- Relatively high cost: the pure-generation cost of about 1 yuan/second is higher than Hailuo (0.8 yuan/second at 2K) and Vidu (reportedly 1/3 of the industry average).
- Supply uncertainty: the API's opening scope is affected by copyright disputes and may change.
- Implicit-label implementation not public: though bound by regulation, the specific technical implementation has no documented basis.
- 4~15s duration: on par with most platforms in this group; long shots require splicing.
7.3. Applicability Boundary
| Suitable for | Not suitable for |
|---|---|
| Teams that already have their own orchestration/asset systems (need only a strong model layer) | Teams wanting the platform to provide a complete consistency solution |
| Shots needing complex four-modality references | Need to connect generation into an AI-native development environment (no CLI/Skills) |
| Content needing native audio-visual joint generation | Need 4K batch production (concurrency constrained) |
| Commercial delivery needing strong compliance-governance backing | Extremely budget-sensitive projects (about 1 yuan/second) |
| High-quality single shots for professional film/advertising | Complex series needing visual fine-grained orchestration |
| Production models needing "edit/extend" cost reduction | Critical businesses needing stable long-term supply assurance (opening scope may change) |
7.4. Selection Recommendations
- The core reason to choose Seedance is model capability and governance maturity, not completeness. If you already have your own orchestration system and asset library and only need a strong model layer, Seedance is a quality choice.
- A complete L4 must be built yourself: this is a prerequisite for using Seedance, not an option. You need a storyboard sheet (structured JSON), character cards (with reference images and text anchors), scene cards, a plot state machine, and re-inject this state on every API call.
- Do cost engineering with differentiated billing: divide shots into "keyframes that must be generated from scratch" (46 yuan tier) and "completion frames that can be edited/extended from existing material" (28 yuan tier), routing them at the orchestration layer.
- Use 4K only for key shots: 4K has a dedicated concurrency of 1 and RPM 0.015k; reserve it for a small number of high-value shots, and run batch production at 1080P.
- Design a multi-model fallback: since the API's opening scope may change with copyright negotiations, the production system should abstract a model-layer interface and reserve the ability to switch to other models (e.g., Hailuo H3, Vidu Q3).
- Verify compliance capability yourself: though bound by the Interim Measures for the Administration of Generative AI Services to "label synthetic media," the specific implementation has no public technical documentation. Works distributed on domestic platforms must self-verify compliance with GB 45438—2025's explicit-label requirements (corner of the opening frame, font height no less than 5% of the shortest side, lasting no less than 2 seconds) and the AIGC-metadata implicit-label requirements.
- Estimate cost at about 1 yuan/second and reserve redundancy: a 15-second clip of 308,880 tokens costs about 15 yuan; calculate the total budget as project shot count × rework factor (recommended no lower than 1.3).
Information Gap Declaration
- Modality-weight mechanism of multimodal input: how the model balances the priority of the four conditions — image, video, audio, text — has no public explanation.
- The technical mechanism and quantitative metrics of "stable maintenance of character features": the Volcano Engine console only gives a qualitative description, with no consistency values or implementation notes.
- Specifications of multi-shot narrative: the maximum supported shot count in a single generation, the shot-switching logic, and whether a shot sequence can be specified have no public explanation.
- Context cache-reuse mechanism: whether context can be reused to reduce token consumption when repeatedly referencing the same character across shots has no public result.
- Whether Seedance has built-in metadata implicit labeling compliant with GB 45438—2025: the law requires labeling, but the specific technical implementation (writing AIGC metadata fields) has no public technical documentation.
- The specific availability-rate figure: only the qualitative "greatly improved" description exists, with no quantified data.
- API opening scope and quota rules: the specific criteria for the phased opening, the quota caps at each stage, and whether it will open further have not been fully disclosed.
- Detailed information on the 8 Cannes films: film titles, screening sections, and jury feedback have not been publicly disclosed.
- Whether official CLI / Skills / MCP support is provided: no relevant public information was found.
- Volcano Ark console's cost-analysis capability: whether it provides usage visualization, budget alerts, and a cost-analysis dashboard is undocumented.
- Capability differences between Seedance 2.5 and 2.0: Jimeng integrated Seedance 2.5 on Aug 5, 2026, but the magnitude of 2.5's parameter improvements over 2.0 has no official quantified disclosure.
- The specific capability boundary of video-editing tasks: which editing operations are supported (changing background, changing character, changing camera movement, changing duration, etc.) has no complete explanation.
8. References
- Volcano Engine Ark console, Doubao-Seedance-2.0 series — https://console.volcengine.com/ark/region:ark+cn-beijing/model/detail?Id=doubao-seedance-2-0
- 21 财经, "1 yuan 1 second! ByteDance Seedance 2.0 pricing released" — https://m.21jingji.com/article/20260305/herald/9065ec418e98b284842d12cab6aa23fb.html
- Baidu Baike, "Seedance 2.0" — https://baike.baidu.com/item/Seedance%202/67291551
- Baidu Baike, "Jimeng AI" — https://baike.baidu.com/item/%E5%8D%B3%E6%A2%A6/64611350
- AI Wiki, "Doubao" — https://aiwiki.ai/wiki/doubao
- Baidu Baike, "AI comic drama" — https://baike.baidu.com/item/AI%E6%BC%AB%E5%89%A7/68788906
- Baidu Baike, "Keyframe animation" — https://baike.baidu.com/item/%E5%85%B3%E9%94%AE%E5%B8%A7%E5%8A%A8%E7%94%BB/10223838
- 360 Baike, "Keyframe" — https://baike.so.com/doc/6737995-32354145.html
- CAC and three other departments, "Measures for the Labeling of AI-Generated Synthetic Content" (Guoxin Ban Tong Zi [2025] No. 2) — https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- Central Government website (Traditional Chinese version), "Measures for the Labeling of AI-Generated Synthetic Content" — https://big5.www.gov.cn/gate/big5/www.gov.cn/zhengce/zhengceku/202503/content_7014286.htm
- Mandatory national standard GB 45438—2025, "Cybersecurity Technology — Methods for Labeling AI-Generated Synthetic Content" — https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
- CCTV news, "How will the labeling, dissemination, and review of AI-generated synthetic content develop?" — https://news.cctv.cn/2025/03/15/ARTI3wMX1ohsE7LsLT4qBVrv250315.shtml
- Xinjiang Production and Construction Corps 12th Division Intermediate People's Court, "New rules for labeling AI synthetic content are here" — http://btd12szy.btcourt.gov.cn/article/detail/2025/08/id/8955655.shtml
- Toutiao, "The monetization logic of AI comic dramas can be summarized as 'one foundation, three wings'" — https://www.toutiao.com/article/7678962433607664163
- KOCPC (English), retelling of DataEye 2026 H1 AI short-drama report — https://en.kocpc.com.tw/archives/25203
- The Paper, "1,914 yuan to produce one episode? Comic dramas still trapped in hidden costs" — https://www.thepaper.cn/newsDetail_forward_33993493
- Epwk, "AI Comic Drama Storyboard Design Guide" — https://gonglue.epwk.com/322844.html