Vidu(生数科技)AI 漫剧平台研究


1. 介绍

Vidu 由北京生数科技股份有限公司联合清华大学发布,官方定位为"中国首个长时长、高一致性、高动态性视频大模型"。在本组 8 个对象中,Vidu 是L4(记忆与状态)形态最明确、最完整的商业平台:其"主体库"把角色、道具、场景保存为可一键引用的持久化条目,多主体一致性支持最多 7 张主体图片。

这直接命中本组核心论断——AI 漫剧是 Harness 六层中 L4 压力最大的场景,谁把 L4 做扎实,谁就能把 AI 漫剧从"能生成"推进到"能连续生产"。Vidu 用"参考生视频"这一范式,把一致性问题从模型能力问题转化为资产工程问题,是本组最有代表性的 L4 样本。

1.1. 开发商与资本

内容
公司北京生数科技股份有限公司
成立时间2023 年
创始团队清华大学朱军团队
技术底座U-ViT 架构(朱军团队 2022 年 9 月提出,全球首个 Diffusion 与 Transformer 融合架构)
用户覆盖全球 200 多个国家和地区
客户与伙伴索尼电影、腾讯动漫、阅文集团
2025 年增长用户和收入超 10 倍增长

融资与资本化:

时间事件
2026 年 2 月完成超 6 亿元一轮融资(中关村科学城、星连资本领投,万兴科技、视觉中国、拓尔思战略投资,启明创投等跟投)
2026 年 3 月 30 日完成股份制改造
近 20 亿元人民币 B 轮融资,由阿里云领投,中网投、九安海棠、好未来、光合创投战略投资;星连资本、达泰资本、建发新兴投资、BV 百度风投、卓源亚洲等老股东追加
融资完成后估值超 20 亿美元;或最快 2026 年上半年启动港股 IPO

战略路线:基座世界模型(Foundation World Model)→ 世界生成模型 WGM(数字世界)+ 世界行动模型 WAM(物理世界);2025 年 12 月开源世界行动模型 Motus

1.2. 定位

Vidu Q3 的定位是"为剧而生"——工业化内容生产,参考生全能力矩阵,声画同出,商用交付就绪。

这个定位在本组中最贴近 AI 漫剧的实际需求:不追求通用创意视频的广度,而是把资源集中在"剧集"这一类内容的工业化交付上。三层能力矩阵(特效、音效、场景)也印证了这一点。

1.3. 定价

Vidu 官方宣称定价"为行业平均水平的 1/3"。

第三方 API 聚合报价(Atlas Cloud),Vidu Q3 每秒价格:

模型/任务原价折后价
Q3-Mix 参考生视频$0.125$0.106
Q3 参考生视频$0.05$0.042
Q3-Pro 首尾帧$0.05$0.042
Q3-Turbo 图生视频$0.04$0.034

C 端定价缺口:viduq3.net 为第三方站($9.9/$29.9/$49.9/$99.9 一次性积分包),vidu.cn 官网未见完整价目表。本组未能获取 Vidu 官方 C 端定价页,标 。

其他成本相关:错峰模式无限积分(错峰时段无限次免费生成视频)。

1.4. 开放形态

形态说明
SaaSVidu Agent、Vidu Claw("雇佣 AI 创意员工"),面向个人/小工作室
MaaSVidu AI 开放平台、Vidu.API,面向企业批量生产/开发者集成
云平台已登陆阿里云百炼平台
实时模型Vidu S1 实时交互模型

2. 名词解释

术语英文 / 缩写释义
AI 漫剧AI Comic Drama介于静态漫画与真人短剧之间的内容形态,以漫画分镜加动态视听语言构成
动态漫Motion Comic以静态漫画素材为基础,通过运镜、缩放、局部动效与配音形成的轻微动态视频形态
分镜 / 分镜脚本Storyboard将文字剧本转化为画面草图,标注每个镜头的构图、动作、时长
角色一致性Character Consistency同一角色在跨镜头、跨集、跨次生成中保持五官、服装、体型、气质稳定的能力
关键帧Keyframe定义动画或运镜变化关键状态的帧(起点与终点),对应二维动画中的"原画"
中间帧 / 过渡帧In-between / Tween关键帧之间通过插值算法自动生成的过渡帧
首尾帧First-Last Frame上传首帧与尾帧,由模型补全中间运动轨迹的图生视频控制法
图生视频Image-to-Video(I2V)输入一张静态图片,由模型生成数秒动画
口型同步 / 唇形同步Lip Sync把音频叠加到生成角色上并驱动嘴部动作匹配发音
镜头语言Camera Language通过景别、角度、运动、构图与剪辑节奏传递叙事信息的视听表达体系
参考生视频Reference-to-VideoVidu 的核心范式:上传 3 张或更多参考图,模型把角色/场景/服化道建模为可复用主体并在输出中保持一致
万物可参Everything ReferenceableVidu 对参考生视频的表述:角色、场景、服化道均可作为参考主体
主体库Subject Bank把角色、道具、场景保存为持久化条目,一键选择参考主体
多主体一致性Multi-Subject Consistency单次生成中保持多个主体一致的能力,Vidu 支持最多 7 张主体图片
原生镜头控制Native Camera Control以帧级导演指令精确控制镜头运动
Smart CutSmart Cut自动场景边界检测与转场
声画同出Native Audio-Visual Output视频与音频在同一次推理中同时产出
U-ViTU-ViT朱军团队 2022 年 9 月提出的架构,全球首个 Diffusion 与 Transformer 融合架构
Vidu Agent / Vidu ClawVidu Agent / Vidu ClawSaaS 层产品;Vidu Claw 定位为"雇佣 AI 创意员工"
MaaSModel as a Service模型即服务;Vidu 的 Vidu AI 开放平台与 Vidu.API
SuperClueSuperClue首个专项参考生能力评测;Vidu Q3 拿下多图参考与单图参考双榜第一
错峰模式Off-Peak Mode错峰时段无限次免费生成视频的模式
显式标识 / 隐式标识Explicit / Implicit LabelAI 生成合成内容的两类法定标识:显式为用户可感知提示;隐式嵌入文件元数据
AIGC 元数据字段AIGC Metadata Field强制性国标 GB 45438—2025 规定的元数据隐式标识字段

3. 功能说明

3.1. 参考生视频

参考生视频(Reference-to-Video)是 Vidu 的核心范式,官方表述为"万物可参"——把角色、场景、服化道提取为可复用建模素材。上传 3 张或更多参考图,模型将其建模为可复用主体并在输出中保持一致。

与同组对象对比:

平台一致性范式参考容量
Vidu参考生视频(万物可参)最多 7 张主体图片
海螺 AI@ 引用系统 + Context-IR最多 12 个参考文件(9 图 + 3 视频 + 3 音频)
PixVerseCharacter 一致性角色未公开 [待填写]
可灵元素引用未公开 [待填写]
白日梦 AI角色库(5 张参考图创建角色)未公开
豆包/Seedance四模态参考最多 9 图 + 3 视频 + 3 音频

Vidu 的特点是参考数量明确、主体概念明确。"主体"这个抽象很关键——它把角色、道具、场景统一为同一类可引用对象,而不是分别做"角色库"和"场景库"。

3.2. 主体库与多主体一致性

主体库(Subject Bank):将角色、道具和场景保存在主体库中,一键选择参考主体。

多主体一致性:上传最多 7 张主体图片,单次生成中保持多个主体一致。

这是本组唯一明确公开容量上限的资产库机制(7 张),也是唯一明确包含三类资产(角色、道具、场景)的机制。

3.3. 特效与音效能力矩阵

Vidu Q3 的三层能力矩阵:

数量明细
特效层6 大粒子、流体、动力学、运镜、转场、光影
音效层5 大环境、动态、氛围、拟音、情绪
场景层4 大短剧、漫剧、影视剧、广告

场景层明确列出"漫剧",是本组唯一在能力矩阵中把漫剧作为独立场景类别的平台。

3.4. 镜头控制与场景边界检测

  • 原生镜头控制:帧级导演指令,可精确控制镜头运动。
  • Smart Cut:自动场景边界检测与转场。

这两项组合起来,实际上提供了"从长素材自动切分为镜头"与"在镜头内精确控制运镜"的双向能力。

3.5. 其他能力与模板

能力说明
单次时长16 秒连续 1080p @ 24fps
原生音视频同步声画同出
漫画图片生成动画漫画图转动画,直接对应动态漫生产
首尾帧支持
错峰模式无限积分错峰时段无限次免费生成
模板库亲吻、拥抱、万物生花、AI 换装等
Vidu S1实时交互模型
Vidu Claw"雇佣 AI 创意员工"
Vidu AgentSaaS 层 Agent 产品

4. 平台架构

图 4-1|Vidu 平台六层架构:从 U-ViT 底座到主体库与商业化分层

Vidu 平台架构:自 U-ViT 底座至商业化分层的六层栈 示意:基于本文总体架构分析绘制 · 资产层为核心差异 商业化分层 SaaS:Vidu Agent · Vidu Claw(个人/小工作室) MaaS:Vidu AI 开放平台 · Vidu.API(企业/开发者)|云平台:阿里云百炼 编排与控制 原生镜头控制(帧级导演指令)· Smart Cut(场景边界检测与转场) 能力矩阵 特效层 6 · 音效层 5 · 场景层 4(含漫剧) 资产层 · 主体库(本图重点) 主体库(角色/道具/场景)· 多主体一致性(最多 7 张) 模型层 Q1(叙事逻辑)· Q2(AI 演技)· Q3(工业化生产)· S1(实时交互)· Motus(世界行动模型) 底座 U-ViT 架构:Diffusion + Transformer 融合(全球首个融合架构) 结构解读:U-ViT 底座支撑 Q 系模型栈,主体库(L4)锚定跨集一致性——这是 Vidu 的核心壁垒。

数据来源:基于本文分析绘制的示意图。

4.1. 总体架构

┌────────────────────────────────────────────────────────────────┐
│  商业化分层  SaaS:Vidu Agent · Vidu Claw(个人/小工作室)         │
│              MaaS:Vidu AI 开放平台 · Vidu.API(企业/开发者)      │
│              云平台:阿里云百炼                                    │
├────────────────────────────────────────────────────────────────┤
│  编排与控制  原生镜头控制(帧级导演指令)                           │
│              Smart Cut(场景边界检测与转场)                       │
├────────────────────────────────────────────────────────────────┤
│  能力矩阵  特效层 6 · 音效层 5 · 场景层 4(含漫剧)                │
├────────────────────────────────────────────────────────────────┤
│  资产层  主体库(角色/道具/场景)· 多主体一致性(最多 7 张)         │
├────────────────────────────────────────────────────────────────┤
│  模型层  Vidu Q1(叙事逻辑)· Q2(AI 演技)· Q3(工业化生产)       │
│          Vidu S1(实时交互)· Motus(世界行动模型,2025-12 开源)   │
├────────────────────────────────────────────────────────────────┤
│  底座  U-ViT 架构(Diffusion + Transformer 融合)                 │
└────────────────────────────────────────────────────────────────┘

4.2. 模型层

模型定位
Vidu Q1重新定义叙事逻辑,夯实基础生成能力与故事线推进框架
Vidu Q2解锁 AI 演技,赋予虚拟角色微表情与肢体表现力
Vidu Q3工业化内容生产,参考生全能力矩阵,声画同出,商用交付就绪;"为剧而生"
Vidu S1实时交互模型
Motus世界行动模型(WAM),2025 年 12 月开源

Q1→Q2→Q3 的演进路径值得注意:先解决"故事能不能推进"(叙事逻辑),再解决"角色会不会表演"(微表情与肢体),最后解决"能不能工业化交付"(参考生能力矩阵 + 声画同出 + 商用就绪)。这是一条从内容到工程的清晰路线。

4.3. 资产层

主体库是本架构的核心差异点,详见 5.4 节 L4 分析。

4.4. 商业化分层

SaaS 层:Vidu Agent、Vidu Claw,面向个人/小工作室。Vidu Claw 的定位表述为"雇佣 AI 创意员工"。

MaaS 层:Vidu AI 开放平台、Vidu.API,面向企业批量生产与开发者集成。已登陆阿里云百炼平台

这种分层让同一套模型能力既能被创作者直接使用,也能被企业集成进自有生产系统。


5. Harness 设计

5.1. L1 上下文工程层

Vidu 的 L1 机制是参考生视频(万物可参):把角色、场景、服化道提取为可复用建模素材,以参考图形式注入上下文。

与同组对象相比,Vidu 的 L1 有两点特征:

  1. 参考即建模:官方表述强调"建模为可复用主体",说明参考图不是简单的条件输入,而是会被建模成主体表示。这比"把图贴上去"更进一步。
  2. 容量明确:最多 7 张主体图片。虽然数量少于 Seedance(9 图 + 3 视频 + 3 音频)与海螺(12 个参考文件),但明确公开上限本身是工程友好性——使用者可以做确定性规划。

缺口

  1. 无公开的上下文压缩机制(对比海螺 H3-Context-IR 的 25:1 压缩)。
  2. 参考是否支持视频与音频模态,未见明确说明(对比 Seedance 与海螺明确支持),标 [待填写]
  3. 无公开的上下文优先级排序与缓存复用机制。

5.2. L2 工具与执行层

Vidu 的 L2 以能力矩阵形式组织,这是本组结构化程度最高的工具层:

类别工具
特效(6)粒子、流体、动力学、运镜、转场、光影
音效(5)环境、动态、氛围、拟音、情绪
生成方式参考生视频、图生视频、首尾帧、漫画图片生成动画
模板亲吻、拥抱、万物生花、AI 换装等
APIVidu.API(MaaS)

判断:Vidu 的 L2 属"强"档。5 大音效层是本组独有的(其他平台多为后期配音或原生音频,不做音效分类)。6 大特效层覆盖了漫剧最常见的视觉需求。

可编程性:提供 Vidu.API(MaaS)与阿里云百炼集成,可编程性明确。但未检索到 CLI 或 Skills 类工具(对比 PixVerse),标 [待填写]

5.3. L3 编排与控制层

Vidu 在 L3 上有三条明确线索:

  1. 原生镜头控制:帧级导演指令。这是本组粒度最细的镜头控制——不是"全景/中景/特写"的粗粒度选择,而是帧级精确指定。
  2. Smart Cut:自动场景边界检测与转场。这解决了长素材自动切分为镜头的问题。
  3. Vidu Agent / Vidu Claw:SaaS 层的 Agent 化编排,Vidu Claw 定位为"雇佣 AI 创意员工"。

判断:Vidu 的 L3 属"强"档。它的编排思路与可灵(分镜指令)、海螺(时间轴)都不同——是帧级精确控制 + 自动边界检测的组合,既给精度又给自动化。

缺口:无可编辑节点图(对比 PixVerse Canvas、ComfyUI);无公开的条件分支与批处理机制;编排结果是否可导出与版本化,未见说明,标 [待填写]

5.4. L4 记忆与状态层

Vidu 的 L4 是本组商业平台中最强,也是本组核心论断的最佳注脚。

机制说明对比优势
主体库角色、道具、场景保存为持久化条目,一键选择参考主体唯一明确包含三类资产的资产库
多主体一致性最多 7 张主体图片,单次生成保持多主体一致唯一明确公开容量上限
参考生视频把角色/场景/服化道建模为可复用素材"主体"抽象统一了三类资产

为什么这很重要:AI 漫剧的跨镜头、跨集一致性本质是跨会话状态保持问题。Vidu 的主体库把这件事产品化了——使用者不需要自己搭一套资产管理系统,平台提供了:

  • 状态载体(主体条目)
  • 状态命名(角色/道具/场景三类)
  • 状态引用(一键选择)
  • 状态容量(最多 7 张)

这是本组唯一把 L4 四要素全部显式化的商业平台。

仍存在的缺口

  1. 容量上限 7 张是否够用:对多角色群戏,7 张主体图片(含角色、道具、场景)可能偏紧。对比海螺的 12 个参考文件与 Seedance 的 9 图 + 3 视频 + 3 音频,Vidu 的容量偏小。
  2. 无剧情状态机:主体库解决"长得一样",不解决"剧情状态连续"。跨集的人物关系、伤势、时间线推进仍需外部维护。
  3. 跨项目复用规则未公开:主体库是否可跨项目、跨团队复用,未见说明,标 [待填写]

5.5. L5 评估与观测层

Vidu 的 L5 是本组商业平台中最有针对性依据的。

榜单/指标成绩
SuperClue(首个专项参考生能力评测)多图参考与单图参考双榜第一(断层优势)
Artificial Analysis 综合排名2026 年 1 月发布时登顶
Artificial Analysis Video Arena ELO1220~1244,全球第 2(第 1 为 Sora 2 约 1250+)

SuperClue 的成绩尤其关键:它是首个专项参考生能力评测,而"参考生"正是 AI 漫剧一致性的核心技术路径。Vidu 在这个专项上拿下双榜第一,说明其 L4 能力有第三方背书——这是本组唯一有"针对一致性"的第三方基准的平台。

缺口:无平台内建的可用率统计面板、回归集、轨迹追踪,标 [待填写]

5.6. L6 治理与安全层

Vidu 的 L6 机制是生态分层:SaaS(个人/小工作室)与 MaaS(企业批量生产/开发者集成)。

产品面向治理含义
SaaSVidu Agent、Vidu Claw个人/小工作室轻量、低门槛
MaaSVidu AI 开放平台、Vidu.API企业/开发者可集成、可批量
云平台阿里云百炼企业依托云厂商合规体系

判断:Vidu 的 L6 属"中"档。生态分层本身提供了使用边界的划分,但缺少明确的权限、配额、审计机制(对比 PixVerse Team Plan 的 RBAC + 积分上限 + 用量分析 + 统一计费)。

缺口

  1. 未检索到 AI 生成合成内容标识的公开说明——是否符合 GB 45438—2025 的显式标识与 AIGC 元数据隐式标识要求,无公开结果,标 [待填写]
  2. 未检索到真人形象校验机制——对比豆包/Seedance,Vidu 在这条红线上无公开信息。
  3. 无公开团队协作与权限管理能力——对比 PixVerse Team Plan。
  4. C 端定价不透明——官网未见完整价目表。
  5. 商用授权规则未公开——"商用交付就绪"是 Q3 的定位表述,但具体商用授权条款未见公开说明。

5.7. 六层能力矩阵

Vidu 的实现成熟度主要缺口
L1 上下文工程参考生视频(万物可参),最多 7 张主体图片无压缩机制;是否支持视频/音频参考未明确
L2 工具与执行6 特效 + 5 音效 + 首尾帧 + 漫画转动画 + 模板 + Vidu.API无 CLI/Skills
L3 编排与控制原生镜头控制(帧级导演指令)+ Smart Cut + Vidu Agent/Claw无节点图;编排不可导出
L4 记忆与状态主体库(角色/道具/场景)+ 多主体一致性(最多 7 张)最强容量偏小;无剧情状态机
L5 评估与观测SuperClue 多图/单图参考双榜第一;AA ELO 1220~1244无平台内建评估面板
L6 治理与安全SaaS/MaaS 生态分层;阿里云百炼集成无 AI 标识公开说明;无真人校验;无权限审计

6. 实际案例

6.1. 案例一:剧集工业化生产

  • 背景:AI 漫剧的核心诉求是"100 集像一部剧",即跨集的角色、道具、场景一致性。
  • 方案:使用 Vidu Q3 的参考生视频与主体库——把主要角色、关键道具、固定场景录入主体库,每次生成时一键选择参考主体;配合原生镜头控制做帧级导演,Smart Cut 做场景边界检测与转场。
  • 效果:Vidu Q3 在首个专项参考生能力评测 SuperClue 中以断层优势拿下多图参考与单图参考双榜第一,2026 年 1 月发布时登顶 Artificial Analysis 综合排名。具体的剧集项目交付数据未见公开披露,标 [待填写]
  • Harness 解读:这是本组唯一有"一致性专项第三方基准"背书的方案。对以连续生产为第一目标的团队,Vidu 的 L4 是当前最省事的选择——平台已经把资产库搭好了。

6.2. 案例二:企业级客户合作

  • 背景:视效与动漫领域的头部客户对一致性与交付标准要求最高。
  • 方案:Vidu 的 SaaS(Vidu Agent、Vidu Claw)+ MaaS(Vidu AI 开放平台、Vidu.API)分层,配合已登陆的阿里云百炼平台。
  • 效果:客户及伙伴包括索尼电影、腾讯动漫、阅文集团;用户与业务覆盖全球 200 多个国家和地区;2025 年实现用户和收入超 10 倍增长。具体项目细节未见公开披露,标 [待填写]
  • Harness 解读:阅文集团作为网文 IP 方、腾讯动漫作为动漫内容方,与 Vidu 的合作指向"小说/漫画 IP → 漫剧"的转化链路,与本组 05-AI-小说组场景有直接衔接。

6.3. 案例三:错峰模式下的成本优化

  • 背景:AI 漫剧的产能需求大,但多数团队现金流紧张;行业回本率不足 1.3%,成本敏感度极高。
  • 方案:Vidu 提供错峰模式无限积分——错峰时段无限次免费生成视频。
  • 效果:具体的错峰时段定义与并发限制未见公开披露,标 [待填写]
  • Harness 解读:错峰模式是典型的 L6 成本控制手段(削峰填谷),与本组中 PixVerse 的 Off-Peak / Preview 模式属同一思路。对非紧急的批量生成任务,把任务排入错峰队列可显著降低边际成本——这是工程上值得采用的调度策略。

7. 总结

7.1. 优势

  1. L4 形态最明确、最完整:主体库(角色/道具/场景)+ 多主体一致性(最多 7 张),是本组唯一把资产库四要素(载体、命名、引用、容量)全部显式化的商业平台。
  2. 一致性有专项第三方背书:SuperClue 多图参考与单图参考双榜第一——本组唯一针对"一致性"的第三方基准成绩。
  3. "为剧而生"的定位最精准:场景层能力矩阵明确列出"漫剧",特效层与音效层的结构化程度本组最高。
  4. 帧级镜头控制 + Smart Cut:精度与自动化兼具。
  5. 技术底座扎实:U-ViT 架构是全球首个 Diffusion 与 Transformer 融合架构;清华朱军团队背景。
  6. 生态分层清晰:SaaS(个人)+ MaaS(企业)+ 阿里云百炼,覆盖不同规模使用者。
  7. 资本与客户验证充分:B 轮近 20 亿元(阿里云领投),估值超 20 亿美元,客户含索尼电影、腾讯动漫、阅文集团,2025 年用户与收入超 10 倍增长。
  8. 错峰模式:对成本敏感团队友好。

7.2. 局限

  1. 参考容量偏小:最多 7 张主体图片,少于海螺的 12 个参考文件与 Seedance 的 9 图 + 3 视频 + 3 音频。多角色群戏可能受限。
  2. 无上下文压缩机制:对比海螺 H3-Context-IR 的 25:1 压缩,Vidu 走的是"明确容量 + 建模复用"路线,容量是硬上限。
  3. 无剧情状态机:主体库解决外观一致性,不解决剧情状态连续。
  4. C 端定价不透明:官网未见完整价目表,仅有第三方站报价。
  5. L6 信息不足:AI 标识、真人形象校验、团队协作权限、商用授权条款均无公开说明。
  6. 无 CLI/Skills:可编程性弱于 PixVerse。
  7. 单次时长 16 秒:与同组多数平台(15~16 秒)持平,长镜头需拼接。
  8. 最高分辨率 1080p:本组唯一未达 2K/4K 的商业平台(对比可灵 4K、PixVerse 4K、海螺 2K、Seedance 4K)。

7.3. 适用边界

适合不适合
100 集以上的连续剧集生产(L4 需求最高)需要 4K 分辨率输出的项目
多角色、多道具、多场景的复杂剧集单次需要大量(>7)参考素材的镜头
小说/漫画 IP 转漫剧(有阅文、腾讯动漫案例)需要把生成接进自有 Agent 环境的团队(无 CLI/Skills)
需要一致性第三方背书的商业交付需要团队协作权限与成本管控的组织
成本敏感、可接受错峰调度的团队需要明确合规能力证明的强监管场景
动态漫(漫画图转动画)需要长镜头单次生成的场景

7.4. 选型建议

  • 选 Vidu 的核心理由是 L4,不是画质或分辨率。若你的首要问题是"角色在第 1 集和第 47 集长得不一样",Vidu 的主体库是当前最直接的答案。
  • 把主体库当作你的主资产库,但不要当作唯一资产库:7 张主体图片的上限意味着你需要对资产做优先级管理——主要角色与固定场景进主体库,次要道具与外部参考由外部系统承载。
  • 务必自建剧情状态机:主体库只管外观。跨集的人物关系、伤势、时间线推进,需要在外部维护结构化状态,并在每次生成时以文本形式注入提示词。
  • 利用错峰模式做批量生成:把非紧急的批量镜头排入错峰队列,可显著降低边际成本。建议把生成任务分为"紧急人工精修"与"错峰批量"两条队列。
  • 分辨率需提前确认:Vidu 最高 1080p,若最终交付要求 2K/4K,需在后期做超分,或混合使用其他平台(如海螺的 In-Context Regeneration 2K)。
  • 合规能力需自行验证:无公开的 AI 标识与真人形象机制说明,国内平台分发的作品须自建符合 GB 45438—2025 的显式标识与 AIGC 元数据隐式标识管线。
  • C 端采购前先索取官方报价:官网未见完整价目表,第三方站报价($9.9/$29.9/$49.9/$99.9 积分包)不可作为预算依据。

信息缺口声明

  1. Vidu 官方 C 端定价:viduq3.net 为第三方站($9.9/$29.9/$49.9/$99.9 一次性积分包),vidu.cn 官网未见完整价目表。
  2. 参考素材是否支持视频与音频模态:官方明确"最多 7 张主体图片",是否支持视频/音频参考未见说明。
  3. 主体库的跨项目/跨团队复用规则:未见公开说明。
  4. 错峰模式的具体规则:错峰时段定义、并发限制、是否有生成质量差异,未见披露。
  5. Vidu 的 AI 生成合成内容标识落实方式:是否符合 GB 45438—2025 的显式标识与 AIGC 元数据隐式标识要求,无公开结果。
  6. Vidu 的真人形象校验机制:是否存在真人校验与真人人脸限制,无公开结果。
  7. Vidu 的团队协作与权限管理能力:是否存在类似 PixVerse Team Plan 的 RBAC、积分上限、用量分析,无公开结果。
  8. 商用授权的具体条款:Q3 定位"商用交付就绪",但授权范围、限制与责任划分未见公开说明。
  9. Vidu.API 的能力覆盖范围:网页端哪些能力可通过 API 调用,未见完整对照说明。
  10. 原生镜头控制的技术规格:"帧级导演指令"的具体语法、支持的控制维度(推拉摇移、景深、焦距等),未见公开说明。
  11. 单次时长 16 秒与 1080p 的数据来源:来自第三方 Atlas Cloud,,未在 Vidu 官网确认。
  12. 平台内建评估能力:是否存在可用率统计、回归集、轨迹追踪,无公开结果。

8. 参考资料

  1. Vidu 官网 — https://www.vidu.cn/
  2. 腾讯云开发者社区《Vidu Q3 参考生视频评测》 — https://cloud.tencent.com/developer/article/2655977
  3. Atlas Cloud《ShengShu Models》 — https://www.atlascloud.ai/providers/shengshu
  4. Atlas Cloud《Vidu Q3 AI Video Generator Now on Atlas Cloud》 — https://www.atlascloud.ai/it/blog/ai-updates/vidu-q3-ai-video-generator-now-on-atlas-cloud-create-16s-cinematic-with-native-audio-sync
  5. AI-Pirates《Vidu: AI Video with Reference Consistency》 — https://www.ai-pirates.com/en/glossar/vidu
  6. 中国经营报《生数科技完成 20 亿融资 前有巨头后有追兵如何突围?》 — https://cj.sina.cn/articles/view/1650111241/625ab30902001g6xy
  7. 科创板日报(转引自 moomoo)《估值超 120 亿 两个月连融 26 亿的生数科技 拟最快上半年启动 IPO》 — https://www.moomoo.com/hant/news/post/68185217
  8. 百度百科《AI漫剧》 — https://baike.baidu.com/item/AI%E6%BC%AB%E5%89%A7/68788906
  9. 百度百科《关键帧动画》 — https://baike.baidu.com/item/%E5%85%B3%E9%94%AE%E5%B8%A7%E5%8A%A8%E7%94%BB/10223838
  10. 360 百科《关键帧》 — https://baike.so.com/doc/6737995-32354145.html
  11. 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  12. 强制性国家标准 GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
  13. 今日头条《AI 漫剧的变现逻辑,可以总结为"一基三翼"》 — https://www.toutiao.com/article/7678962433607664163
  14. 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — https://www.thepaper.cn/newsDetail_forward_33993493
  15. 一品威客《AI 漫剧分镜设计指南》 — https://gonglue.epwk.com/322844.html

Vidu (ShengShu Technology) AI Comic Drama Platform Research

1. Introduction

Vidu was released by Beijing ShengShu Technology Co., Ltd. together with Tsinghua University, and is officially positioned as "China's first long-duration, high-consistency, high-dynamics video foundation model". Among the 8 objects in this group, Vidu is the commercial platform with the most explicit and complete L4 (memory and state) form: its "Subject Bank" saves characters, props, and scenes as persistent entries that can be referenced with one click, and multi-subject consistency supports up to 7 subject images.

This directly hits the group's core thesis — AI comic drama is the scenario with the greatest L4 pressure among the six Harness layers; whoever makes L4 solid can push AI comic drama from "being able to generate" to "being able to produce continuously." Through the "Reference-to-Video" paradigm, Vidu turns the consistency problem from a model-capability problem into an asset-engineering problem, making it the most representative L4 sample in this group.

1.1. Developer and Capital

ItemContent
CompanyBeijing ShengShu Technology Co., Ltd.
Founded2023
Founding teamTsinghua University Zhu Jun team
Technology foundationU-ViT architecture (proposed by the Zhu Jun team in September 2022; the world's first architecture fusing Diffusion with Transformer)
User coverageOver 200 countries and regions worldwide
Customers and partnersSony Pictures, Tencent Animation, China Literature (Yuewen)
2025 growthUsers and revenue grew more than 10x

Financing and capitalization:

TimeEvent
February 2026Completed a funding round of over RMB 600 million (led by Zhongguancun Science City and Xinglian Capital, with strategic investment from Wanxing Technology, VCG, and TRS, and follow-on investment from Qiming Venture Partners and others)
March 30, 2026Completed joint-stock reform
Nearly RMB 2 billion Series B funding round, led by Alibaba Cloud, with strategic investment from CSIC, Jiuan Haitang, TAL Education, and Guanghe Venture Capital; existing shareholders including Xinglian Capital, Data Capital, C&D Emerging Investment, BV Baidu Ventures, and Zhuoyuan Asia added more
Valuation exceeded USD 2 billion after funding; a Hong Kong IPO may be launched as early as the first half of 2026

Strategic roadmap: foundation world model → world generation model WGM (digital world) + world action model WAM (physical world); in December 2025, the world action model Motus was open-sourced.

1.2. Positioning

Vidu Q3 is positioned as "made for drama" — industrialized content production, full Reference-to-Video capability matrix, native audio-visual output, and commercial delivery ready.

This positioning is the closest in this group to the real needs of AI comic drama: it does not pursue the breadth of general creative video, but concentrates resources on the industrialized delivery of the single content category of "episodic drama." The three-layer capability matrix (VFX, sound effects, and scenes) also confirms this.

1.3. Pricing

Vidu officially claims pricing of "1/3 of the industry average."

Third-party API aggregator pricing (Atlas Cloud, [To be verified: third-party API aggregator pricing]), Vidu Q3 per-second prices:

Model / TaskList PriceDiscounted Price
Q3-Mix Reference-to-Video$0.125$0.106
Q3 Reference-to-Video$0.05$0.042
Q3-Pro First-Last Frame$0.05$0.042
Q3-Turbo Image-to-Video$0.04$0.034

C-end pricing gap: viduq3.net is a third-party site (one-time credit packs of $9.9/$29.9/$49.9/$99.9); the vidu.cn official site shows no complete price list. This group was unable to obtain Vidu's official C-end pricing page, marked [To be verified].

Other cost-related points: off-peak mode with unlimited credits (unlimited free video generation during off-peak hours).

1.4. Open Forms

FormDescription
SaaSVidu Agent, Vidu Claw ("hire AI creative staff"), for individuals / small studios
MaaSVidu AI Open Platform, Vidu.API, for enterprise batch production / developer integration
Cloud platformAvailable on the Alibaba Cloud Bailian platform
Real-time modelVidu S1 real-time interactive model

2. Glossary

TermEnglish / AbbreviationDefinition
AI 漫剧AI Comic DramaA content form between static comics and live-action short dramas, composed of comic storyboards plus dynamic audiovisual language
动态漫Motion ComicA slightly dynamic video form based on static comic material, produced through camera movement, zoom, localized motion effects, and voiceover
分镜 / 分镜脚本StoryboardTurning a written script into picture sketches, annotating each shot's composition, action, and duration
角色一致性Character ConsistencyThe ability to keep the same character's facial features, clothing, physique, and temperament stable across shots, episodes, and generations
关键帧KeyframeFrames (starting and ending points) that define key states of animation or camera changes, corresponding to "original drawings" in 2D animation
中间帧 / 过渡帧In-between / TweenTransition frames automatically generated between keyframes through interpolation algorithms
首尾帧First-Last FrameAn Image-to-Video control method in which a first frame and a last frame are uploaded and the model fills in the intermediate motion trajectory
图生视频Image-to-Video(I2V)Inputting a single static image and having the model generate several seconds of animation
口型同步 / 唇形同步Lip SyncOverlaying audio onto a generated character and driving the mouth movements to match the speech
镜头语言Camera LanguageAn audiovisual expression system that conveys narrative information through shot size, angle, movement, composition, and editing rhythm
参考生视频Reference-to-VideoVidu's core paradigm: upload 3 or more reference images, and the model models the characters/scenes/wardrobe-and-props as reusable subjects and keeps them consistent in the output
万物可参Everything ReferenceableVidu's description of Reference-to-Video: characters, scenes, and wardrobe-and-props can all serve as reference subjects
主体库Subject BankSaving characters, props, and scenes as persistent entries, with one-click selection of reference subjects
多主体一致性Multi-Subject ConsistencyThe ability to keep multiple subjects consistent in a single generation; Vidu supports up to 7 subject images
原生镜头控制Native Camera ControlPrecisely controlling camera movement with frame-level director commands
Smart CutSmart CutAutomatic scene-boundary detection and transitions
声画同出Native Audio-Visual OutputVideo and audio are produced simultaneously in the same inference
U-ViTU-ViTAn architecture proposed by the Zhu Jun team in September 2022; the world's first architecture fusing Diffusion with Transformer
Vidu Agent / Vidu ClawVidu Agent / Vidu ClawSaaS-layer products; Vidu Claw is positioned as "hire AI creative staff"
MaaSModel as a ServiceModel as a Service; Vidu's Vidu AI Open Platform and Vidu.API
SuperClueSuperClueThe first dedicated Reference-to-Video capability benchmark; Vidu Q3 ranked first on both the multi-image reference and single-image reference leaderboards
错峰模式Off-Peak ModeA mode with unlimited free video generation during off-peak hours
显式标识 / 隐式标识Explicit / Implicit LabelThe two statutory labeling types for AI-generated synthetic content: explicit labels are user-perceptible prompts; implicit labels are embedded in file metadata
AIGC 元数据字段AIGC Metadata FieldThe metadata implicit-labeling field mandated by mandatory national standard GB 45438—2025

3. Feature Description

3.1. Reference-to-Video

Reference-to-Video is Vidu's core paradigm, officially described as "Everything Referenceable" — extracting characters, scenes, and wardrobe-and-props as reusable modeling material. Upload 3 or more reference images, and the model models them as reusable subjects and keeps them consistent in the output.

Comparison with other objects in the group:

PlatformConsistency ParadigmReference Capacity
ViduReference-to-Video (Everything Referenceable)Up to 7 subject images
Hailuo AI@ reference system + Context-IRUp to 12 reference files (9 images + 3 videos + 3 audio)
PixVerseCharacter-consistent charactersNot disclosed [To be filled]
KlingElement referenceNot disclosed [To be filled]
Daydream AI (白日梦 AI)Character library (create a character from 5 reference images)Not disclosed
Doubao / SeedanceFour-modality referenceUp to 9 images + 3 videos + 3 audio

Vidu's characteristics are clear reference count and clear subject concept. The "subject" abstraction is key — it unifies characters, props, and scenes into the same class of referenceable objects, rather than building separate "character libraries" and "scene libraries."

3.2. Subject Bank and Multi-Subject Consistency

Subject Bank: saves characters, props, and scenes in the Subject Bank, with one-click selection of reference subjects.

Multi-subject consistency: upload up to 7 subject images and keep multiple subjects consistent within a single generation.

This is the only asset-library mechanism in this group that explicitly discloses a capacity limit (7 images), and the only one that explicitly covers three asset categories (characters, props, and scenes).

3.3. VFX and Sound-Effect Capability Matrix

Vidu Q3's three-layer capability matrix:

LayerCountDetails
VFX layer6 majorParticles, fluids, dynamics, camera movement, transitions, lighting
Sound-effect layer5 majorEnvironment, dynamics, atmosphere, foley, emotion
Scene layer4 majorShort drama, comic drama, film/TV drama, ads

The scene layer explicitly lists "comic drama", making Vidu the only platform in this group that treats comic drama as an independent scene category in its capability matrix.

3.4. Camera Control and Scene-Boundary Detection

  • Native camera control: frame-level director commands for precise control of camera movement.
  • Smart Cut: automatic scene-boundary detection and transitions.

Together, these two provide the bidirectional capability of "automatically splitting long material into shots" and "precisely controlling camera movement within a shot."

3.5. Other Capabilities and Templates

CapabilityDescription
Single-run duration16 seconds continuous 1080p @ 24fps [To be verified: third-party source]
Native audio-video syncNative audio-visual output
Animate comic imagesTurning comic images into animation, directly corresponding to motion-comic production
First-last frameSupported
Off-peak mode unlimited creditsUnlimited free generation during off-peak hours
Template libraryKiss, hug, everything blooms, AI costume change, etc.
Vidu S1Real-time interactive model
Vidu Claw"Hire AI creative staff"
Vidu AgentSaaS-layer Agent product

4. Platform Architecture

图 4-1|Vidu 平台六层架构:从 U-ViT 底座到主体库与商业化分层

Vidu 平台架构:自 U-ViT 底座至商业化分层的六层栈 示意:基于本文总体架构分析绘制 · 资产层为核心差异 商业化分层 SaaS:Vidu Agent · Vidu Claw(个人/小工作室) MaaS:Vidu AI 开放平台 · Vidu.API(企业/开发者)|云平台:阿里云百炼 编排与控制 原生镜头控制(帧级导演指令)· Smart Cut(场景边界检测与转场) 能力矩阵 特效层 6 · 音效层 5 · 场景层 4(含漫剧) 资产层 · 主体库(本图重点) 主体库(角色/道具/场景)· 多主体一致性(最多 7 张) 模型层 Q1(叙事逻辑)· Q2(AI 演技)· Q3(工业化生产)· S1(实时交互)· Motus(世界行动模型) 底座 U-ViT 架构:Diffusion + Transformer 融合(全球首个融合架构) 结构解读:U-ViT 底座支撑 Q 系模型栈,主体库(L4)锚定跨集一致性——这是 Vidu 的核心壁垒。

数据来源:基于本文分析绘制的示意图。

4.1. Overall Architecture

┌────────────────────────────────────────────────────────────────┐
│  商业化分层  SaaS:Vidu Agent · Vidu Claw(个人/小工作室)         │
│              MaaS:Vidu AI 开放平台 · Vidu.API(企业/开发者)      │
│              云平台:阿里云百炼                                    │
├────────────────────────────────────────────────────────────────┤
│  编排与控制  原生镜头控制(帧级导演指令)                           │
│              Smart Cut(场景边界检测与转场)                       │
├────────────────────────────────────────────────────────────────┤
│  能力矩阵  特效层 6 · 音效层 5 · 场景层 4(含漫剧)                │
├────────────────────────────────────────────────────────────────┤
│  资产层  主体库(角色/道具/场景)· 多主体一致性(最多 7 张)         │
├────────────────────────────────────────────────────────────────┤
│  模型层  Vidu Q1(叙事逻辑)· Q2(AI 演技)· Q3(工业化生产)       │
│          Vidu S1(实时交互)· Motus(世界行动模型,2025-12 开源)   │
├────────────────────────────────────────────────────────────────┤
│  底座  U-ViT 架构(Diffusion + Transformer 融合)                 │
└────────────────────────────────────────────────────────────────┘

4.2. Model Layer

ModelPositioning
Vidu Q1Redefining narrative logic, strengthening basic generation capability and the storyline-advancement framework
Vidu Q2Unlocking AI acting, giving virtual characters micro-expressions and physical expressiveness
Vidu Q3Industrialized content production, full Reference-to-Video capability matrix, native audio-visual output, commercial delivery ready; "made for drama"
Vidu S1Real-time interactive model
MotusWorld Action Model (WAM), open-sourced in December 2025

The Q1→Q2→Q3 evolution path is worth noting: first solve "whether the story can advance" (narrative logic), then "whether characters can perform" (micro-expressions and body language), and finally "whether it can be delivered industrially" (Reference-to-Video capability matrix + native audio-visual output + commercial readiness). This is a clear route from content to engineering.

4.3. Asset Layer

The Subject Bank is the core differentiator of this architecture; see the L4 analysis in Section 5.4.

4.4. Commercialization Layers

SaaS layer: Vidu Agent, Vidu Claw, for individuals / small studios. Vidu Claw is positioned as "hire AI creative staff."

MaaS layer: Vidu AI Open Platform, Vidu.API, for enterprise batch production and developer integration. It is available on the Alibaba Cloud Bailian platform.

This layering lets the same set of model capabilities be used directly by creators and also integrated by enterprises into their own production systems.


5. Harness Design

5.1. L1 Context Engineering Layer

Vidu's L1 mechanism is Reference-to-Video (Everything Referenceable): extracting characters, scenes, and wardrobe-and-props as reusable modeling material and injecting them into the context in the form of reference images.

Compared with other objects in the group, Vidu's L1 has two characteristics:

  1. Reference-as-modeling: the official wording emphasizes "modeled as a reusable subject," meaning the reference image is not a simple conditional input but is modeled into a subject representation. This goes a step beyond "pasting an image on."
  2. Clear capacity: up to 7 subject images. Although fewer than Seedance (9 images + 3 videos + 3 audio) and Hailuo (12 reference files), the explicitly publicized limit itself is engineering-friendly — users can plan deterministically.

Gaps:

  1. No public context-compression mechanism (compared with Hailuo H3-Context-IR's 25:1 compression).
  2. Whether references support video and audio modalities is not clearly explained (compared with Seedance and Hailuo, which explicitly support them), marked [To be filled].
  3. No public context-priority sorting or cache-reuse mechanism.

5.2. L2 Tools and Execution Layer

Vidu's L2 is organized as a capability matrix, the most structured tool layer in this group:

CategoryTools
VFX (6)Particles, fluids, dynamics, camera movement, transitions, lighting
Sound effects (5)Environment, dynamics, atmosphere, foley, emotion
Generation methodsReference-to-Video, Image-to-Video, first-last frame, animate comic images
TemplatesKiss, hug, everything blooms, AI costume change, etc.
APIVidu.API (MaaS)

Assessment: Vidu's L2 ranks "strong." The 5-major sound-effect layer is unique in this group (other platforms mostly do post-production dubbing or native audio without sound-effect categorization). The 6-major VFX layer covers the most common visual needs of comic drama.

Programmability: provides Vidu.API (MaaS) and Alibaba Cloud Bailian integration, with clear programmability. However, no CLI or Skills-type tools were found (compared with PixVerse), marked [To be filled].

5.3. L3 Orchestration and Control Layer

Vidu has three clear threads at L3:

  1. Native camera control: frame-level director commands. This is the finest-grained camera control in this group — not a coarse-grained choice of "wide / medium / close-up" but precise frame-level specification.
  2. Smart Cut: automatic scene-boundary detection and transitions. This solves the problem of automatically splitting long material into shots.
  3. Vidu Agent / Vidu Claw: Agent-like orchestration at the SaaS layer; Vidu Claw is positioned as "hire AI creative staff."

Assessment: Vidu's L3 ranks "strong." Its orchestration approach differs from both Kling (storyboard commands) and Hailuo (timeline) — it is a combination of frame-level precise control + automatic boundary detection, offering both precision and automation.

Gaps: no editable node graph (compared with PixVerse Canvas, ComfyUI); no public conditional-branch and batch-processing mechanisms; whether orchestration results can be exported and versioned is not explained, marked [To be filled].

5.4. L4 Memory and State Layer

Vidu's L4 is the strongest among the commercial platforms in this group, and the best illustration of the group's core thesis.

MechanismDescriptionComparative Advantage
Subject BankSaves characters, props, and scenes as persistent entries, with one-click selection of reference subjectsThe only asset library that explicitly covers three asset categories
Multi-subject consistencyUp to 7 subject images, keeping multiple subjects consistent within a single generationThe only one that explicitly discloses a capacity limit
Reference-to-VideoModels characters/scenes/wardrobe-and-props as reusable materialThe "subject" abstraction unifies the three asset categories

Why this matters: the cross-shot, cross-episode consistency of AI comic drama is essentially a cross-session state-persistence problem. Vidu's Subject Bank productizes this — users do not need to build their own asset-management system; the platform provides:

  • State carrier (subject entries)
  • State naming (three categories: characters / props / scenes)
  • State reference (one-click selection)
  • State capacity (up to 7)

This is the only commercial platform in this group that makes all four L4 elements explicit.

Remaining gaps:

  1. Whether the 7-image capacity limit is enough: for multi-character ensemble scenes, 7 subject images (including characters, props, and scenes) may be tight. Compared with Hailuo's 12 reference files and Seedance's 9 images + 3 videos + 3 audio, Vidu's capacity is on the small side.
  2. No plot state machine: the Subject Bank solves "looking the same" but not "plot-state continuity." Cross-episode character relationships, injuries, and timeline progression still need external maintenance.
  3. Cross-project reuse rules not disclosed: whether the Subject Bank can be reused across projects and teams is not explained, marked [To be filled].

5.5. L5 Evaluation and Observation Layer

Vidu's L5 has the most targeted evidence in this group of commercial platforms.

Benchmark / MetricResult
SuperClue (first dedicated Reference-to-Video capability benchmark)Ranked first on both the multi-image and single-image reference leaderboards (decisive advantage)
Artificial Analysis overall rankingTopped it at the January 2026 release
Artificial Analysis Video Arena ELO1220~1244, 2nd worldwide (1st is Sora 2 at ~1250+) [To be verified: third-party source]

The SuperClue result is especially critical: it is the first dedicated Reference-to-Video capability benchmark, and "reference-to-video" is precisely the core technical path for AI comic drama consistency. Vidu's double first-place finish on this benchmark shows third-party endorsement of its L4 capability — it is the only platform in this group with a third-party benchmark specifically targeting "consistency."

Gaps: no platform-built-in availability-statistics dashboard, regression set, or trajectory tracking, marked [To be filled].

5.6. L6 Governance and Safety Layer

Vidu's L6 mechanism is ecosystem layering: SaaS (individuals / small studios) and MaaS (enterprise batch production / developer integration).

LayerProductsAudienceGovernance Implication
SaaSVidu Agent, Vidu ClawIndividuals / small studiosLightweight, low barrier to entry
MaaSVidu AI Open Platform, Vidu.APIEnterprises / developersIntegratable, batch-capable
Cloud platformAlibaba Cloud BailianEnterprisesRelies on the cloud vendor's compliance system

Assessment: Vidu's L6 ranks "medium." The ecosystem layering itself provides a division of usage boundaries, but explicit permission, quota, and audit mechanisms are lacking (compared with PixVerse Team Plan's RBAC + credit cap + usage analytics + unified billing).

Gaps:

  1. No public explanation of AI-generated synthetic-content labeling found — whether it meets the explicit-labeling and AIGC metadata implicit-labeling requirements of GB 45438—2025 has no public result, marked [To be filled].
  2. No real-person image verification mechanism found — compared with Doubao/Seedance, Vidu has no public information on this red line.
  3. No public team-collaboration and permission-management capability — compared with PixVerse Team Plan.
  4. C-end pricing is opaque — no complete price list on the official site.
  5. Commercial-licensing rules not disclosed — "commercial delivery ready" is Q3's positioning statement, but the specific commercial-licensing terms have no public explanation.

5.7. Six-Layer Capability Matrix

LayerVidu's ImplementationMaturityMain Gaps
L1 Context EngineeringReference-to-Video (Everything Referenceable), up to 7 subject imagesStrongNo compression mechanism; whether video/audio references are supported is unclear
L2 Tools & Execution6 VFX + 5 sound effects + first-last frame + comic-to-animation + templates + Vidu.APIStrongNo CLI/Skills
L3 Orchestration & ControlNative camera control (frame-level director commands) + Smart Cut + Vidu Agent/ClawStrongNo node graph; orchestration cannot be exported
L4 Memory & StateSubject Bank (characters/props/scenes) + multi-subject consistency (up to 7)StrongestCapacity on the small side; no plot state machine
L5 Evaluation & ObservationSuperClue first on both multi-image/single-image reference; AA ELO 1220~1244StrongNo platform-built-in evaluation dashboard
L6 Governance & SafetySaaS/MaaS ecosystem layering; Alibaba Cloud Bailian integrationMediumNo public AI-labeling explanation; no real-person verification; no permission auditing

6. Case Studies

6.1. Case 1: Industrialized Episodic Production

  • Background: the core requirement of AI comic drama is "100 episodes feel like one drama" — that is, cross-episode consistency of characters, props, and scenes.
  • Approach: use Vidu Q3's Reference-to-Video and Subject Bank — register main characters, key props, and fixed scenes in the Subject Bank, and one-click select reference subjects for each generation; pair with native camera control for frame-level directing and Smart Cut for scene-boundary detection and transitions.
  • Results: Vidu Q3 won first place on both the multi-image and single-image reference leaderboards with a decisive advantage in SuperClue, the first dedicated Reference-to-Video capability benchmark, and topped the Artificial Analysis overall ranking at its January 2026 release. Specific episodic-project delivery data has not been publicly disclosed, marked [To be filled].
  • Harness interpretation: this is the only solution in this group backed by a "consistency-specific third-party benchmark." For teams whose first priority is continuous production, Vidu's L4 is currently the most effort-saving choice — the platform has already built the asset library.

6.2. Case 2: Enterprise-Grade Customer Collaboration

  • Background: top customers in the visual-effects and animation field have the highest requirements for consistency and delivery standards.
  • Approach: Vidu's SaaS (Vidu Agent, Vidu Claw) + MaaS (Vidu AI Open Platform, Vidu.API) layering, combined with the already-available Alibaba Cloud Bailian platform.
  • Results: customers and partners include Sony Pictures, Tencent Animation, and China Literature (Yuewen); users and business cover over 200 countries and regions worldwide; in 2025 users and revenue grew more than 10x. Specific project details have not been publicly disclosed, marked [To be filled].
  • Harness interpretation: with China Literature as an online-novel IP party and Tencent Animation as an animation content party, the collaboration with Vidu points to the "novel/comic IP → comic drama" conversion chain, directly connecting to this group's 05-AI-Novel scenario.

6.3. Case 3: Cost Optimization in Off-Peak Mode

  • Background: AI comic drama has large production-demand, but most teams face tight cash flow; the industry payback rate is below 1.3%, so cost sensitivity is extremely high.
  • Approach: Vidu provides off-peak mode with unlimited credits — unlimited free video generation during off-peak hours.
  • Results: the specific definition of off-peak hours and concurrency limits have not been publicly disclosed, marked [To be filled].
  • Harness interpretation: off-peak mode is a typical L6 cost-control measure (peak shaving and valley filling), following the same approach as PixVerse's Off-Peak / Preview modes in this group. For non-urgent batch generation tasks, queueing tasks into the off-peak queue can significantly reduce marginal cost — a scheduling strategy worth adopting from an engineering standpoint.

7. Summary

7.1. Strengths

  1. Most explicit and complete L4 form: Subject Bank (characters/props/scenes) + multi-subject consistency (up to 7), and the only commercial platform in this group that makes all four asset-library elements (carrier, naming, reference, capacity) explicit.
  2. Consistency backed by a dedicated third-party benchmark: first on both the multi-image and single-image reference leaderboards in SuperClue — the only third-party benchmark result in this group aimed specifically at "consistency."
  3. Most precise "made for drama" positioning: the scene-layer capability matrix explicitly lists "comic drama," and the VFX and sound-effect layers are the most structured in this group.
  4. Frame-level camera control + Smart Cut: offers both precision and automation.
  5. Solid technology foundation: the U-ViT architecture is the world's first architecture fusing Diffusion with Transformer; backed by the Tsinghua Zhu Jun team.
  6. Clear ecosystem layering: SaaS (individuals) + MaaS (enterprises) + Alibaba Cloud Bailian, covering users of different scales.
  7. Strong capital and customer validation: nearly RMB 2 billion Series B (led by Alibaba Cloud), valuation over USD 2 billion, customers including Sony Pictures, Tencent Animation, and China Literature, with 2025 users and revenue growing more than 10x.
  8. Off-peak mode: friendly to cost-sensitive teams.

7.2. Limitations

  1. Reference capacity on the small side: up to 7 subject images, fewer than Hailuo's 12 reference files and Seedance's 9 images + 3 videos + 3 audio. Multi-character ensemble scenes may be constrained.
  2. No context-compression mechanism: compared with Hailuo H3-Context-IR's 25:1 compression, Vidu follows the "clear capacity + modeling reuse" route, so capacity is a hard limit.
  3. No plot state machine: the Subject Bank solves appearance consistency, not plot-state continuity.
  4. C-end pricing is opaque: no complete price list on the official site, only third-party-site quotes.
  5. Insufficient L6 information: AI labeling, real-person image verification, team collaboration permissions, and commercial-licensing terms all lack public explanation.
  6. No CLI/Skills: programmability is weaker than PixVerse.
  7. Single-run duration of 16 seconds: on par with most platforms in this group (15~16 seconds); long shots require stitching.
  8. Maximum resolution 1080p: the only commercial platform in this group not reaching 2K/4K (compared with Kling 4K, PixVerse 4K, Hailuo 2K, Seedance 4K).

7.3. Applicability Boundary

Suitable ForNot Suitable For
Continuous episodic production of 100+ episodes (highest L4 demand)Projects requiring 4K resolution output
Complex episodes with many characters, props, and scenesShots needing a large number of reference materials in a single run (>7)
Novel/comic IP to comic drama (with China Literature and Tencent Animation cases)Teams that need to integrate generation into their own Agent environments (no CLI/Skills)
Commercial delivery requiring third-party consistency endorsementOrganizations needing team-collaboration permissions and cost controls
Cost-sensitive teams that can accept off-peak schedulingHeavily regulated scenarios requiring explicit compliance-proof
Motion comics (comic image to animation)Scenarios requiring single-run long-shot generation

7.4. Selection Recommendations

  • The core reason to choose Vidu is L4, not image quality or resolution. If your primary problem is "characters look different in episode 1 and episode 47," Vidu's Subject Bank is currently the most direct answer.
  • Treat the Subject Bank as your main asset library, but not your only one: the 7-subject-image limit means you need to prioritize your assets — main characters and fixed scenes go into the Subject Bank, while secondary props and external references are carried by external systems.
  • Be sure to build your own plot state machine: the Subject Bank only handles appearance. Cross-episode character relationships, injuries, and timeline progression need to be maintained externally as structured state and injected into the prompt as text at each generation.
  • Use off-peak mode for batch generation: queueing non-urgent batch shots into the off-peak queue can significantly reduce marginal cost. We recommend dividing generation tasks into two queues: "urgent manual refinement" and "off-peak batch."
  • Confirm resolution in advance: Vidu maxes out at 1080p; if final delivery requires 2K/4K, you need to do upscaling in post, or mix in other platforms (e.g., Hailuo's In-Context Regeneration 2K).
  • Verify compliance capabilities yourself: there is no public explanation of AI-labeling and real-person-image mechanisms; works distributed on domestic platforms must build their own pipeline meeting GB 45438—2025's explicit-labeling and AIGC metadata implicit-labeling requirements.
  • Request an official quote before C-end purchasing: the official site has no complete price list, so third-party-site quotes ($9.9/$29.9/$49.9/$99.9 credit packs) cannot be used as a budgeting basis.

Information-Gap Statement

  1. Vidu's official C-end pricing: viduq3.net is a third-party site (one-time credit packs of $9.9/$29.9/$49.9/$99.9); the vidu.cn official site shows no complete price list.
  2. Whether reference materials support video and audio modalities: officially "up to 7 subject images," whether video/audio references are supported is not explained.
  3. Cross-project / cross-team reuse rules for the Subject Bank: no public explanation.
  4. Specific rules of off-peak mode: the definition of off-peak hours, concurrency limits, and whether there is a generation-quality difference have not been disclosed.
  5. How Vidu implements AI-generated synthetic-content labeling: whether it meets the explicit-labeling and AIGC metadata implicit-labeling requirements of GB 45438—2025 has no public result.
  6. Vidu's real-person image verification mechanism: whether real-person verification and real-face restrictions exist has no public result.
  7. Vidu's team-collaboration and permission-management capability: whether RBAC, credit caps, and usage analytics similar to PixVerse Team Plan exist has no public result.
  8. Specific commercial-licensing terms: Q3 is positioned as "commercial delivery ready," but the licensing scope, restrictions, and responsibility division have no public explanation.
  9. Vidu.API's capability coverage: which web-side capabilities can be called through the API has no complete comparison explanation.
  10. Technical specifications of native camera control: the specific syntax of "frame-level director commands" and the supported control dimensions (push-pull, par-away-tilt, pan-tilt, depth-of-field, focal length, etc.) have no public explanation.
  11. Data source for the 16-second single-run duration and 1080p: from a third party, Atlas Cloud, not confirmed on the Vidu official site.
  12. Platform-built-in evaluation capability: whether availability statistics, regression sets, and trajectory tracking exist has no public result.

8. References

  1. Vidu official site — https://www.vidu.cn/
  2. Tencent Cloud Developer Community "Vidu Q3 Reference-to-Video Review" — https://cloud.tencent.com/developer/article/2655977
  3. Atlas Cloud "ShengShu Models" — https://www.atlascloud.ai/providers/shengshu
  4. Atlas Cloud "Vidu Q3 AI Video Generator Now on Atlas Cloud" — https://www.atlascloud.ai/it/blog/ai-updates/vidu-q3-ai-video-generator-now-on-atlas-cloud-create-16s-cinematic-with-native-audio-sync
  5. AI-Pirates "Vidu: AI Video with Reference Consistency" — https://www.ai-pirates.com/en/glossar/vidu
  6. China Business Journal "ShengShu Technology Completes RMB 2 Billion Funding; with Giants Ahead and Pursuers Behind, How Will It Break Through?" — https://cj.sina.cn/articles/view/1650111241/625ab30902001g6xy
  7. Science and Technology Innovation Board Daily (reprinted via moomoo) "ShengShu Technology, Valued at Over RMB 12 Billion and Raising RMB 2.6 Billion in Two Months, Plans IPO Possibly in First Half of the Year" — https://www.moomoo.com/hant/news/post/68185217
  8. Baidu Baike "AI Comic Drama" — https://baike.baidu.com/item/AI%E6%BC%AB%E5%89%A7/68788906
  9. Baidu Baike "Keyframe Animation" — https://baike.baidu.com/item/%E5%85%B3%E9%94%AE%E5%B8%A7%E5%8A%A8%E7%94%BB/10223838
  10. 360 Baike "Keyframe" — https://baike.so.com/doc/6737995-32354145.html
  11. Measures for Labeling AI-Generated Synthetic Content (issued by the Cyberspace Administration of China and three other departments, CAC Decree No. 2 of 2025) — https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  12. Mandatory national standard GB 45438—2025 "Network Security Technology — Labeling Methods for AI-Generated Synthetic Content" — https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
  13. Toutiao "The Monetization Logic of AI Comic Drama Can Be Summarized as 'One Foundation, Three Wings'" — https://www.toutiao.com/article/7678962433607664163
  14. The Paper "RMB 1,914 to Produce One Episode? Comic Drama Is Still Trapped in Hidden Costs" — https://www.thepaper.cn/newsDetail_forward_33993493
  15. Epwk.com "AI Comic Drama Storyboard Design Guide" — https://gonglue.epwk.com/322844.html