AI 网剧(AI Drama)中的 AI Harness
1. 介绍
1.1 背景
AI 网剧(AI 微短剧)是创意产业组中规模化最快、规范建设最滞后的方向。据中国网络视听协会《微短剧创作指引》(引自人民日报,2026),2026 年一季度全行业上线微短剧约 12.8 万部,其中 AI 微短剧约 12.2 万部,占比超 95%。同期研究机构数据显示,抖音每月 TOP5000 短剧中全 AI 生成微短剧从 2025 年 1 月的 4 部增长到 10 月的 69 部、11 月的 217 部。
这个渗透率带来的直接后果是:质量的方差被放大到极致。成本结构上,60 集、每集 2 分钟的中端真人短剧约需 50~60 万元,同规格 AI 短剧每集成本约 2000~3000 元、整部不到 20 万元(从业者访谈口径,引自人民日报,2026)。成本下降一个数量级的另一面是,大量作品在"故事连贯性"和"角色一致性"上不达标。
21 世纪经济报道(2025-10-23)对行业问题的归纳可作为本方向的"常见失败"基线:
- 内容层:尚未诞生现象级作品,多数依赖平台托底;过度追求画面华丽与动作流畅而忽视故事连贯性与角色塑造,沦为"幻灯片堆砌";严重同质化,扎堆玄幻、科幻题材。
- 技术层:情感表达与人物塑造能力不足;人物与场景一致性问题频发。
- 商业层:整体盈利能力未充分验证,多数依赖平台补贴与流量扶持;IP 属性不足,用户留存率低。
- 版权与伦理:部分剧组用 AI 一键生成剧本、换脸明星,甚至售卖"定制亲密戏"。
1.2 定义与范围
本方向的 AI Harness 指:在 AI 微短剧 / 网剧制作场景下,承接剧本 → 分镜 → 生成 → 剪辑 → 配音五段流水线,把分散的视频生成能力组织为可控、可复现、可回归验证的成片生产系统的工程化承载层。
| 流水线阶段 | 典型决策 | 输出工件 | 一致性风险 |
|---|---|---|---|
| 剧本 | 题材、结构、爽点密度、人物小传 | 剧本、分集大纲、人物小传 | 人设前后矛盾 |
| 分镜 | 镜头序列、景别、运镜 | 分镜脚本、镜头表 | 镜头语言不连贯 |
| 生成 | 角色形象、场景、表演、时长 | 镜头素材(多候选) | 角色换脸、场景漂移 |
| 剪辑 | 节奏、转场、音画同步 | 粗剪版本 | 帧间不一致 |
| 配音 | 音色、情绪、BGM、音效 | 成片音轨 | 音色跨集漂移 |
1.3 在 AI Harness 体系中的定位
图 1-1|AI 网剧 Harness:六层体系与五段流水线
数据来源:基于本文分析绘制的示意图。
主导层:L3 编排与控制层 + L4 记忆与状态层。瓶颈层:L4 记忆与状态层。
| 层 | 在 AI 网剧方向的体现 | 关键约束 |
|---|---|---|
| L1 上下文工程 | 剧本、人物小传、角色标准三视图、分镜脚本作为锚定上下文 | 三视图与色卡是角色一致性的物理锚点 |
| L2 工具与执行 | 剧本模型、角色生成、分镜生成、视频生成、TTS、剪辑、BGM 各阶段工具 | 按镜头类型动态调用不同参数量级模型 |
| L3 编排与控制 | 剧本 → 分镜 → 生成 → 剪辑 → 配音 五段流水线;分段生成 + 一致性校验 + 重生成 | 重生成回路必须内置,不能靠人工发现后再补 |
| L4 记忆与状态 | 角色资产库(三视图 + 表情集 + 色卡 + 音色)与镜头状态机;跨集连续性 | 本方向的质量瓶颈 |
| L5 评估与观测 | 角色一致性(面部/服饰/道具)、风格稳定性、叙事连贯性、成本与周期 | 四项缺一不可 |
| L6 治理与安全 | 训练数据版权合规机制;前置审核;显著标识;不得虚假标注主创信息;肖像权与名誉权 | 前置审核是本方向的特有强约束 |
本方向的 L4 瓶颈表现得最为典型。原因有三:
- 跨越的时间跨度最长:一部 60 集短剧的制作周期以周计,模型会话不可能保持那么久的状态。
- 一致性的维度最多:面部、服饰、道具、场景、音色五个维度都要一致,任何一个漂移都会出戏。
- 重生成的代价高:发现第 40 集角色漂移时,重生成意味着重新走一遍流水线。
因此,本方向 L4 的核心结构不是"记忆",而是镜头状态机:以镜头为单位记录生成状态、版本、校验结果与失败原因,使任何一个镜头都可以被单独重生成而不影响其他镜头。
1.4 产业现状与已公开的量化口径
| 指标 | 数值 | 来源与口径 |
|---|---|---|
| 2026 Q1 微短剧上线量 | 约 12.8 万部 | 中国网络视听协会《微短剧创作指引》,引自人民日报,2026 |
| 其中 AI 微短剧 | 约 12.2 万部(占比 > 95%) | 同上 |
| 抖音 TOP5000 中全 AI 微短剧 | 2025-01: 4 部 → 2025-10: 69 部 → 2025-11: 217 部 | 研究机构数据,引自人民日报,2025 |
| 中端真人短剧成本(60 集 × 2 分钟) | 约 50~60 万元 | 从业者访谈,引自人民日报,2026 |
| 同规格 AI 短剧成本 | 每集 2000~3000 元,整部 < 20 万元 | 同上 |
| 《白狐》制作周期 | 3 个月 → 2 周 | 21 世纪经济报道,2025 |
| 《白狐》每分钟成本 | 数万元 → 万元级(另一口径为"万元以内") | 同上 |
| 《山海奇镜之劈波斩浪》周期 / 成本 | 3~6 个月 → 2 个月;成本为传统制作的 1/4 | 同上 |
| 《奶团太后宫心计》(68 集) | 累计播放 2.1 亿 | 公开报道,2025 |
| 《兴安岭诡事》 | 上线不到 21 小时播放破千万;截至 2025-10-17 累计 5613.3 万;抖音原生端收益超 30 万元;涨粉 10 万+ | 公开报道,2025 |
视频生成模型规格(券商研报汇总,海通,2025-10-29):
| 模型 | 规格 |
|---|---|
| Sora 2 | App 默认 10 秒竖屏;标准档 720×1280 / 1280×720 |
| 可灵 Kling | 官方称 1080p、30 fps、最长 2 分钟 |
| Veo 3.1 | API 预览 4 / 6 / 8 秒;Flow"场景续写"最长约 1 分钟 |
| 即梦 JiMeng | 面向图片/视频一体"智能画布";规格未统一公开 |
规格差异直接决定了编排策略:单次生成长度受限,必然要求"分段生成 + 一致性校验 + 重生成"的流水线结构,这是本方向 L3 编排层设计的物理前提。
2. 名词解释
| 术语 | 英文/缩写 | 释义 |
|---|---|---|
| AI 微短剧 | AI Micro-drama | 主要或全部由 AI 生成的单集时长极短的剧集形态 |
| 角色三视图 | Character Turnaround | 角色正面、侧面、背面三个标准视角的参考图,是一致性的物理锚点 |
| 角色资产库 | Character Asset Library | 承载角色三视图、表情集、色卡、音色与属性标签的可复用资产库 |
| 镜头状态机 | Shot State Machine | 以镜头为单位记录生成状态、版本、校验结果与失败原因的状态机 |
| 人物小传 | Character Profile | 角色的身份、性格、动机与关系的结构化描述,供后续角色设计使用 |
| 分镜生成 | Storyboard Generation | 由剧本自动生成镜头序列、景别与运镜的过程 |
| 爽点 | Hook / Satisfaction Point | 短剧中制造情绪释放的情节节点,是剧本模型优化的重点 |
| 情感音色库 | Emotional Voice Bank | 按情绪分类的角色音色集合,用于角色与场景音频精准匹配 |
| 中间帧生成 | Motion In-Betweening, MIB | 在关键帧之间生成过渡帧,消除滑步与抖动 |
| 智能动捕 | AI Motion Capture | 用任意摄像头或视频驱动 AI 演员表演的技术 |
| 跨集连续性 | Cross-episode Continuity | 角色、场景与剧情在多集之间保持一致的能力 |
| 前置审核 | Pre-release Review | 内容上线前完成的审核环节,AI 短剧被明确要求建立此机制 |
| 显著标识 | Conspicuous Label | 在显著位置添加的生成合成内容提示标识,AI 短剧的强制要求 |
| 虚假标注主创信息 | False Credit Attribution | 不实标注导演、编剧、演员等主创信息的行为,行业规范明令禁止 |
| 训练数据版权合规机制 | Training Data Copyright Compliance | 开发运营方需建立的训练数据授权与合规管理制度 |
| 分类分层审核 | Tiered Content Review | 按投资额与题材分级的内容审核制度 |
| "AI 魔改" | AI Mashup / Remix | 用 AI 对既有影视内容进行歪曲性改编的内容形态,已被纳入专项治理 |
3. 案例
3.1 昆仑万维 SkyReels:AI 短剧全流水线平台
3.1.1 背景
2024 年 8 月,昆仑万维推出 SkyReels,定位为全球首个集成视频大模型与 3D 大模型的 AI 短剧创作平台。其要解决的核心问题是:AI 短剧制作涉及剧本、角色、分镜、视频、音频、剪辑六类能力,分属不同工具,中间的状态传递全靠人工,导致一致性在每次交接时衰减。
3.1.2 方案
SkyReels 构建了一条完整的流水线,其结构本身就是一份 AI 网剧 Harness 的参考实现:
| 模块 | 能力 | 对应 Harness 层 |
|---|---|---|
| SkyScript 剧本大模型 | 快速创作结构完整、情节丰富的短剧剧本;优化的"爽点"生成能力,人工评级稳定达到 A/S 级;新增海量爆款创意模板;可摘要出人物小传供后续角色设计 | L1 → L3 |
| 角色库 | 引入写实 AI 演员并配备细粒度属性标签,可智能匹配剧本人设;支持一键生成角色形象与配音(情感音色库) | L4 |
| StoryboardGen 分镜生成大模型 | 利用分镜师专业经验和上下文增强技术,提升分镜连贯性与人物一致性 | L3 |
| Sky3DGen | 生成多样化 3D 元素与场景,具备实时 3D 场景交互能力 | L2 |
| WorldEngine | AI 3D 引擎与视频大模型深度融合,解决传统视频创作中物理现象不协调问题 | L2 |
| 视频生成 | 1080P / 60 fps,单次生成长度达 180 秒(另有口径称"已实现 60 秒以上视频生成");结合不同镜头类型动态调用不同参数量级模型 | L2 |
| 音频 | 构建情感音色库与短剧 BGM 库,实现角色与场景音频精准匹配 | L2 + L4 |
| Web 端 | 3D 交互编辑(分镜画面与人物实时调整);AI 动捕(任意摄像头或视频驱动 AI 演员表演);集成百万级影视动作数据库并开放自定义 | L2 |
| 开源模型 SkyReels-A1 | 给定输入视频序列与参考人像,提取面部表情感知特征点作为运动描述符迁移到人像,基于 DiT 条件视频生成框架 | L2 |
从 Harness 视角看,SkyReels 最关键的设计是角色库与分镜生成的耦合:分镜生成直接引用角色库中的资产,而不是每次重新描述角色。这正是把 L4 从"记忆"变成"状态载体"的工程体现。
3.1.3 效果
- SkyScript 生成的剧本人工评级稳定达到 A/S 级。
- 视频生成规格达到 1080P / 60 fps / 单次 180 秒。
- 昆仑万维 SkyReels 平台 + DramaWave 分发平台,截至 2025 年一季度月流水达 1000 万美元。
来源:昆仑万维官网;证券时报《昆仑万维 2024 年年度报告摘要》,2025。月流水数据为企业披露口径。
3.2 《白狐》与《山海奇镜》:周期与成本的压缩
3.2.1 背景
传统微短剧的周期与成本结构由三部分决定:剧组人力、场地与拍摄、后期。对于玄幻、科幻、古装等题材,场地与后期占比尤其高——一位导演曾表示,某部戏 90% 的戏份发生在兴安岭,"如果要去林海雪原实拍,造价非常高,危险系数很大"。
3.2.2 方案
- 《白狐》(全 AI 制作微短剧):4 人团队,用 ChatGPT 完成剧本快速迭代 + AI 绘图 + 智能剪辑。
- 《山海奇镜之劈波斩浪》(快手,导演陈坤,2024):采用 AI 全流程制作。
- 《兴安岭诡事》(国内首部付费 AI 短剧,导演丁宽):用 AI 替代"最烦琐、烧钱以及最危险的工作"。
从 Harness 视角拆解:
- L2:极小团队(4 人)意味着工具链必须高度集成,任何需要人工搬运的中间格式都会成为瓶颈。
- L3:剧本 → 分镜 → 生成 → 剪辑 → 配音 的流水线在 4 人团队中必须高度自动化,否则周期压缩无从谈起。
- L4:小团队更容易犯错——没有专职的角色资产管理角色,一致性风险更高。这解释了为何 AI 短剧"人物与场景一致性问题频发"。
3.2.3 效果
| 作品 | 周期变化 | 成本变化 | 其他 |
|---|---|---|---|
| 《白狐》 | 传统 3 个月 → 2 周 | 每分钟成本从数万元降至万元级(另一口径为"万元以内") | 4 人团队 |
| 《山海奇镜之劈波斩浪》 | 通常 3~6 个月 → 2 个月 | 成本仅为传统制作的 1/4 | 快手,2024 |
| 《兴安岭诡事》 | — | — | 上线不到 21 小时播放破千万;累计 5613.3 万(截至 2025-10-17);抖音原生端收益超 30 万元;涨粉 10 万+ |
| 《奶团太后宫心计》 | — | — | 68 集,累计播放 2.1 亿 |
来源:21 世纪经济报道,2025-10-23;公开报道,2025。
3.3 上市公司与平台生态的 AI 短剧布局
3.3.1 背景
2025 年,AI 短剧从"创作者实验"进入"机构化生产"阶段。上市公司、平台方与工具方同时入场,形成了制作—工具—分发—激励的完整生态。
3.3.2 方案
| 主体 | 布局 |
|---|---|
| 博纳影业 | 2023 年底成立 AIGMS 制作中心(融合 AIGC 与电影工业化制片体系);2024-07 与抖音联合出品国内首部 AIGC 科幻短剧《三星堆:未来启示录》;推出"博纳一键 AI 短剧生成平台" |
| 昆仑万维 | SkyReels 平台 + DramaWave 分发平台 |
| 掌阅科技 | 未来每月上线约 3 部 AI 短剧 |
| 国脉文化 | 累计完成 240 集 AI 短剧制作,全网播放量突破 7000 万次 |
| 中文在线 | 2024 年用 AI 制作近百部漫画与动态漫,累计观看量超 30 亿次 |
| 华谊兄弟 | 储备 7 部 AI 短剧和 1 部 AI 电影,将推出国内影视行业首个全 AI 制作片单 |
| 光线传媒 | AI 智能系统批量评估和诊断上千个积压剧本,缩减剧本孵化周期 |
| 井英科技(海外) | 《After Divorce: My Five Brothers Paved My Way to the Billionaire Throne》登顶短剧周榜,累计热力值超过 500 万,被称为全球首部跻身短剧票房畅销榜的 AI 短剧 |
| 平台激励 | 即梦 AI 与抖音启动"AIGC 短剧联合招募计划";快手推出"星芒短剧 × 可灵 AI 大模型「AI 创想剧场」",均提供千万级现金激励与亿级流量扶持 |
从 Harness 视角看,这一阶段的特征是工具与分发的一体化:SkyReels 配 DramaWave,即梦配抖音,可灵配快手。分发端的反馈数据(完播率、留存)回流到制作端,构成 L5 评估闭环——这在小团队阶段是无法实现的。
3.3.3 效果
- 国脉文化:240 集,全网播放 7000 万次。
- 中文在线:近百部 AI 漫画与动态漫,累计观看量超 30 亿次。
- 井英科技:海外 AI 短剧登顶周榜,热力值超 500 万。
- 行业层面:横店影视城信号显示,2023 年为微短剧剧组定制数十个场景,2026 年同期来实景拍摄的剧组数量大幅下降(人民日报调查)。
来源:21 世纪经济报道,2025-10-23;人民日报,2026。
4. 实践标准
4.1 AGENTS.md 规范
4.1.1. AGENTS.md(AI 网剧 · AI Drama 方向)
# AGENTS.md —— AI 网剧(AI Drama)
## 角色与边界
- 你运行在 AI 短剧制作流水线之上,负责剧本、分镜、生成、剪辑、配音五段任务。
- 你负责执行与制作,不负责创意方向与内容价值判断。
- 你不得在无具名人类导演或制片人签核的情况下输出成片。
- 你不得生成任何真实自然人(明星、演员、公众人物、普通个人)的肖像与声音。
- 你不得虚假标注主创信息。
## 环境假设
- 存在角色资产库:每个角色有三视图(正/侧/背)、表情集、色卡、音色与细粒度属性标签。
- 存在剧本与分镜系统;存在视频生成、TTS、剪辑、BGM 工具。
- 存在非编与渲染农场,支持批量渲染与导出封装。
- 存在镜头状态机:以镜头为单位记录生成状态、版本、校验结果与失败原因。
- 存在前置审核流程(AI 生成内容必须前置审核,而非事后抽查)。
- 存在 scripts/ 目录承载确定性操作:抽帧、色彩校正、元数据写入、标识渲染、一致性比对、导出封装。
## 上下文加载顺序(Context Budget)
1. 用户显式指令与本次镜头的验收标准
2. 本文件(方向级)与组级 AGENTS.md
3. 角色资产:三视图、表情集、色卡、音色及其版本号(不可裁剪)
4. 当前镜头在镜头状态机中的上下文:前一镜头收尾状态、本镜头分镜描述(不可裁剪)
5. 剧本当前段落与人物小传
6. 风格参考集与历史素材(可裁剪)
规则:角色三视图与前一镜头状态属于锚定上下文,任何情况下不得被裁剪。角色漂移的根因几乎总是锚定信息被挤出上下文。
## 工具契约
- 只读类(自由调用):查询角色库、查询剧本、查询分镜、查询镜头状态机、查询素材。
- 生成类(输出必须进校验关卡):剧本生成、角色形象生成、分镜生成、视频生成、TTS、BGM 匹配。
- 校验类(阻断式):一致性比对(面部/服饰/道具/场景/音色)、标识校验、版权与肖像核查、内容前置审核。
- 写操作类(二次确认 + 留痕):覆盖资产库版本、提交成片、提交分发。
- 抽帧、色彩校正、元数据写入、标识渲染、一致性比对、导出封装一律走 scripts/。
## 任务执行流程(SOP)
1. 确认本次任务:集数、镜头号、分镜描述、时长与规格。
2. 加载角色资产与前一镜头状态,锁定版本号。
3. 生成:按镜头类型选择模型与参数;输出 2~3 条候选。
4. 一致性校验:与三视图做面部/服饰/道具比对;与前一镜头做场景与色调比对。
5. 重生成回路:不通过则调整锚定输入后重生成;**同一镜头连续 2 次不通过则升级人工**。
6. 音频:从角色音色库取音色,按情绪匹配 BGM。
7. 标识处理:显著标识写入;隐式元数据写入。
8. 前置审核:内容合规、版权、肖像、主创信息真实性。
9. 人工签核:导演或制片人确认。
10. 导出与回读:导出封装后回读,确认元数据与标识在位。
11. 更新镜头状态机:记录版本、校验结果、失败原因、成本。
## 验证与证据要求
- 每个镜头必须留存:候选素材路径、比对分数、通过/不通过判定、失败原因(若失败)。
- 跨集交付时必须提供角色一致性抽查报告(抽查比例与判定标准由制片人确定)。
- 标识必须给出证据:显著标识位置 + 隐式元数据回读结果。
- 主创信息必须真实:不得将 AI 生成标注为真人导演/演员作品。
- 引用效果数据必须标注口径层级。
## 失败与升级策略
- 角色漂移(面部/服饰/道具):回滚到镜头状态机最近通过点,重新加载三视图后重生成;连续 2 次失败升级人工。
- 场景漂移:把前一镜头的收尾帧作为参考图注入。
- 音色漂移:强制从角色音色库取值,禁止临时生成。
- 情感表达不足或人物塑造单薄:不强行加特效,回到剧本与表演设计环节。
- 内容同质化:在剧本阶段增加题材与结构约束,而非在生成阶段补救。
- 前置审核不通过:不得进入分发环节,按审核意见修改后重审。
- 升级必须携带:集数、镜头号、状态机记录、比对证据、已尝试处理。
## 安全与合规红线
- 建立 AI 训练数据版权合规机制;训练与参考素材必须有授权台账。
- 生成内容需符合社会主义核心价值观,不得生成违法违规内容。
- 不得侵害他人知识产权、隐私权、名誉权;禁止 AI 换脸明星。
- AI 生成内容必须进行**显著标识**;并写入隐式元数据。
- **不得虚假标注主创信息**。
- 建立内容审核机制,**对 AI 生成内容进行前置审核**。
- 不得恶意删除、篡改、伪造、隐匿标识(第十条);去标识日志留存不少于六个月(第九条)。
## 禁止事项
- 禁止使用真实自然人肖像与声音,包括换脸与音色克隆。
- 禁止虚假标注导演、编剧、演员等主创信息。
- 禁止跳过一致性校验与前置审核。
- 禁止把角色状态存放在模型会话里;必须写入镜头状态机。
- 禁止让模型逐 token 生成抽帧结果、调色参数、元数据与导出文件。
- 禁止编造规范编号、条款与案例数值。
## 输出格式
集数 / 镜头号 / 使用角色资产版本 / 分镜描述 / 候选数量与路径 / 一致性比对结果 / 标识证据 / 前置审核结果 / 待签核项 / 责任人 / 状态机更新记录 / 遗留问题
## 评估与自检
- 角色三视图是否进入了上下文且未被裁剪?
- 本镜头的状态是否已写入状态机(不依赖会话记忆)?
- 一致性比对分数是否留存?
- 标识是否可回读?
- 主创信息标注是否真实?
- 本次失败样本是否已纳入评估集? 4.2 SKILL.md 规范
4.2.1. SKILL.md(AI 网剧 · 单镜头合规交付)
---
name: ai-drama-shot-delivery
description: AI 网剧单镜头交付技能。当需要从分镜描述出发生成一个镜头素材,并完成角色一致性比对、场景与色调连续性校验、显著标识与隐式元数据写入、内容前置审核、导演签核与镜头状态机更新时使用。适用于 AI 微短剧、AI 网剧的分段批量生产场景。
version: 1.0
created: 2026-09-12
---
# AI 网剧 · 单镜头合规交付
## 适用场景
- 从分镜描述与角色资产出发生成一个镜头的可交付素材。
- 需要在分段生成模式下保证角色、服饰、道具、场景与音色的跨镜头一致性。
- 需要把每个镜头的生成状态、校验结果与失败原因写入镜头状态机,支持单独重生成。
## 前置条件
- 已加载方向级 AGENTS.md 与组级 AGENTS.md。
- 角色资产库可用:三视图(正/侧/背)、表情集、色卡、音色,且带版本号。
- 镜头状态机可用,能读到前一镜头的收尾状态与色调基线。
- scripts/ 中存在抽帧、色彩校正、一致性比对、元数据写入、标识渲染、导出封装脚本。
- 已确定一致性判定阈值与抽查比例,并经制片人确认。
- 已确定具名签核人(导演或制片人)与前置审核流程。
## 输入
| 输入项 | 说明 | 必需 |
|---|---|---|
| 集数与镜头号 | 用于状态机定位 | 是 |
| 分镜描述 | 景别、运镜、时长、表演要求 | 是 |
| 角色资产引用 | 角色 ID + 资产版本号 | 是 |
| 前一镜头状态 | 收尾帧、色调基线、场景 ID | 是 |
| 输出规格 | 分辨率、帧率、时长、画幅 | 是 |
| 一致性阈值 | 面部/服饰/道具/场景的判定阈值 | 是 |
## 输出
- 镜头素材(含 2~3 条候选)
- 一致性比对报告(面部/服饰/道具/场景/音色)
- 标识证据(显著标识位置 + 隐式元数据回读结果)
- 前置审核结果
- 签核记录与镜头状态机更新记录
## 执行步骤
1. 从镜头状态机读取本镜头任务与前一镜头收尾状态。
2. 加载角色三视图、表情集、色卡与音色,锁定版本号。
3. 按镜头类型选择模型与参数,生成 2~3 条候选。
4. 一致性比对:与三视图比对面部/服饰/道具;与前一镜头收尾帧比对场景与色调。
5. 重生成回路:不通过则调整锚定输入(补参考图、加约束条件)后重生成;连续 2 次不通过升级人工。
6. 音频处理:从角色音色库取音色,按情绪匹配 BGM。
7. 标识处理:脚本渲染显著标识并写入隐式元数据。
8. 前置审核:内容合规、版权、肖像、主创信息真实性。
9. 导演或制片人签核。
10. 导出封装并回读校验。
11. 更新镜头状态机:版本、比对分数、失败原因、成本、责任人。
## 质量标准(DoD)
- 面部/服饰/道具/场景/音色五项一致性判定全部通过,且留存比对分数。
- 显著标识在位;隐式元数据可回读;去标识日志留存不少于六个月。
- 前置审核通过并留痕。
- 主创信息标注真实,无虚假标注。
- 镜头状态机已更新,本镜头可单独重生成而不影响其他镜头。
- 引用效果数据标注口径层级(如"成本 50~60 万 → 不到 20 万"为从业者访谈口径)。
## 常见失败与处理
- 角色换脸:三视图未入上下文或被裁剪 → 重新加载并固化为不可裁剪区。
- 场景漂移:未引用前一镜头收尾帧 → 把收尾帧作为参考图注入。
- 帧间不一致(多出的轮子、错位的轴):缩短单次生成长度,增加逐帧人工精修节点。
- 情感表达不足:不要靠加大特效掩盖,回到剧本与表演设计环节。
- 内容同质化(扎堆玄幻/科幻):在剧本阶段增加题材约束与结构多样性要求。
- 音色跨集漂移:强制从角色音色库取值,禁止临时生成新音色。
- 沦为"幻灯片堆砌":把叙事连贯性纳入判定标准,限制单镜头华丽度权重。
- 前置审核不通过:按审核意见修改后重审,不得绕过。
## 示例
任务:第 7 集第 12 镜头,中景,角色 A 在雨中回身,时长 6 秒,1080P / 30 fps / 竖屏。
输入:角色 A 资产版本 v3.2;前一镜头(SC-0711)收尾帧与色调基线;分镜描述;一致性阈值。
执行:加载三视图 → 生成 3 条候选 → 比对(候选 1、2 服饰偏差超阈值,候选 3 通过)→ 音色库取音色 → 标识渲染 → 前置审核 → 导演签核 → 导出回读 → 更新状态机。
输出:成片素材 + 比对报告(含 2 条失败样本)+ 标识证据 + 签核记录 + 状态机记录(SC-0712 = passed)。 4.3 落地检查清单
| 编号 | 检查项 | 层级 | 判定 | 说明 |
|---|---|---|---|---|
| E-01 | 角色资产库已建立:三视图 + 表情集 + 色卡 + 音色 | L4 | 必备 | 本方向瓶颈的核心结构 |
| E-02 | 每个角色资产有版本号与授权状态 | L4 | 必备 | — |
| E-03 | 镜头状态机已建立,支持单镜头独立重生成 | L4 | 必备 | 不得以"整段重做"替代 |
| E-04 | 跨集连续性有抽查机制与抽查比例 | L4 | 必备 | 抽查比例由制片人确定 |
| E-05 | 角色三视图与前一镜头状态列为不可裁剪上下文 | L1 | 必备 | 漂移的主要根因 |
| E-06 | 五段流水线(剧本→分镜→生成→剪辑→配音)工具齐备 | L2 | 必备 | — |
| E-07 | 按镜头类型动态调用不同参数量级模型 | L2 | 建议 | 昆仑万维采用此策略 |
| E-08 | 抽帧、色彩校正、元数据写入、标识渲染已脚本化 | L2 | 必备 | — |
| E-09 | 一致性校验内置在流水线中,非事后人工发现 | L3 | 必备 | 校验后接重生成回路 |
| E-10 | 重生成回路支持连续 2 次失败自动升级人工 | L3 | 必备 | — |
| E-11 | 单次生成长度受模型规格约束时,已做分段生成设计 | L3 | 必备 | 如 Veo 3.1 API 为 4/6/8 秒 |
| E-12 | 一致性判定覆盖面部/服饰/道具/场景/音色五项 | L5 | 必备 | 缺一项即会出现出戏点 |
| E-13 | 叙事连贯性纳入判定标准 | L5 | 必备 | 防止"幻灯片堆砌" |
| E-14 | 成本与周期纳入观测 | L5 | 必备 | 参考:整部 < 20 万元 |
| E-15 | AI 训练数据版权合规机制已建立 | L6 | 必备 | 行业规范明确要求 |
| E-16 | 内容前置审核流程已建立并留痕 | L6 | 必备 | 非事后抽查 |
| E-17 | 显著标识已由脚本渲染,隐式元数据可回读 | L6 | 必备 | — |
| E-18 | 主创信息标注真实,无虚假标注 | L6 | 必备 | 行业规范明确禁止 |
| E-19 | 真实自然人肖像与声音已被拦截(含换脸与音色克隆) | L6 | 必备 | 硬性红线 |
| E-20 | 去标识日志留存不少于六个月 | L6 | 必备 | 《标识办法》第九条 |
5. 总结
AI 网剧方向把创意产业组的 L4 瓶颈暴露得最为彻底。原因很简单:它是唯一一个需要把"同一个角色"保持几十集、上百个镜头、数周制作周期的方向。面部、服饰、道具、场景、音色五个维度任何一个漂移,观众立刻出戏。
行业给出的答案不是"更好的模型",而是"更好的状态载体"。昆仑万维 SkyReels 的设计中,最具参考价值的不是 1080P / 60 fps / 180 秒的规格,而是角色库与分镜生成的耦合——分镜直接引用角色库资产,而不是每次重新描述角色。这一设计把一致性从"模型能力问题"变成了"工程结构问题"。
成本与周期的数据是真实的:整部 AI 短剧从 50~60 万元降到不到 20 万元,《白狐》从 3 个月压到 2 周,《山海奇镜》从 3~6 个月压到 2 个月、成本降至 1/4。但必须同时看到行业尚未解决的问题:尚未诞生现象级作品、严重同质化、整体盈利能力未充分验证、以及"人物与场景一致性问题频发"。
规范侧,中国网络视听协会于 2025 年 4 月发布《网络视听 AI 短剧创作生产与传播规范》,要求 AI 生成内容显著标识、不得虚假标注主创信息、建立训练数据版权合规机制、对 AI 生成内容进行前置审核。需要如实说明:该规范的正式文号与条款全文未能在协会官网核实到原文,本文仅按二手来源引述,标注 。
信息缺口声明
| 缺口项 | 处理方式 |
|---|---|
| 《网络视听 AI 短剧创作生产与传播规范》的正式文号与条款全文 | 仅检索到二手页面(aiww.cn)称"2025 年 4 月由中国网络视听节目服务协会正式发布",未在协会官网核实到原文、文号与条款 → 只写时间与机构,标注 ,不编造文号 |
| 广电总局《微短剧发展管理办法(征求意见稿)》(11 项禁播红线、AI 短剧每集标识要求) | 仅见于券商研报汇编,称 2026-06 发布,未核实到原文 → 标注"据券商研报汇编"并加 |
| 广电总局"微短剧+"行动计划、分类分层审核制度 | 同上,仅见于研报汇编 → 标注 |
| 《网络安全技术 人工智能生成合成内容标识方法》的 GB 编号 | 已确认:GB 45438—2025(强制性国标,2025-02-28 发布、2025-09-01 实施,与《标识办法》同步;来源:国家标准全文公开系统、TC260 官方文本) |
| SkyReels 视频生成规格的两个口径(180 秒 vs 60 秒以上) | 两种口径并存 → 本文采用"1080P / 60 fps / 单次 180 秒"并注明另一口径 |
| 《白狐》每分钟成本的两个口径(万元级 vs 万元以内) | 两种口径并存 → 并列呈现,不择一 |
| 井英科技海外 AI 短剧的收入绝对值 | 仅检索到"热力值超过 500 万"与"登顶周榜",无收入绝对值 → 不补充 |
| AI 微短剧的完播率、留存率等分发侧指标 | 未检索到公开数据 → [待填写] |
6. 参考资料
- 人工智能生成合成内容标识办法 — 国家互联网信息办公室、工业和信息化部、公安部、国家广播电视总局,2025。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- AI 短剧:资本追逐的新风口 — 21 世纪经济报道,2025-10-23。https://www.21jingji.com/article/20251023/herald/89aa15c81c1bf4ae8b48bf94060ed94e.html
- AI 要给微短剧"洗牌"?— 人民日报,2026。https://kpzg.people.com.cn/n1/2026/0511/c404214-40717117.html
- AI 要给微短剧"洗牌"?(湖南频道转载)— 人民网,2026。https://hn.people.com.cn/BIG5/n2/2026/0511/c208814-41577065.html
- AI 短剧规范发布(中国网络视听协会)— AIWW。注:二手来源,正式文号未核实。 https://www.aiww.cn/s/ff3ac3ba84deb97b
- 昆仑万维 2024 年年度报告摘要(SkyReels 技术栈)— 证券时报,2025。https://epaper.stcn.com/pic/202504/26/8eec6686f3528439d13a5db995a8f55b.pdf
- Kunlun Tech 官网(SkyReels-A1 等开源模型)— 昆仑万维。https://www.kunlun.com/en
- 多措并举推进标识体系建设,助力新时代人工智能健康发展(标识落地方法论与 6 项标准实践指南)— 国家互联网应急中心,2025。https://www.cac.gov.cn/2025-09/06/c_1758880709361356.htm
- 主流媒体所办新媒体发展研究报告(2024-2025)— 人民网,2025。https://sc.people.com.cn/BIG5/n2/2025/1030/c345167-41396739.html
- AGENTS.md 官方站 — Agentic AI Foundation(Linux Foundation)。https://agents.md/
- Agent Skills Specification — agentskills.io。https://agentskills.io/specification
- Equipping agents for the real world with Agent Skills — Anthropic,2025(2025-12-18 更新)。https://claude.com/blog/equipping-agents-for-the-real-world-with-agent-skills
AI Harness for AI Drama
1. Introduction
1.1 Background
AI dramas (AI micro-dramas) are the direction within the Creative Industries cluster that is scaling the fastest while being the most lagging in specification-building. According to the China Netcasting Services Association's Guidelines for Micro-drama Creation (as cited in People's Daily, 2026), in Q1 2026 the industry went live with about 128,000 micro-dramas, of which roughly 122,000 were AI micro-dramas, a share exceeding 95%. In the same period, research-institution data shows that among Douyin's monthly TOP 5000 short dramas, the number of fully AI-generated micro-dramas grew from 4 in January 2025 to 69 in October and 217 in November.
The direct consequence of this penetration rate is that the variance in quality has been amplified to the extreme. On the cost side, a mid-tier live-action short drama of 60 episodes at 2 minutes each costs roughly 500,000–600,000 CNY; an AI short drama of the same spec costs about 2,000–3,000 CNY per episode and under 200,000 CNY for the whole series (practitioner-interview basis, as cited in People's Daily, 2026). The flip side of an order-of-magnitude cost reduction is that many works fail to meet standards on "story coherence" and "character consistency."
The 21st Century Business Herald (2025-10-23) summary of industry problems can serve as this direction's "common failures" baseline:
- Content layer: no phenomenon-level hit has yet emerged; most rely on platforms to carry them; excessive pursuit of visual splendor and smooth motion while neglecting story coherence and character building, devolving into a "slide-show pile-up"; severe homogenization, clumping around xuanhuan (fantasy) and sci-fi themes.
- Tech layer: insufficient emotional expression and character-portrayal capability; character-and-scene consistency problems occur frequently.
- Commercial layer: overall profitability is not yet fully validated; most rely on platform subsidies and traffic support; weak IP attributes and low user retention.
- Copyright and ethics: some production crews use AI to generate scripts in one click, deepfake celebrities, and even sell "customized intimate scenes."
1.2 Definition and Scope
In this direction, AI Harness refers to the engineered carrier layer that, in AI micro-drama / online-drama production scenarios, takes on the five-stage pipeline of script → storyboard → generation → editing → dubbing, organizing scattered video-generation capabilities into a controllable, reproducible, regression-verifiable finished-production system.
| Pipeline Stage | Typical Decisions | Output Artifacts | Consistency Risk |
|---|---|---|---|
| Script | Genre, structure, hook density, character profile | Script, episode outline, character profile | Character setting contradicting itself |
| Storyboard | Shot sequence, shot size, camera movement | Storyboard script, shot list | Incoherent shot language |
| Generation | Character appearance, scene, performance, duration | Shot footage (multiple candidates) | Character face-swapping, scene drift |
| Editing | Rhythm, transitions, audio-sync | Rough-cut version | Inter-frame inconsistency |
| Dubbing | Voice timbre, emotion, BGM, sound effects | Final audio track | Voice-timbre drift across episodes |
1.3 Position within the AI Harness System
图 1-1|AI 网剧 Harness:六层体系与五段流水线
数据来源:基于本文分析绘制的示意图。
Dominant layer: L3 Orchestration and Control layer + L4 Memory and State layer. Bottleneck layer: L4 Memory and State layer.
| Layer | Manifestation in AI Online Dramas | Key Constraint |
|---|---|---|
| L1 Context Engineering | Script, character profile, standard character turnaround views, and storyboard script serve as anchor context | The turnaround views and color card are the physical anchors of character consistency |
| L2 Tools and Execution | Per-stage tools for the scripting model, character generation, storyboard generation, video generation, TTS, editing, and BGM | Dynamically invoke models of different parameter scales by shot type |
| L3 Orchestration and Control | Five-stage pipeline: script → storyboard → generation → editing → dubbing; segmented generation + consistency checking + regeneration | The regeneration loop must be built in; it cannot be patched in only after humans discover a problem |
| L4 Memory and State | Character asset library (turnaround views + expression set + color card + vocal timbre) and shot state machine; cross-episode continuity | The quality bottleneck of this direction |
| L5 Evaluation and Observation | Character consistency (face/apparel/props), style stability, narrative coherence, cost and cycle time | All four are indispensable |
| L6 Governance and Safety | Training-data copyright compliance mechanism; pre-release review; conspicuous labeling; no false credit attribution; portrait rights and reputation rights | Pre-release review is this direction's distinctive hard constraint |
This direction exhibits the L4 bottleneck most typically. There are three reasons:
- The longest time span crossed: a 60-episode short drama's production cycle is measured in weeks; a model session cannot retain state that long.
- The greatest number of consistency dimensions: face, apparel, props, scene, and vocal timbre—five dimensions all must stay consistent; a drift in any one of them breaks immersion.
- High regeneration cost: when character drift is discovered at episode 40, regenerating means rerunning the entire pipeline.
Therefore, the core structure of L4 in this direction is not "memory," but the shot state machine: recording generation state, version, verification results, and failure reasons on a per-shot basis, so that any single shot can be regenerated independently without affecting other shots.
1.4 Industry Status and Published Quantitative Metrics
| Indicator | Value | Source & Basis |
|---|---|---|
| Q1 2026 micro-dramas launched | About 128,000 | China Netcasting Services Association, Guidelines for Micro-drama Creation, as cited in People's Daily, 2026 |
| Of which AI micro-dramas | About 122,000 (share > 95%) | Same as above |
| Fully AI micro-dramas within Douyin TOP 5000 | 2025-01: 4 → 2025-10: 69 → 2025-11: 217 | Research-institution data, as cited in People's Daily, 2025 |
| Mid-tier live-action short drama cost (60 episodes × 2 min) | About 500,000–600,000 CNY | Practitioner interviews, as cited in People's Daily, 2026 |
| AI short drama cost of the same spec | 2,000–3,000 CNY per episode; whole series < 200,000 CNY | Same as above |
| Bai Hu production cycle | 3 months → 2 weeks | 21st Century Business Herald, 2025 |
| Bai Hu per-minute cost | Tens of thousands → ten-thousand-level (another figure is "under ten thousand") | Same as above |
| Shan Hai Qi Jing: Ji Po Zhan Lang cycle / cost | 3–6 months → 2 months; cost is 1/4 of traditional production | Same as above |
| Nai Tuan Tai Hou Gong Xin Ji (68 episodes) | Cumulative views 210 million | Public reports, 2025 |
| Xing'anling Gui Shi | Over 10 million views within 21 hours of launch; cumulative 56.133 million as of 2025-10-17; Douyin native-channel revenue over 300,000 CNY; 100,000+ new followers | Public reports, 2025 |
Video-generation model specifications (consolidated from securities research reports, Haitong, 2025-10-29):
| Model | Specification |
|---|---|
| Sora 2 | App default 10-second portrait; standard tier 720×1280 / 1280×720 |
| Kling | Officially claimed 1080p, 30 fps, up to 2 minutes |
| Veo 3.1 | API preview 4 / 6 / 8 seconds; Flow "scene continuation" up to about 1 minute |
| JiMeng | Targets an image/video-integrated "smart canvas"; specifications not uniformly published |
The spec differences directly determine the orchestration strategy: single-generation length is limited, which necessarily requires a "segmented generation + consistency checking + regeneration" pipeline structure — this is the physical premise of this direction's L3 orchestration-layer design.
2. Glossary
| Term | English / Abbreviation | Definition |
|---|---|---|
| AI micro-drama | AI Micro-drama | An episodic format with very short single-episode duration, generated mainly or entirely by AI |
| Character turnaround | Character Turnaround | Reference images of a character from three standard views (front, side, back); the physical anchor of consistency |
| Character asset library | Character Asset Library | A reusable asset library that carries character turnaround views, expression sets, color cards, vocal timbres, and attribute tags |
| Shot state machine | Shot State Machine | A state machine that records generation state, version, verification results, and failure reasons on a per-shot basis |
| Character profile | Character Profile | A structured description of a character's identity, personality, motivations, and relationships, used for subsequent character design |
| Storyboard generation | Storyboard Generation | The process of automatically generating shot sequences, shot sizes, and camera movements from a script |
| Hook | Hook / Satisfaction Point | Plot nodes in short dramas that create emotional release; a key focus of script-model optimization |
| Emotional voice bank | Emotional Voice Bank | A collection of character vocal timbres classified by emotion, used to precisely match character and scene audio |
| Motion in-betweening | Motion In-Betweening, MIB | Generating transition frames between keyframes to eliminate sliding and jitter |
| AI motion capture | AI Motion Capture | A technology that drives AI actors' performance using any camera or video |
| Cross-episode continuity | Cross-episode Continuity | The ability to keep characters, scenes, and plot consistent across episodes |
| Pre-release review | Pre-release Review | A review stage completed before content goes live; AI short dramas are explicitly required to establish this mechanism |
| Conspicuous label | Conspicuous Label | A prompt label added in a conspicuous location to indicate generated/synthetic content; a mandatory requirement for AI short dramas |
| False credit attribution | False Credit Attribution | The act of falsely attributing credit to directors, screenwriters, actors, and other key creators; explicitly prohibited by industry standards |
| Training data copyright compliance | Training Data Copyright Compliance | The training-data licensing and compliance management system that developers/operators must establish |
| Tiered content review | Tiered Content Review | A content review system tiered by investment amount and theme |
| "AI mashup" | AI Mashup / Remix | A content form that uses AI to make distorting adaptations of existing film/TV content; has been brought under special governance |
3. Case Studies
3.1 Kunlun Tech SkyReels: A Full-Pipeline Platform for AI Short Dramas
3.1.1 Background
In August 2024, Kunlun Tech launched SkyReels, positioned as the world's first AI short-drama creation platform integrating a video foundation model with a 3D foundation model. The core problem it set out to solve: AI short-drama production involves six types of capabilities—script, character, storyboard, video, audio, and editing—spread across different tools, with all intermediate state transfer done manually, causing consistency to degrade at every handover.
3.1.2 Approach
SkyReels built a complete pipeline whose structure is itself a reference implementation of an AI online-drama Harness:
| Module | Capability | Corresponding Harness Layer |
|---|---|---|
| SkyScript script foundation model | Quickly creates structurally complete, plot-rich short-drama scripts; optimized "hook" generation, with human ratings consistently reaching A/S grade; adds a vast library of proven hit creative templates; can summarize character profiles for subsequent character design | L1 → L3 |
| Character library | Introduces realistic AI actors with fine-grained attribute tags that intelligently match script character settings; supports one-click generation of character appearance and dubbing (emotional voice bank) | L4 |
| StoryboardGen storyboard generation foundation model | Leverages storyboard artists' professional experience and context-enhancement techniques to improve storyboard coherence and character consistency | L3 |
| Sky3DGen | Generates diverse 3D elements and scenes, with real-time 3D scene interaction capability | L2 |
| WorldEngine | Deeply integrates an AI 3D engine with a video foundation model to solve the problem of incoherent physics in traditional video creation | L2 |
| Video generation | 1080P / 60 fps, with single-generation length up to 180 seconds (another figure says "video generation of over 60 seconds has been achieved"); dynamically invokes models of different parameter scales by shot type | L2 |
| Audio | Builds an emotional voice bank and a short-drama BGM library to precisely match character and scene audio | L2 + L4 |
| Web client | 3D interactive editing (real-time adjustment of storyboard frames and characters); AI motion capture (driving AI actor performance with any camera or video); integrates a million-scale film/TV motion database and opens it to customization | L2 |
| Open-source model SkyReels-A1 | Given an input video sequence and a reference portrait, extracts facial-expression-aware feature points as motion descriptors to transfer onto the portrait, based on a DiT conditional video-generation framework | L2 |
From a Harness perspective, SkyReels' most critical design is the coupling of the character library with storyboard generation: storyboard generation directly references assets in the character library rather than re-describing the character each time. This is precisely the engineering embodiment of turning L4 from "memory" into a "state carrier."
3.1.3 Results
- Scripts generated by SkyScript consistently achieve a human rating of A/S grade.
- Video-generation specifications reach 1080P / 60 fps / 180 seconds per generation.
- The Kunlun Tech SkyReels platform plus the DramaWave distribution platform reached US $10 million in monthly revenue as of Q1 2025.
Sources: Kunlun Tech official website; Securities Times, Summary of Kunlun Tech's 2024 Annual Report, 2025. The monthly revenue figure is the enterprise-disclosed basis.
3.2 Bai Hu and Shan Hai Qi Jing: Compressing Cycle Time and Cost
3.2.1 Background
The cycle time and cost structure of traditional micro-dramas are determined by three parts: crew labor, location and shooting, and post-production. For xuanhuan (fantasy), sci-fi, and costume genres in particular, location and post-production account for an especially high share—a director once said that 90% of one drama's scenes took place in Xing'anling, and "if we were to shoot on location in the forested snowfields, the cost would be extremely high and the danger factor very large."
3.2.2 Approach
- Bai Hu (a fully AI-produced micro-drama): a 4-person team, using ChatGPT for rapid script iteration + AI image generation + intelligent editing.
- Shan Hai Qi Jing: Ji Po Zhan Lang (Kuaishou, director Chen Kun, 2024): produced through an all-AI end-to-end process.
- Xing'anling Gui Shi (China's first paid AI short drama, director Ding Kuan): uses AI to take over "the most tedious, money-burning, and most dangerous work."
Breaking it down from a Harness perspective:
- L2: an extremely small team (4 people) means the toolchain must be highly integrated; any intermediate format requiring manual handling becomes a bottleneck.
- L3: the script → storyboard → generation → editing → dubbing pipeline must be highly automated in a 4-person team, otherwise cycle-time compression is out of the question.
- L4: small teams are more prone to mistakes—with no dedicated character-asset-management role, consistency risk is higher. This explains why AI short dramas "frequently suffer character-and-scene consistency problems."
3.2.3 Results
| Work | Cycle-Time Change | Cost Change | Other |
|---|---|---|---|
| Bai Hu | Traditionally 3 months → 2 weeks | Per-minute cost reduced from tens of thousands to ten-thousand-level (another figure is "under ten thousand") | 4-person team |
| Shan Hai Qi Jing: Ji Po Zhan Lang | Usually 3–6 months → 2 months | Cost is only 1/4 of traditional production | Kuaishou, 2024 |
| Xing'anling Gui Shi | — | — | Over 10 million views within 21 hours of launch; cumulative 56.133 million (as of 2025-10-17); Douyin native-channel revenue over 300,000 CNY; 100,000+ new followers |
| Nai Tuan Tai Hou Gong Xin Ji | — | — | 68 episodes, cumulative views 210 million |
Sources: 21st Century Business Herald, 2025-10-23; public reports, 2025.
3.3 AI Short Drama Strategies of Listed Companies and Platform Ecosystems
3.3.1 Background
In 2025, AI short dramas moved from the "creator experiment" stage into the "institutionalized production" stage. Listed companies, platforms, and tool providers entered at the same time, forming a complete ecosystem of production—tools—distribution—incentives.
3.3.2 Approach
| Player | Deployment |
|---|---|
| Bona Film Group | Established the AIGMS production center at the end of 2023 (integrating AIGC with a film-industrialization production system); co-produced with Douyin in 2024-07 China's first AIGC sci-fi short drama Sanxingdui: Future Revelation; launched the "Bona one-click AI short-drama generation platform" |
| Kunlun Tech | SkyReels platform + DramaWave distribution platform |
| Zhangyue Technology | Plans to launch about 3 AI short dramas per month |
| Guomai Culture | Cumulatively completed 240 episodes of AI short-drama production, with all-network views surpassing 70 million |
| ChineseAll | In 2024 produced nearly a hundred comics and dynamic comics with AI, with cumulative views exceeding 3 billion |
| Huayi Brothers | Has 7 AI short dramas and 1 AI film in reserve, and will release the first all-AI-produced slate in the domestic film/TV industry |
| Enlight Media | An AI intelligent system batch-evaluates and diagnoses over a thousand backlogged scripts, shortening the script-incubation cycle |
| Jingying Technology (overseas) | After Divorce: My Five Brothers Paved My Way to the Billionaire Throne topped the short-drama weekly chart, with cumulative heat value exceeding 5 million; hailed as the world's first AI short drama to enter the short-drama box-office best-seller chart |
| Platform incentives | JiMeng AI and Douyin launched the "AIGC short-drama joint recruitment program"; Kuaishou launched the "Xingmang Short Drama × Kling AI Foundation Model 'AI Creative Theater'", both providing tens-of-millions cash incentives and hundreds-of-millions traffic support |
From a Harness perspective, the defining feature of this stage is the integration of tools and distribution: SkyReels paired with DramaWave, JiMeng with Douyin, and Kling with Kuaishou. Feedback data from the distribution side (completion rate, retention) flows back to the production side, forming an L5 evaluation loop—something impossible to achieve at the small-team stage.
3.3.3 Results
- Guomai Culture: 240 episodes, 70 million all-network views.
- ChineseAll: nearly a hundred AI comics and dynamic comics, with cumulative views exceeding 3 billion.
- Jingying Technology: overseas AI short drama topped the weekly chart, with heat value exceeding 5 million.
- Industry level: signals from Hengdian World Studios show that in 2023 dozens of scenes were customized for micro-drama crews, while in the same period of 2026 the number of crews coming for on-location shooting dropped sharply (People's Daily investigation).
Sources: 21st Century Business Herald, 2025-10-23; People's Daily, 2026.
4. Practice Standards
4.1 AGENTS.md Specification
4.1.1. AGENTS.md (AI Drama Direction)
# AGENTS.md —— AI 网剧(AI Drama)
## 角色与边界
- 你运行在 AI 短剧制作流水线之上,负责剧本、分镜、生成、剪辑、配音五段任务。
- 你负责执行与制作,不负责创意方向与内容价值判断。
- 你不得在无具名人类导演或制片人签核的情况下输出成片。
- 你不得生成任何真实自然人(明星、演员、公众人物、普通个人)的肖像与声音。
- 你不得虚假标注主创信息。
## 环境假设
- 存在角色资产库:每个角色有三视图(正/侧/背)、表情集、色卡、音色与细粒度属性标签。
- 存在剧本与分镜系统;存在视频生成、TTS、剪辑、BGM 工具。
- 存在非编与渲染农场,支持批量渲染与导出封装。
- 存在镜头状态机:以镜头为单位记录生成状态、版本、校验结果与失败原因。
- 存在前置审核流程(AI 生成内容必须前置审核,而非事后抽查)。
- 存在 scripts/ 目录承载确定性操作:抽帧、色彩校正、元数据写入、标识渲染、一致性比对、导出封装。
## 上下文加载顺序(Context Budget)
1. 用户显式指令与本次镜头的验收标准
2. 本文件(方向级)与组级 AGENTS.md
3. 角色资产:三视图、表情集、色卡、音色及其版本号(不可裁剪)
4. 当前镜头在镜头状态机中的上下文:前一镜头收尾状态、本镜头分镜描述(不可裁剪)
5. 剧本当前段落与人物小传
6. 风格参考集与历史素材(可裁剪)
规则:角色三视图与前一镜头状态属于锚定上下文,任何情况下不得被裁剪。角色漂移的根因几乎总是锚定信息被挤出上下文。
## 工具契约
- 只读类(自由调用):查询角色库、查询剧本、查询分镜、查询镜头状态机、查询素材。
- 生成类(输出必须进校验关卡):剧本生成、角色形象生成、分镜生成、视频生成、TTS、BGM 匹配。
- 校验类(阻断式):一致性比对(面部/服饰/道具/场景/音色)、标识校验、版权与肖像核查、内容前置审核。
- 写操作类(二次确认 + 留痕):覆盖资产库版本、提交成片、提交分发。
- 抽帧、色彩校正、元数据写入、标识渲染、一致性比对、导出封装一律走 scripts/。
## 任务执行流程(SOP)
1. 确认本次任务:集数、镜头号、分镜描述、时长与规格。
2. 加载角色资产与前一镜头状态,锁定版本号。
3. 生成:按镜头类型选择模型与参数;输出 2~3 条候选。
4. 一致性校验:与三视图做面部/服饰/道具比对;与前一镜头做场景与色调比对。
5. 重生成回路:不通过则调整锚定输入后重生成;**同一镜头连续 2 次不通过则升级人工**。
6. 音频:从角色音色库取音色,按情绪匹配 BGM。
7. 标识处理:显著标识写入;隐式元数据写入。
8. 前置审核:内容合规、版权、肖像、主创信息真实性。
9. 人工签核:导演或制片人确认。
10. 导出与回读:导出封装后回读,确认元数据与标识在位。
11. 更新镜头状态机:记录版本、校验结果、失败原因、成本。
## 验证与证据要求
- 每个镜头必须留存:候选素材路径、比对分数、通过/不通过判定、失败原因(若失败)。
- 跨集交付时必须提供角色一致性抽查报告(抽查比例与判定标准由制片人确定)。
- 标识必须给出证据:显著标识位置 + 隐式元数据回读结果。
- 主创信息必须真实:不得将 AI 生成标注为真人导演/演员作品。
- 引用效果数据必须标注口径层级。
## 失败与升级策略
- 角色漂移(面部/服饰/道具):回滚到镜头状态机最近通过点,重新加载三视图后重生成;连续 2 次失败升级人工。
- 场景漂移:把前一镜头的收尾帧作为参考图注入。
- 音色漂移:强制从角色音色库取值,禁止临时生成。
- 情感表达不足或人物塑造单薄:不强行加特效,回到剧本与表演设计环节。
- 内容同质化:在剧本阶段增加题材与结构约束,而非在生成阶段补救。
- 前置审核不通过:不得进入分发环节,按审核意见修改后重审。
- 升级必须携带:集数、镜头号、状态机记录、比对证据、已尝试处理。
## 安全与合规红线
- 建立 AI 训练数据版权合规机制;训练与参考素材必须有授权台账。
- 生成内容需符合社会主义核心价值观,不得生成违法违规内容。
- 不得侵害他人知识产权、隐私权、名誉权;禁止 AI 换脸明星。
- AI 生成内容必须进行**显著标识**;并写入隐式元数据。
- **不得虚假标注主创信息**。
- 建立内容审核机制,**对 AI 生成内容进行前置审核**。
- 不得恶意删除、篡改、伪造、隐匿标识(第十条);去标识日志留存不少于六个月(第九条)。
## 禁止事项
- 禁止使用真实自然人肖像与声音,包括换脸与音色克隆。
- 禁止虚假标注导演、编剧、演员等主创信息。
- 禁止跳过一致性校验与前置审核。
- 禁止把角色状态存放在模型会话里;必须写入镜头状态机。
- 禁止让模型逐 token 生成抽帧结果、调色参数、元数据与导出文件。
- 禁止编造规范编号、条款与案例数值。
## 输出格式
集数 / 镜头号 / 使用角色资产版本 / 分镜描述 / 候选数量与路径 / 一致性比对结果 / 标识证据 / 前置审核结果 / 待签核项 / 责任人 / 状态机更新记录 / 遗留问题
## 评估与自检
- 角色三视图是否进入了上下文且未被裁剪?
- 本镜头的状态是否已写入状态机(不依赖会话记忆)?
- 一致性比对分数是否留存?
- 标识是否可回读?
- 主创信息标注是否真实?
- 本次失败样本是否已纳入评估集? 4.2 SKILL.md Specification
4.2.1. SKILL.md (AI Drama Single-Shot Compliant Delivery)
---
name: ai-drama-shot-delivery
description: AI 网剧单镜头交付技能。当需要从分镜描述出发生成一个镜头素材,并完成角色一致性比对、场景与色调连续性校验、显著标识与隐式元数据写入、内容前置审核、导演签核与镜头状态机更新时使用。适用于 AI 微短剧、AI 网剧的分段批量生产场景。
version: 1.0
created: 2026-09-12
---
# AI 网剧 · 单镜头合规交付
## 适用场景
- 从分镜描述与角色资产出发生成一个镜头的可交付素材。
- 需要在分段生成模式下保证角色、服饰、道具、场景与音色的跨镜头一致性。
- 需要把每个镜头的生成状态、校验结果与失败原因写入镜头状态机,支持单独重生成。
## 前置条件
- 已加载方向级 AGENTS.md 与组级 AGENTS.md。
- 角色资产库可用:三视图(正/侧/背)、表情集、色卡、音色,且带版本号。
- 镜头状态机可用,能读到前一镜头的收尾状态与色调基线。
- scripts/ 中存在抽帧、色彩校正、一致性比对、元数据写入、标识渲染、导出封装脚本。
- 已确定一致性判定阈值与抽查比例,并经制片人确认。
- 已确定具名签核人(导演或制片人)与前置审核流程。
## 输入
| 输入项 | 说明 | 必需 |
|---|---|---|
| 集数与镜头号 | 用于状态机定位 | 是 |
| 分镜描述 | 景别、运镜、时长、表演要求 | 是 |
| 角色资产引用 | 角色 ID + 资产版本号 | 是 |
| 前一镜头状态 | 收尾帧、色调基线、场景 ID | 是 |
| 输出规格 | 分辨率、帧率、时长、画幅 | 是 |
| 一致性阈值 | 面部/服饰/道具/场景的判定阈值 | 是 |
## 输出
- 镜头素材(含 2~3 条候选)
- 一致性比对报告(面部/服饰/道具/场景/音色)
- 标识证据(显著标识位置 + 隐式元数据回读结果)
- 前置审核结果
- 签核记录与镜头状态机更新记录
## 执行步骤
1. 从镜头状态机读取本镜头任务与前一镜头收尾状态。
2. 加载角色三视图、表情集、色卡与音色,锁定版本号。
3. 按镜头类型选择模型与参数,生成 2~3 条候选。
4. 一致性比对:与三视图比对面部/服饰/道具;与前一镜头收尾帧比对场景与色调。
5. 重生成回路:不通过则调整锚定输入(补参考图、加约束条件)后重生成;连续 2 次不通过升级人工。
6. 音频处理:从角色音色库取音色,按情绪匹配 BGM。
7. 标识处理:脚本渲染显著标识并写入隐式元数据。
8. 前置审核:内容合规、版权、肖像、主创信息真实性。
9. 导演或制片人签核。
10. 导出封装并回读校验。
11. 更新镜头状态机:版本、比对分数、失败原因、成本、责任人。
## 质量标准(DoD)
- 面部/服饰/道具/场景/音色五项一致性判定全部通过,且留存比对分数。
- 显著标识在位;隐式元数据可回读;去标识日志留存不少于六个月。
- 前置审核通过并留痕。
- 主创信息标注真实,无虚假标注。
- 镜头状态机已更新,本镜头可单独重生成而不影响其他镜头。
- 引用效果数据标注口径层级(如"成本 50~60 万 → 不到 20 万"为从业者访谈口径)。
## 常见失败与处理
- 角色换脸:三视图未入上下文或被裁剪 → 重新加载并固化为不可裁剪区。
- 场景漂移:未引用前一镜头收尾帧 → 把收尾帧作为参考图注入。
- 帧间不一致(多出的轮子、错位的轴):缩短单次生成长度,增加逐帧人工精修节点。
- 情感表达不足:不要靠加大特效掩盖,回到剧本与表演设计环节。
- 内容同质化(扎堆玄幻/科幻):在剧本阶段增加题材约束与结构多样性要求。
- 音色跨集漂移:强制从角色音色库取值,禁止临时生成新音色。
- 沦为"幻灯片堆砌":把叙事连贯性纳入判定标准,限制单镜头华丽度权重。
- 前置审核不通过:按审核意见修改后重审,不得绕过。
## 示例
任务:第 7 集第 12 镜头,中景,角色 A 在雨中回身,时长 6 秒,1080P / 30 fps / 竖屏。
输入:角色 A 资产版本 v3.2;前一镜头(SC-0711)收尾帧与色调基线;分镜描述;一致性阈值。
执行:加载三视图 → 生成 3 条候选 → 比对(候选 1、2 服饰偏差超阈值,候选 3 通过)→ 音色库取音色 → 标识渲染 → 前置审核 → 导演签核 → 导出回读 → 更新状态机。
输出:成片素材 + 比对报告(含 2 条失败样本)+ 标识证据 + 签核记录 + 状态机记录(SC-0712 = passed)。 4.3 Implementation Checklist
| No. | Check Item | Layer | Requirement | Notes |
|---|---|---|---|---|
| E-01 | Character asset library established: turnaround views + expression set + color card + vocal timbre | L4 | Required | The core structure of this direction's bottleneck |
| E-02 | Each character asset has a version number and licensing status | L4 | Required | — |
| E-03 | Shot state machine established, supporting independent single-shot regeneration | L4 | Required | Must not be replaced by "re-doing entire segments" |
| E-04 | Cross-episode continuity has a sampling mechanism and sampling ratio | L4 | Required | Sampling ratio determined by the producer |
| E-05 | Character turnaround views and previous-shot state are listed as non-trimmable context | L1 | Required | The main root cause of drift |
| E-06 | All five-stage pipeline tools (script → storyboard → generation → editing → dubbing) in place | L2 | Required | — |
| E-07 | Dynamically invoke models of different parameter scales by shot type | L2 | Recommended | Kunlun Tech adopts this strategy |
| E-08 | Frame extraction, color correction, metadata writing, and label rendering are scripted | L2 | Required | — |
| E-09 | Consistency checking is built into the pipeline, not discovered manually afterwards | L3 | Required | Checking feeds into the regeneration loop |
| E-10 | Regeneration loop supports auto-escalation to human after 2 consecutive failures | L3 | Required | — |
| E-11 | When single-generation length is constrained by model spec, segmented-generation design is in place | L3 | Required | e.g., Veo 3.1 API is 4/6/8 seconds |
| E-12 | Consistency determination covers all five: face/apparel/props/scene/vocal timbre | L5 | Required | Missing any one creates a break-immersion point |
| E-13 | Narrative coherence included in the determination criteria | L5 | Required | Prevents "slide-show pile-up" |
| E-14 | Cost and cycle time included in observation | L5 | Required | Reference: whole series < 200,000 CNY |
| E-15 | AI training-data copyright compliance mechanism established | L6 | Required | Explicitly required by industry standards |
| E-16 | Content pre-release review process established and kept on record | L6 | Required | Not post-hoc sampling |
| E-17 | Conspicuous label rendered by script; implicit metadata readable back | L6 | Required | — |
| E-18 | Credit information truthful, with no false attribution | L6 | Required | Explicitly prohibited by industry standards |
| E-19 | Real natural persons' portraits and voices blocked (incl. face-swapping and voice cloning) | L6 | Required | Hard red line |
| E-20 | De-labeling logs retained for no less than six months | L6 | Required | Article 9 of the Labeling Measures |
5. Summary
The AI online-drama direction exposes the Creative Industries cluster's L4 bottleneck most thoroughly. The reason is simple: it is the only direction that must keep "the same character" consistent across dozens of episodes, hundreds of shots, and a weeks-long production cycle. If any of the five dimensions—face, apparel, props, scene, vocal timbre—drifts, the audience is immediately thrown out of the story.
The industry's answer is not "a better model," but "a better state carrier." In Kunlun Tech's SkyReels design, the most valuable reference is not the 1080P / 60 fps / 180-second spec, but the coupling of the character library with storyboard generation—the storyboard directly references character-library assets rather than re-describing the character each time. This design turns consistency from a "model-capability problem" into an "engineering-structure problem."
The cost and cycle-time data are real: a whole AI short drama dropped from 500,000–600,000 CNY to under 200,000 CNY, Bai Hu compressed from 3 months to 2 weeks, and Shan Hai Qi Jing compressed from 3–6 months to 2 months with cost falling to 1/4. But one must also see the problems the industry has not yet solved: no phenomenon-level hit has yet emerged, severe homogenization, overall profitability not fully validated, and "frequently occurring character-and-scene consistency problems."
On the regulatory side, the China Netcasting Services Association released the Network Audio-Visual AI Short Drama Creation, Production, and Distribution Specification in April 2025, requiring conspicuous labeling of AI-generated content, no false credit attribution, establishment of a training-data copyright compliance mechanism, and pre-release review of AI-generated content. It must be stated truthfully: the official document number and full text of this specification could not be verified on the association's official website; this document cites it only from secondary sources, marked as [To be verified].
Information Gap Statement
| Gap Item | Handling |
|---|---|
| The official document number and full text of the Network Audio-Visual AI Short Drama Creation, Production, and Distribution Specification | Only a secondary page (aiww.cn) was found claiming it was "officially published by the China Netcasting Services Association in April 2025"; the original text, document number, and clauses were not verified on the association's official website → only state the time and institution, mark as [To be verified], and do not fabricate a document number |
| NRTA Measures for the Administration of Micro-drama Development (Draft for Comments) (11 banned-broadcast red lines, per-episode labeling requirement for AI short dramas) | Found only in consolidated securities research reports, said to be released 2026-06, original text not verified → label as "per consolidated securities research reports" and add |
| NRTA "Micro-drama+" action plan, tiered content review system | Same as above, only in consolidated research reports → mark as [To be verified] |
| The GB number of Network Security Technology — Methods for Labeling AI-Generated and Synthetic Content | Confirmed: GB 45438—2025 (a mandatory national standard, published 2025-02-28, effective 2025-09-01, synchronized with the Labeling Measures; sources: National Public Service Platform for Standards Information, TC260 official text) |
| The two figures for SkyReels video-generation specs (180 seconds vs. over 60 seconds) | Both figures exist → this document adopts "1080P / 60 fps / 180 seconds per generation" and notes the other figure |
| The two figures for Bai Hu's per-minute cost (ten-thousand-level vs. under ten thousand) | Both figures exist → present them side by side without choosing one |
| The absolute revenue of Jingying Technology's overseas AI short drama | Only "heat value exceeding 5 million" and "topped the weekly chart" were found; no absolute revenue → not supplemented |
| Distribution-side metrics such as completion rate and retention rate for AI micro-dramas | No public data found → [To be filled] |
6. References
- Procedures for the Labeling of AI-Generated and Synthetic Content — Cyberspace Administration of China, Ministry of Industry and Information Technology, Ministry of Public Security, National Radio and Television Administration, 2025. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- AI Short Dramas: The New Hot Spot Being Chased by Capital — 21st Century Business Herald, 2025-10-23. https://www.21jingji.com/article/20251023/herald/89aa15c81c1bf4ae8b48bf94060ed94e.html
- Will AI "Shuffle" Micro-dramas? — People's Daily, 2026. https://kpzg.people.com.cn/n1/2026/0511/c404214-40717117.html
- Will AI "Shuffle" Micro-dramas? (Hunan Channel reprint) — People's Daily Online, 2026. https://hn.people.com.cn/BIG5/n2/2026/0511/c208814-41577065.html
- AI Short Drama Specification Published (China Netcasting Services Association) — AIWW. Note: secondary source; the official document number has not been verified. https://www.aiww.cn/s/ff3ac3ba84deb97b
- Summary of Kunlun Tech's 2024 Annual Report (SkyReels technology stack) — Securities Times, 2025. https://epaper.stcn.com/pic/202504/26/8eec6686f3528439d13a5db995a8f55b.pdf
- Kunlun Tech Official Website (SkyReels-A1 and other open-source models) — Kunlun Tech. https://www.kunlun.com/en
- Promoting the Construction of the Labeling System Through Multiple Measures to Support the Healthy Development of AI in the New Era (labeling implementation methodology and 6 standard practice guides) — National Internet Emergency Center, 2025. https://www.cac.gov.cn/2025-09/06/c_1758880709361356.htm
- Research Report on the Development of New Media Run by Mainstream Media (2024-2025) — People's Daily Online, 2025. https://sc.people.com.cn/BIG5/n2/2025/1030/c345167-41396739.html
- AGENTS.md Official Website — Agentic AI Foundation (Linux Foundation). https://agents.md/
- Agent Skills Specification — agentskills.io. https://agentskills.io/specification
- Equipping agents for the real world with Agent Skills — Anthropic, 2025 (updated 2025-12-18). https://claude.com/blog/equipping-agents-for-the-real-world-with-agent-skills