产业格局:从模型能力竞争到 Harness 完整度竞争


1. 引言:四个赛道与一个共同变量

1.1. 研究范围与分析框架

图 1-1|Harness 六层能力模型:竞争重心从模型能力到完整度

Harness 六层能力模型:四赛道竞争格局分析框架 依据白皮书第 3 章六层能力模型 · 跨赛道成熟度状态取自本文 6.2 节 L1 · 上下文工程层 相对最强:rules 层级 / 索引 / 压缩在四赛道均成熟 L2 · 工具与执行层 强且快速趋同:沙箱 / MCP / 节点化执行普及 L3 · 编排与控制层 同质化中:多智能体编排 / 计划模式已成标配 L4 · 记忆与状态层(本图重点) 共同短板:资产库无统一基准、记忆持久化薄弱 L5 · 评估与观测层(本图重点) 共同短板:弱评估为主、客观标尺与基准缺位 L6 · 治理与安全层 两极分化:合规驱动型强、工具型弱 结构解读:竞争重心已从「接哪个模型」转向「六层完整度」;L4 / L5 是共同短板,亦是未来两年的差异化窗口。

数据来源:基于本文分析绘制的示意图。

本章基于对四个内容与工具赛道的系统调研,给出 AI Harness 视角下的产业格局全景:

  1. AI IDE:以编码场景为核心的智能体开发环境(Cursor、Claude Code、GitHub Copilot、Codex CLI、Trae 等);
  2. AI Agents 平台:智能体框架、低代码平台与垂直自主智能体(LangGraph、OpenAI Agents SDK、Claude Agent SDK、Dify、Coze、Devin、Manus 等);
  3. AI 图像与内容生成:文生图、图像编辑与设计工具(Midjourney、即梦、可灵、Nano Banana、Leonardo、ComfyUI 等);
  4. AI 漫剧与小说:长程内容生产(即梦 AI 漫剧、可灵、Vidu、白日梦;阅文妙笔、番茄、NovelAI、Sudowrite 等)。

分析框架沿用本白皮书第 3 章的六层能力模型:不把平台当作"工具"罗列功能,而是把每个平台当作一套已经落地的 Harness 实现来解剖——看它在哪一层做了真投入、在哪一层留了缺口、缺口由谁来补。所有厂商经营数据均标注 self-reported 或来源等级;口径冲突处并列呈现并标 。

1.2. 全章核心判断

四个赛道的产品形态、客户群体与商业模式差异巨大,但在 2025 至 2026 年呈现出同一个结构性变化:竞争的重心正从"接了哪个模型"转向"Harness 六层的完整度"。本章将论证:这一转变在每个赛道都有独立证据;L4(记忆与状态层,可复现性与资产持久化)与 L5(评估与观测层)是全行业共同的短板;而中国与海外市场将在合规、部署形态与开源策略三个维度上走出不同的收敛路径。


2. AI IDE 赛道

2.1. 市场全景:采用率饱和,信任度未跟上

AI IDE 是竞争最激烈、也最能验证"模型能力不等于工程可用性"的赛道。三组权威调研数据勾出市场的基本形态:

指标数值来源(注明口径)
开发者采用率84% 正在使用或计划使用 AI 工具(2024 年为 76%);51% 专业开发者每日使用Stack Overflow 2025 开发者调查,2025-07-30,49,000+ 份回答
信任度46% 不信任 AI 输出准确性(2024 年为 31%);仅 3.1% 高度信任Stack Overflow 2025
智能体形态渗透约 31% 已在使用 AI agent;37.9% 不打算用Stack Overflow 2025
主要挫败66% 认为"AI 方案几乎对但不完全对";45.2% 认为调试 AI 代码更耗时Stack Overflow 2025
企业侧采用90% 以上使用 AI(另有 95% 口径);80% 感知个体生产力提升;30% 有单次超 4 小时的长任务DORA 2025,2025-11-12
实测效率资深开发者使用 AI 后实测慢 19%,自评快 20%METR 随机对照试验,2025-07-10

这组数据的含义是:市场教育已经完成,信任建设尚未开始。采用率与不信任率同步上升,说明工具价值已被承认,而产出尚未获得工程上的可预期性。"66% 认为几乎对但不完全对"描述的不是能力问题,而是可验证性问题——恰恰是 Harness 的 L5 层要解决的。

2.2. 代表平台对比矩阵

平台开发商形态开源Harness 六层的显著特征
CursorAnysphereVS Code 分支 IDEL1 规则层级 + 代码库索引;L6 团队市场 + SSO + 审计日志 + 代码追踪 API
Claude CodeAnthropic终端 + IDE + WebL2 细粒度沙箱 + 凭据掩码;L1 就近规则 + 技能渐进披露 + 压缩;L3 计划模式 + 子智能体 + 钩子
GitHub CopilotMicrosoft / GitHub多 IDE 插件 + CLI + WebL6 组织策略 + 内容排除 + 审计 + IP 赔付;L5 平台侧代码审查与用量看板;共享 agentic harness(Copilot SDK 单一组件)驱动多条产品线
Codex CLIOpenAI终端 + IDE 扩展 + Web是(Apache 2.0)L2 OS 级沙箱与审批正交配置;L1 AGENTS.md 层级;L3 非交互执行与会话恢复
WindsurfCognitionVS Code 分支 IDEL4 记忆系统 + 检查点回退为差异化核心;L3 计划模式 + 并行会话
Trae字节跳动VS Code 分支 IDE + WebL1 上下文压缩 + 规则工程;L3 多智能体;以价格与国内模型生态切入

关于该赛道的定调性证据:GitHub 官方在 2026 年明确表述"模型提供原始智能,而 harness 决定这份智能被应用得多有效"(the harness shapes how effectively that intelligence is applied),并公开了其 agentic harness 的受控对比方法学;Terminal-Bench 官方亦明确其排行榜"排的是系统而非模型"。官方榜单与头部厂商的双重确认,使"Harness 完整度决定产品差距"从行业直觉升级为可引用的工程结论(方法论细节见本白皮书第 5 章)。

2.3. 竞争焦点

AI IDE 赛道的竞争焦点已明确迁移:

  1. L1 上下文工程是第一竞争维度。一个中型仓库几十万行代码,而上下文窗口再大也装不下。rules 文件层级(就近加载)、代码库语义索引、上下文压缩(compaction)三件套决定单次任务能承载多复杂的需求。
  2. L6 治理是准入项而非加分项。自主度越高,事故爆炸半径越大——2025-08-11 的 Amazon Q 提示注入事件(恶意指令试图诱导智能体删除 AWS 资源)是"被处理的外部内容可以直接成为指令"这一结构性弱点的首次大规模公开暴露。文件系统隔离与网络隔离缺一不可,凭据默认不可见,破坏性命令需硬编码拦截。
  3. 正面证据表明治理与自主性是正和:Anthropic 官方披露其沙箱化使内部权限提示减少 84%,同时提升安全性——约束让规模化速度成为可能。

2.4. 商业模式:三种计费逻辑

计费逻辑代表优点风险
订阅额度制Claude Code(随 Claude 订阅)、Codex CLI(随 ChatGPT 订阅)成本可预测高峰期撞限,长任务被迫中断
信用点 / 用量制Cursor(额度池 + 按用量)、GitHub Copilot(2026-06-01 起 AI Credits,1 credit = 0.01 美元)一次对话与一时代理会话不再同价,更公平成本不确定性,需盯消耗速率
配额刷新制Windsurf(2026-03 起按日 / 周刷新配额)便于预算复杂任务单次上限受约束

三种逻辑正向"订阅 + 用量"混合收敛。计费模式的迁移意味着采购方需要新的管理能力:组织级消耗速率观测与预算告警。GitHub Copilot 的案例还显示产品线整合趋势——其 agentic harness 作为 Copilot SDK 的单一共享组件,同时驱动 CLI、App、代码审查等多条体验,"Harness 即平台资产"的定位日益清晰。


3. AI Agents 平台赛道

3.1. 市场全景:三层供给结构

AI Agents 平台供给呈清晰的三层结构,且三层遵循不同的 Harness 完整度逻辑:

代表供给逻辑Harness 覆盖
代码框架LangGraph、OpenAI Agents SDK、Claude Agent SDK、CrewAI、Microsoft Agent Framework、AutoGen面向工程团队,开源或免费,主要覆盖 L2 / L3L4 / L5 / L6 大多留白,由使用者自建或外接(如 LangSmith 提供追踪与评估)
低代码平台Dify、Coze(扣子)、腾讯元器面向业务团队,可视化编排 + 托管运行时L1 至 L3 产品化封装,L5 观测与 L6 租户治理随企业版分级
垂直自主智能体Devin(Cognition)、Manus面向端到端任务交付,按工作量计费六层闭环封装于产品内,黑盒程度最高

增长信号方面(均为厂商侧或第三方口径,需谨慎引用):Devin 的 ARR 据第三方汇编从 3,700 万美元(2025-05)增至 4.92 亿美元(2026-05),估值口径从 102 亿美元升至 260 亿美元——来源为第三方汇总(低置信),,本白皮书不作为行业规模测算依据;Devin 官方披露的运营里程碑(上线 8 个月处理 147 万亿 token、驱动 8,000 万虚拟计算机)同样为 self-reported。框架侧的确定性事实是:核心框架普遍开源免费(LangGraph、CrewAI、AutoGen 均 MIT),收入转向托管与观测层(LangSmith 免费档 5,000 traces/月,Developer 档 39 美元/用户/月)。

3.2. 代表平台对比矩阵

平台类别开源 / 许可计费形态治理与合规亮点
LangGraph代码框架MIT 开源框架免费,仅付模型 API 费检查点(checkpointing)能力为同类框架中最完整
OpenAI Agents SDK代码框架开源框架免费;默认追踪看板需 OpenAI 平台账号与模型厂商运行时深度耦合
Claude Agent SDK代码框架闭源 SDK无许可费,按 Claude API token 计费内置 compaction、子智能体上下文隔离、hooks
CrewAI代码框架MIT 开源AMP 平台免费档 50 次工作流执行/月,Enterprise 定制企业档提供 FedRAMP High、SSO、RBAC(官方口径)
Dify低代码平台开源(部分) + 商业版免费 + 云订阅私有化部署成熟;2026-03 获 3,000 万美元 Series Pre-A(第三方转述)
Coze(扣子)低代码平台国内版闭源、海外版有开源版本(口径冲突)免费档 + 按 token 超量企业版订阅约 2 万至 20 万元/年(第三方评测口径)
Devin垂直自主智能体Core 20 美元/月 + 2.25 美元/ACU(1 ACU 约 15 分钟主动工作);Team 500 美元/月含 250 ACUEnterprise 档提供 VPC 部署、SAML/OIDC SSO、集中管控(官方口径)
Manus垂直自主智能体否(L1/L4/L5 机制闭源,不可验证)Freemium + credit 消耗制(Starter 20 美元/月起)公开治理细节最少

3.3. 竞争焦点与商业模式

该赛道的竞争焦点呈现两个方向:

  1. 框架层的竞争在"控制平面"。当编排原语(循环、子智能体、工具调用)趋同后,差异化转向状态管理(LangGraph 的 checkpointing)、上下文治理(Claude Agent SDK 的 compaction)与可观测性(LangSmith)。这与六层模型中 L3 与 L4 的权重上升一致。
  2. 平台层与垂直层的竞争在 L6 与成本计量。企业采购的核心问题已从"能不能编排"变为"权限怎么管、审计怎么导出、成本怎么封顶"。Devin 按 ACU(主动工作时间)计费、Manus 按 credit 计费、Dify / Coze 按执行与 token 计费——三家形态殊途同归:把不可预测的 token 消耗翻译成可预算的工作量单位,这本质上是 L5 观测能力的产品化。

商业模式的结构性事实是:框架层不挣钱,挣钱的是运行时、观测与治理。这一分布印证了本白皮书的判断——价值正在从模型接入层向 Harness 层转移。

当日增量(2026-09-12;2026-09-13 修正):OpenAI Agents API 于 2026-09-10 进入公测(A级,OpenAI 官方 Changelog 口径;2026-09-12 快照曾依媒体报道记为 09-11),把既有 Agents SDK 的编排能力托管化——这与 AWS Bedrock AgentCore GA(2026-06)构成同一方向的两次落子,标志着该赛道的竞争焦点正从"编排原语是否好用"转向"托管运行时是否可治理"。详见调研库 03-市场研究/02-AI-Agents组/21-openai-agents-api.md


4. AI 图像与内容生成赛道

4.1. 市场全景

图像赛道是四个赛道中商业化最成熟的。以美图公司为例(港股公告口径,极高置信):2026 上半年总收入 22.1 亿元(同比 +22.1%),其中影像与设计产品收入 17.7 亿元(+30.9%,占 80%);经调整归母净利润 6.5 亿元(+39.5%);MAU 2.82 亿,付费订阅用户超 1,844 万(+19.7%),订阅渗透率 6.5%;生产力应用 MAU 3,300 万(+43.5%),付费订阅 235 万,ARR 约 6.2 亿元。美图的可观测设计尤其值得注意:其以"AI 算力点消费环比增幅"(Q1、Q2 均超 46%)作为价值命中率代理指标——用消费深度衡量价值命中率,是 L5 观测在内容赛道的罕见成熟实践

供给侧的结构变化同样显著:Google 以 Nano Banana(Gemini 系图像模型)把顶级图像能力注入通用 API(初代 1K/2K 单图 0.039 美元、Pro 系列 1K/2K 0.134 美元,官方定价页截图整理口径),国内通义万相以 0.2 元/张的 API 定价(阿里云官方)展开价格竞争——基础图像生成正在快速商品化,利润向工作流、资产管理与行业方案(Harness 层)迁移

4.2. 代表平台对比矩阵

平台开发商形态付费形态Harness 六层定位
MidjourneyMidjourneyWeb / Discord订阅 10 至 120 美元/月强模型弱 Harness 的代表:L5 依赖人眼主观评测与社区反馈,未见官方回归集机制
Nano Banana(Gemini 系)GoogleAPI + 产品内嵌按 token / 按图计费以 API 分发能力,Harness 由上层应用承担
即梦 AI字节跳动SaaS + APIC 端会员(两套第三方口径冲突)模型 + 创作工具 + 抖音分发闭环
可灵 AI快手SaaS + API国际站 6.99 至 127.99 美元/月(官方页);国内站为二次来源口径元素引用 + 系列生成支撑一致性;商用权按档位
Leonardo.aiLeonardoWeb + API订阅 12 至 60 美元/月 + PAYG API(新账户赠 5 美元,最高 10 并发)三类容量约束显式文档化 + Webhook + 定价计算器,工程化程度高
RunwayRunwayWeb + API订阅 12 至 76 美元/月(年付);API 0.01 美元/credit视频 / 图像混合工作流,Adobe 集成
ComfyUIComfy Org(开源社区)本地节点工作流免费(GPL-3.0)60,000+ 节点生态;工作流 JSON 即状态、可版本化;六层全部可编程自建
WeShop / 美图设计室美图等SaaS订阅 + 算力点Agent Teams 多智能体协同 + 创作资产复用(self-reported)

4.3. 竞争焦点与商业模式

该赛道的两个结构性观察:

  1. Harness 化程度的分水岭在 L3 与 L4。美图设计室(多 Agent 协同 + 创作资产复用 + 算力点计量)、ComfyUI(工作流即 JSON 图、可版本控制、可 SDK 化)、Leonardo(容量约束显式化 + Webhook + 成本预估)代表第三代形态;Midjourney、妙鸭代表"强模型弱 Harness"——模型质量领先但工程承载薄弱,且妙鸭已因此出局(详见 4.4 节)。
  2. 商业模式从单次付费转向订阅 + 消耗混合。Midjourney 纯订阅、Runway 纯 credit、Leonardo 订阅 + PAYG 双轨、美图订阅 + 算力点。消耗制的前提是精准的用量计量与告警——又一个 L5 能力成为商业基础设施的实例。

4.4. 典型失败样本:妙鸭相机

妙鸭相机是本赛道唯一具备完整生命周期的负面样本,也是"强模型弱 Harness"的教科书案例:

  • 爆发:2023-07-17 上线,2023 年 8 月登顶 App Store 总榜,日活突破 60 万,高峰期 4,000 余人排队出片;技术底座为阿里大文娱自研"提香"模型。
  • 工程短板:L2 / L3 / L5 三层几乎空白——每次交互都是原子操作(上传、等待、选模板、出片),无中间状态可干预、无参数可调节;用户唯一的质量修正入口是"更像我一点"的相似度微调。L6 在早期发生用户协议信任事故(官方道歉并修订)。
  • 商业短板:数字分身一次性 9.9 元(推广价)买断 + 10 张成片,是典型引流品定价,缺乏后续消耗场景与订阅锚点;对比同类身份类产品(可灵按会员档位限制 30 至 500 个主体)缺少资产运营纵深。
  • 技术路线固化:训练式身份管线(每用户 ≥20 张照片、完整 per-user 训练、数小时排队)在 2024 至 2025 年零样本身份注入方案(InstantID、PuLID)成熟后未完成迁移,成本结构被固化。从 Harness 视角看:训练式路线把身份资产做成了"产品内的私有状态",不可迁移、不可组合、不可审计;零样本路线把身份资产做成"可跨管线流通的上下文",获得真正的资产流动性。
  • 结局:团队于 2025 年 9 月底正式解散(该事实的权威信源目前仅有第三方工具站表述),产品仅维持最低限度运营。

妙鸭的失败命题值得写进所有内容类产品的立项评审:在生成质量已经足够好的前提下,决定产品生死的不是模型,而是承载模型的 Harness


5. AI 漫剧与小说赛道

5.1. AI 漫剧:产能过剩与确定性稀缺

AI 漫剧是 AIGC 内容产业中工程化程度提升最快、也最早暴露"模型能力不等于生产可用"矛盾的场景。2026 年上半年的行业数据(DataEye 报告与中国网络视听协会数据,经二手转引,原始报告全文未交叉验证):

指标数值说明
微短剧上线总量36.7 万部,AI 内容占比超 74%中国网络视听协会数据(转引)
抖音端原生新增 AI 剧及漫剧22.19 万部,累计播放 5,157.38 亿次
爆款率播放量破亿仅 1,055 部,约 0.47%DataEye
回本率不足 1.3%,约每 77 部 1 部回本按 5,000 万播放盈亏平衡线测算
上线速度平均每 36 秒一部DataEye
万次播放收入从峰值 30 至 100 元跌至 5 至 10 元,跌幅超九成DataEye
市场规模2026 年预计突破 400 亿元(+138%);出海预计超 40 亿美元中研普华 / DataEye 转引,

产能以每 36 秒一部的速度堆积、爆款率不足 0.5%、单价跌去九成——行业从"抢产能"进入"抢确定性"。而确定性恰恰是模型给不了的:同一个角色在第 1 集和第 47 集长得不一样、同一场景在两个镜头之间光影跳变、剧情状态跨会话续写时丢失。这些问题没有一个能靠换更强的模型解决,只能靠 Harness 的状态管理(L4)、编排(L3)与回归评估(L5)解决。

5.2. AI 漫剧代表平台对比矩阵

平台开发商竞争抓手L4(记忆 / 资产)实现L6 治理
即梦 AI字节跳动模型 + 工具 + 抖音分发闭环;多镜头叙事角色特征稳定保持;资产库口径未公开真人分身认证;2026-04-28 因未有效落实 AI 生成合成内容标识被网信部门查处(公开处罚事件)
可灵 AI快手智能分镜 + 元素引用;海外收入占比约 70%(第三方口径)系列生成 + 一致性延长商用权按档位解锁
PixVerse爱诗科技CLI / Skills 兼容 Claude Code、Codex 等编码智能体;覆盖 175 个国家Character 一致性角色 + Team 资产库Team Plan RBAC + 积分上限 + 用量分析
海螺 AIMiniMaxH3-Context-IR 上下文压缩(约 100k token 压缩至约 4k)+ 开源权重@ 引用系统 + In-Context Regeneration开源权重可自审;完成多款国产芯片适配
Vidu生数科技参考生视频"万物可参"主体库(角色 / 道具 / 场景),最多 7 张参考图SaaS / MaaS 分层
白日梦 AI光魔科技永久角色库,宣称跨集相似度 95%+角色库为核心卖点公开治理信息缺失
ComfyUI开源社区60,000+ 节点、完全本地离线角色 LoRA + 风格锚点;工作流 JSON 即状态、可版本化GPL-3.0,标识管线可自建(合规自建范例)
豆包 / Seedance字节 SeedMaaS API 输出无原生资产库,需上层自建真人人脸禁用 + 数字分身真人校验(公开的合规最严样本之一)

该赛道的分层规律清晰:L1 至 L3(四模态输入、智能分镜、多镜头叙事、节点工作流)在 2026 年已成为头部平台标配、高度同质化;L4 尚未同质化——从"最多 7 张参考图"(Vidu)到"约 100k token 压缩至约 4k token"(海螺 H3)再到"工作流 JSON 即状态"(ComfyUI),实现路径完全不同。谁把 L4 做扎实(角色锚定、资产库、镜头状态机),谁就能把 AI 漫剧从"能生成"推进到"能连续生产一百集"。

5.3. AI 小说:长程状态管理竞赛

AI 小说是六层模型中 L4 与 L1 压力最大的内容场景:长篇连载的核心难题是"几十万字后仍不崩人设、不忘伏笔"(网文行业以"吃书"命名这一必现故障)。行业收敛出三种 L4 范式:

范式机制代表实现
A · 结构化滚动摘要上下文超长时自动压缩为结构化摘要,保留关键线索Claude 长文工作流的滚动摘要层
B · 关键词触发条目库设定拆为条目,仅在相关时注入上下文NovelAI Lorebook;AI Dungeon Story Cards
C · 版本化状态快照每章完成后抽取状态并存档为版本化快照,按需回溯Claude Book 的 state/current/ 符号链接 + 逐章归档;灵蟹创作的"项目宪法"

平台格局呈现"中文平台与海外平台同题异解":阅文妙笔以"妙笔通鉴"切入"第二大脑"定位(专做伏笔挖掘与细节检索),其披露的业务指标为妙笔日活增长超一倍、日均 token 消耗增长超 90%、作家周使用率超 75%(厂商 self-reported);番茄小说以最强 L6 著称(强制 AI 申报、保底书禁用扩写续写、低质内容机器判定);NovelAI 以 Lorebook 与隐私优先(XSalsa20 客户端加密、不训练)立足订阅市场;Sudowrite 以 Story Bible + 读前 20,000 词 + 25 文档章节连续性支撑英文长篇。社区侧,GitHub 开源项目 Claude-Code-Novel-Writer(v4.1,2026-08-21)以 AGENTS.md + 7 角色 + 5 技能的多智能体编排,给出了"权威源 / 派生索引分离"(手稿为权威源,状态跟踪可重生成)的工程范本。

5.4. 治理是前置条件而非事后检查

内容双赛道的 L6 已经从"加分项"变为"上线前置条件":

  • 法规基线:《人工智能生成合成内容标识办法》(2025-03-07 印发,2025-09-01 施行)要求视频在起始画面添加显著显式标识、在文件元数据中写入隐式标识;强制性国标 GB 45438—2025 进一步量化:显式标识文字高度不低于画面最短边长度的 5%、正常播放速度下持续不少于 2 秒,元数据字段含 Label / ContentProducer / ProduceID 等。
  • 执法实例:2026-04-28,即梦 AI 因未有效落实标识规定被网信部门依法查处——即便一线大厂平台,标识管线也可能断裂。
  • 平台治理分化:番茄 2025-09-23 起强制作者申报"是否使用 AI";AI 续写、扩写仅限非保底书作者使用(按合同类型分级授权);晋江文学城仅允许文字校对、创意要素辅助、创意粗纲辅助三种 AI 使用场景。海外平台则主要围绕版权归属(是否主张作品权利、是否用用户作品训练)与隐私(加密、不训练承诺)展开治理。

对工程团队的直接含义:合规标识必须做成生成管线末端的自动节点,与分辨率、帧率同为产物的一等属性,不能依赖人工后期补加。


6. 跨赛道共性判断

6.1. 竞争正从模型能力转向 Harness 完整度

四个赛道的独立证据汇总:

赛道"模型能力决定论"失效的证据竞争重心的迁移
AI IDE同一模型在不同脚手架下分差可达两位数百分点;GitHub 官方明确 harness 决定智能的应用效率;Terminal-Bench 排行榜"排系统不排模型"模型列表长度 → rules 层级、沙箱粒度、审批策略、可观测性
AI Agents 平台编排原语趋同后,框架差异化转向状态管理与可观测;框架开源免费、收入来自托管与治理编排能力 → 控制平面(checkpointing、上下文治理、审计、成本计量)
AI 图像基础生成能力快速商品化(0.2 元/张级 API 定价);"强模型弱 Harness"的妙鸭出局生成质量 → 工作流、资产管理、用量计量
AI 漫剧与小说产能过剩 + 爆款率 0.47%;一致性、连续性问题无法靠换模型解决单次生成上限(L1 至 L3)→ 连续生产上限(L4)

四个赛道以完全不同的方式收敛到同一个结论:当模型能力越过"可用"阈值后,产品之间的可感知差距主要由 Harness 六层的完成度决定。这与本白皮书第 5 章的评测发现互为表里——既然评测分数本身是"模型 × Harness"的复合产物,市场竞争自然也以同样的复合方式展开。

6.2. L4 与 L5 是全行业共同短板

把四个赛道的六层成熟度做横向归并,可以得出一个高度一致的结论:

跨赛道状态证据
L1 上下文工程相对最强rules 文件、检索、压缩在各赛道均有成熟实践(AI IDE 三件套、Lorebook、参考图注入、H3-Context-IR)
L2 工具与执行强且快速趋同沙箱、MCP、CLI 化、节点化执行原语普及
L3 编排与控制强、同质化中多智能体编排、计划模式、节点 DAG 已成标配
L4 记忆与状态共同短板AI 漫剧的资产库形态各异且无统一基准;小说的"吃书"是必现故障;妙鸭的身份资产不可迁移;低代码平台的长期记忆普遍薄弱
L5 评估与观测共同短板多数平台仅有"多候选重生成"式弱评估;本白皮书第 5 章证明 AI SRE / DevOps / 多智能体方向连公开基准都缺位;内容质量无客观标尺
L6 治理与安全两极分化合规驱动型(内容赛道、企业版)强,工具型(Midjourney、部分框架)弱

L4 短板的共性根因是:模型本身无状态,所有"一致性"问题本质都是跨调用、跨会话、跨周期的状态注入与持久化问题——角色一致(漫剧)、设定一致(小说)、身份资产可迁移(图像)、会话可恢复(IDE)是同一类工程问题在不同赛道的投影。L5 短板的共性根因则是:内容与开放域任务的产出缺乏客观判定标尺,且公共基准缺位。因此,L4 与 L5 既是全行业的短板,也是未来两年最明确的差异化机会窗口——谁先把"可复现性与资产持久化"(L4)和"效果评估"(L5)做成产品能力而非营销表述,谁就能在各自的赛道拿到 Vidu 主体库、海螺 Context-IR、美图算力点观测那样的差异化位置。


7. 中国与海外市场差异

7.1. 生态结构差异

  • 海外:以"模型厂商 + 订阅制 + API 经济"为主轴。分配逻辑是产品力与全球市场覆盖(PixVerse 覆盖 175 个国家、可灵海外收入占比约七成为其增长引擎),定价以美元订阅与用量计费为主。
  • 中国:以"大厂渠道闭环 + 免费策略 + 分账生态"为主轴。即梦对接抖音、可灵对接快手,模型 + 工具 + 分发一体;部分国内工具以完全免费策略切入(如字节 Trae、腾讯元宝"目前完全免费使用,暂无收费计划"的公开表态);内容侧形成独特的分账体系(漫剧按"时长 × 单价 × 类型系数 × 版权系数"分成)与保底激励(2026 年抖音集团全年真人短剧保底预算超 15 亿元)。

7.2. 合规环境差异

  • 中国:合规是明确的硬约束且已有执法实例——《人工智能生成合成内容标识办法》与 GB 45438—2025 构成量化基线,即梦 2026-04-28 被查处是标志性事件;内容平台另有备案门槛(2026 年 1 月重点微短剧备案门槛由 100 万元提高至 300 万元)与题材审核。
  • 海外:EU AI Act 的透明度义务自 2026-08-02 起适用于面向公众的 AI 生成文本标注(开源许可不构成豁免);美国以平台自律与诉讼驱动为主。总体上,中国是"事前标识 + 平台核验"模式,欧盟是"透明度义务"模式,美国是"事后追责"模式。

7.3. 部署形态差异

  • 中国:信创与国产芯片适配是真实需求——MiniMax H3 开源发布首日即完成华为昇腾、摩尔线程、沐曦、海光、昆仑芯、天数智芯、壁仞等适配;私有化部署在企业级采购中权重高(Dify 的私有化能力是其核心卖点)。
  • 海外:云优先、SaaS 优先,本地部署主要出现在受监管行业与开源社区(ComfyUI 的完全本地离线能力使其同时满足两类需求)。

7.4. 开源权重策略差异

  • 中国厂商更积极地把模型权重开源:MiniMax H3(视频模型开源权重)、通义 Qwen 系(编码与图像)、智谱 GLM 系(NovelAI 的 Xialong 即基于 GLM-4.6 微调)——开源权重成为获取开发者生态与国产芯片适配的杠杆。
  • 海外头部厂商分层处理:OpenAI 将 Codex CLI 以 Apache 2.0 开源(CLI 层开源、模型闭源);框架层普遍 MIT(LangGraph、CrewAI、AutoGen);前沿模型权重基本闭源。总体上,中国呈现"权重开源换生态"策略,海外呈现"工具与框架开源换标准、权重闭源保利润"策略

7.5. 差异总览

维度中国市场海外市场
生态结构大厂渠道闭环(模型 + 工具 + 分发一体);免费策略切入;分账与保底生态模型厂商 + 订阅制 + API 经济;全球市场覆盖;美元订阅与用量计费
合规事前标识 + 平台核验(标识办法、GB 45438—2025、备案门槛),已有执法实例欧盟透明度义务(EU AI Act,2026-08-02 适用);美国以平台自律与事后追责为主
部署形态信创与国产芯片适配为真实需求;私有化部署权重高云优先、SaaS 优先;本地部署集中于受监管行业与开源社区
开源策略权重开源换生态(MiniMax H3、Qwen 系、GLM 系)工具与框架开源换标准(Codex CLI、LangGraph、AGENTS.md),权重闭源保利润
治理关注点内容标识、备案、题材审核、训练数据授权(番茄协议风波)版权归属、隐私承诺(加密、不训练)、青少年与内容政策
内容变现分账体系 + 保底激励 + 广告投放(投放约占制作链路成本 70%)订阅、打赏、版权授权;分发以自有渠道与社区为主

需要强调:差异是程度差异而非本质差异。两个市场在"竞争重心从模型能力转向 Harness 完整度"(6.1 节)与"L4 / L5 共同短板"(6.2 节)两个结构性判断上完全一致——差异体现在收敛的速度与约束条件,而非方向。


8. 格局演变的三种可能路径

8.1. 路径一:平台收敛

大厂依托模型、渠道与算力的一体化优势,把 Harness 六层全部产品化,中小玩家退化为模型供应商或垂直插件。支持信号:GitHub 用单一 agentic harness 统一多条产品线;即梦 / 可灵的"模型 + 工具 + 分发"闭环;美图以平台级算力点体系整合多产品线。约束条件:L4 与 L5 恰恰是大厂最难做深的层——资产持久化与效果评估需要长期客户陪伴而非流量打法,垂直玩家仍有机会守住在位优势。

8.2. 路径二:垂直深耕

行业知识与状态层成为护城河,通用平台难以替代。支持信号:AI 漫剧的连续生产能力取决于行业特有的资产库与分镜状态机;网文的"项目宪法"与伏笔台账依赖领域语义;企业级 Agents 采购中 L6(权限、审计、私有化)权重持续上升。妙鸭的失败从反面支持该路径——没有状态层与运营纵深的"通用爆品"生命周期极短。约束条件:垂直市场规模有限,难以摊薄基础模型研发成本,最终多依赖大厂模型 API,议价能力受制于人。

8.3. 路径三:开源生态侵蚀

开源工作流与开源权重从两端挤压闭源 SaaS:ComfyUI 以 60,000+ 节点生态覆盖专业生产,Codex CLI 证明闭源厂商也会主动开源 Harness 层以争夺标准,中国厂商的开源权重策略持续降低自建门槛。若 MCP 等开放协议与 AGENTS.md 等开放标准继续渗透(AGENTS.md 已被 60,000+ 开源项目采用,由 Linux Foundation 旗下 Agentic AI Foundation 托管),六层能力将越来越多地以"可组装的开源组件"形态存在,闭源平台被迫向"托管与治理"(L5 / L6)收窄盈利面。约束条件:开源方案默认缺失评估与治理,企业自建成本高——这反过来又给了商业产品以 L5 / L6 为核心的生存空间。

8.4. 观察指标

三条路径并非互斥,最终格局更可能是三者的混合。对格局走向的判断可依据以下可观察指标:

观察指标指向路径一(收敛)指向路径二(深耕)指向路径三(开源侵蚀)
六层能力的开源替代速度无关快(L1 至 L3 陆续开源化)
企业采购的决策重心平台品牌与集成领域状态层与合规能力可组合性与退出成本
头部厂商 Harness 层的开放程度封闭闭环部分开放(协议兼容)全面开放(标准输出)
L4 / L5 能力的产品化进度平台内置且绑定垂直行业深度绑定开源组件 + 商业托管

9. 总结

本章基于四个赛道的系统调研,给出三点结论:

  1. 竞争正从"模型能力"转向"Harness 完整度"。AI IDE 赛道由官方榜单与头部厂商双重确认("排系统不排模型"、harness 决定智能的应用效率);Agents 赛道框架免费、收入流向托管与治理;图像赛道基础生成快速商品化、利润流向工作流与资产管理;内容双赛道产能过剩、确定性稀缺。四个赛道以不同方式收敛到同一结构性变化,其评测侧的镜像就是本白皮书第 5 章的发现——分数从来是"模型 × Harness"的复合产物。
  2. L4(记忆与状态)与 L5(评估与观测)是全行业共同短板。模型无状态,一切"一致性"都是状态管理问题;内容与开放域任务缺乏客观标尺,公共基准在多个高价值方向缺位。妙鸭相机的完整生命周期失败(强模型、弱 Harness、无状态资产流动性、无评估闭环)与 AI 漫剧 0.47% 的爆款率,从正反两面标定了这两层短板的商业代价;而 Vidu 主体库、海螺 Context-IR、美图算力点观测等案例则标定了补齐短板的差异化回报。
  3. 中国与海外市场将在三条路径的混合演化中走出不同形状:中国以合规硬约束、渠道闭环与开源权重策略为特征,海外以订阅与 API 经济、透明度义务与工具开源为特征。无论哪种形状,格局的最终决定变量是同一个——谁能把模型的不确定性转化为工程上的可预期性,谁就掌握下一个竞争周期的定价权

信息缺口声明

  1. Devin 的 ARR 与估值数据(3,700 万至 4.92 亿美元、102 亿至 260 亿美元)来源为第三方汇总(低置信),全部标注 ,未用于任何行业规模测算。
  2. 妙鸭相机团队解散(2025-09 底)的权威信源目前仅有第三方工具站表述,未定位到科技媒体报道或工商信息佐证,。
  3. AI 漫剧市场规模(400 亿元 / 40 亿美元出海)与爆款率数据主要来自 DataEye 报告经二手转引(百度百科、KOCPC 等),未获取原始报告全文交叉验证。
  4. 各平台 C 端定价存在多套冲突口径(即梦、可灵国内站、美图设计室、Coze 企业版、Vidu 等),本白皮书仅采信官方页可核验口径,其余标注 或不予引用。
  5. 可灵 3.0 的发布时间与"最长 3 分钟"、白日梦"跨集相似度 95%+"等仅见二手来源,。
  6. 阅文妙笔业务指标(日活 +100%、token +90%、周使用率超 75%)与美图设计室 Agent Teams 效果均为厂商 self-reported 口径。
  7. AI 小说与图像赛道多数平台无官方披露的可验证效果指标;第三方测评数据多为营销软文,未予引用。
  8. 本白皮书未采用任何国内大厂 AI 编码提效的自媒体数据作为硬数据(该类数据在本工程检索中未找到一手出处)。

10. 参考资料

  1. 2025 Stack Overflow Developer Survey — Stack Overflow,2025-07-30。https://survey.stackoverflow.co/2025/
  2. State of AI-assisted Software Development 2025 — DORA / Google Cloud,2025-11-12。https://dora.dev/research/2025/dora-report/
  3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity(随机对照试验) — METR,2025-07-10。https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  4. Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks — GitHub Blog,2026。https://github.blog/ai-and-ml/github-copilot/evaluating-performance-and-efficiency-of-the-github-copilot-agentic-harness-across-models-and-tasks/
  5. Terminal-Bench 官方站与 Leaderboard — Stanford / Laude Institute,2025 至 2026。https://www.tbench.ai/
  6. AGENTS.md 官方站 — Agentic AI Foundation(Linux Foundation 旗下)。https://agents.md/
  7. Sandboxing: a safer and more autonomous approach — Anthropic,2025。https://www.anthropic.com/engineering/claude-code-sandboxing
  8. MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities — MiniMax 官方博客。https://www.minimax.io/blog/minimax-h3
  9. 美图公司 2026 年中期业绩公告 — 香港交易所披露易(hkexnews)。https://www.hkexnews.hk/
  10. 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — 2025-03-07 印发,2025-09-01 施行。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  11. 强制性国家标准 GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — 全国网络安全标准化技术委员会。https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
  12. DataEye《2026 上半年 AI 剧 / 漫剧数据报告》 — DataEye研究院(经 KOCPC 等二手转引)。https://en.kocpc.com.tw/archives/25203
  13. 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — 澎湃新闻,2026。https://www.thepaper.cn/newsDetail_forward_33993493
  14. 中国青年报《"妙笔通鉴""漫剧助手"发布,AI 赋能网文创作和 IP 改编》 — 中国青年报,2025。https://new.qq.com/rain/a/20251017A08J6J00
  15. GitHub — forsonny/Claude-Code-Novel-Writer(Multi-Agent Novel Writer v4.1) — 2026。https://github.com/forsonny/Claude-Code-Novel-Writer
  16. Apache Software Foundation《AI Agents and AGENTS.md Files (Draft)》 — ASF legal-discuss,2026。https://www.apache.org/legal/generative-tooling-agents.html
  17. 通义万相与 Qwen 系模型阿里云官方定价页 — 阿里云。https://help.aliyun.com/
  18. Leonardo.ai 官方定价与 API 文档 — Leonardo。https://leonardo.ai/pricing
  19. Runway 官方定价页 — Runway。https://runwayml.com/pricing
  20. Comfy 官网 — Comfy Org。https://comfy.org/

Industry Landscape: From Model-Capability Competition to Harness-Completeness Competition


1. Introduction: Four Tracks and One Common Variable

1.1. Research Scope and Analytical Framework

图 1-1|Harness 六层能力模型:竞争重心从模型能力到完整度

Harness 六层能力模型:四赛道竞争格局分析框架 依据白皮书第 3 章六层能力模型 · 跨赛道成熟度状态取自本文 6.2 节 L1 · 上下文工程层 相对最强:rules 层级 / 索引 / 压缩在四赛道均成熟 L2 · 工具与执行层 强且快速趋同:沙箱 / MCP / 节点化执行普及 L3 · 编排与控制层 同质化中:多智能体编排 / 计划模式已成标配 L4 · 记忆与状态层(本图重点) 共同短板:资产库无统一基准、记忆持久化薄弱 L5 · 评估与观测层(本图重点) 共同短板:弱评估为主、客观标尺与基准缺位 L6 · 治理与安全层 两极分化:合规驱动型强、工具型弱 结构解读:竞争重心已从「接哪个模型」转向「六层完整度」;L4 / L5 是共同短板,亦是未来两年的差异化窗口。

数据来源:基于本文分析绘制的示意图。

This chapter, based on systematic research across four content and tool tracks, presents a panoramic view of the industry landscape from the AI Harness perspective:

  1. AI IDE: agent development environments centered on coding scenarios (Cursor, Claude Code, GitHub Copilot, Codex CLI, Trae, and others);
  2. AI Agents Platforms: agent frameworks, low-code platforms, and vertical autonomous agents (LangGraph, OpenAI Agents SDK, Claude Agent SDK, Dify, Coze, Devin, Manus, and others);
  3. AI Image and Content Generation: text-to-image, image editing, and design tools (Midjourney, 即梦, 可灵, Nano Banana, Leonardo, ComfyUI, and others);
  4. AI Comic-Drama and Novels: long-horizon content production (即梦 AI comic drama, 可灵, Vidu, 白日梦; 阅文妙笔, 番茄, NovelAI, Sudowrite, and others).

The analytical framework follows the Six-Layer Capability Model from Chapter 3 of this whitepaper: instead of treating platforms as "tools" and listing their features, each platform is dissected as a Harness implementation that has already been deployed in production — seeing where it made genuine investment, where it left gaps, and who fills those gaps. All vendor operating data are annotated as self-reported or with a source grade; where measurement scopes conflict, they are presented side by side and marked [To be verified].

1.2. Core Judgment of the Whole Chapter

The four tracks differ enormously in product form, customer base, and business model, yet between 2025 and 2026 they exhibit the same structural shift: the center of gravity of competition is moving from "which model is plugged in" to "the completeness of the Harness six layers". This chapter will argue: this shift has independent evidence in every track; L4 (Memory & State Layer, reproducibility and asset persistence) and L5 (Evaluation & Observability Layer) are shortfalls common to the whole industry; and the Chinese and overseas markets will follow different convergence paths along three dimensions — compliance, deployment form, and open-source strategy.


2. AI IDE Track

2.1. Market Overview: Adoption Saturated, Trust Has Not Kept Pace

The AI IDE is the most fiercely contested track, and the one that best verifies that "model capability does not equal engineering usability". Three sets of authoritative survey data sketch out the basic shape of the market:

MetricValueSource (scope noted)
Developer adoption84% are using or planning to use AI tools (76% in 2024); 51% of professional developers use them dailyStack Overflow 2025 Developer Survey, 2025-07-30, 49,000+ responses
Trust46% do not trust the accuracy of AI output (31% in 2024); only 3.1% trust it highlyStack Overflow 2025
Agent-form penetrationAbout 31% already use AI agents; 37.9% do not plan toStack Overflow 2025
Main frustrations66% think "AI solutions are almost right but not entirely right"; 45.2% think debugging AI code takes longerStack Overflow 2025
Enterprise-side adoptionOver 90% use AI (a separate 95% figure exists); 80% perceive individual productivity gains; 30% have single tasks exceeding 4 hoursDORA 2025, 2025-11-12
Measured efficiencyExperienced developers measured 19% slower when using AI, but self-assessed 20% fasterMETR randomized controlled trial, 2025-07-10

The implication of these numbers is: market education is complete, but trust-building has not yet begun. The simultaneous rise in adoption and distrust shows that the value of the tools is acknowledged, while their output has not yet attained engineering predictability. "66% think it is almost right but not entirely right" describes not a capability problem but a verifiability problem — precisely what the L5 layer of the Harness is meant to solve.

2.2. Representative Platform Comparison Matrix

PlatformDeveloperFormOpen SourceNotable Six-Layer Harness Characteristics
CursorAnysphereVS Code fork IDENoL1 rules hierarchy + codebase indexing; L6 team marketplace + SSO + audit logs + code-tracking API
Claude CodeAnthropicTerminal + IDE + WebNoL2 fine-grained sandboxing + credential masking; L1 proximity-based rules + progressive disclosure of skills + compaction; L3 plan mode + sub-agents + hooks
GitHub CopilotMicrosoft / GitHubMulti-IDE plugins + CLI + WebNoL6 organization policies + content exclusion + auditing + IP indemnification; L5 platform-side code review and usage dashboards; a shared agentic harness (the single Copilot SDK component) drives multiple product lines
Codex CLIOpenAITerminal + IDE extension + WebYes (Apache 2.0)L2 orthogonal configuration of OS-level sandboxing and approvals; L1 AGENTS.md hierarchy; L3 non-interactive execution and session recovery
WindsurfCognitionVS Code fork IDENoL4 memory system + checkpoint rollback as the core differentiator; L3 plan mode + parallel sessions
Trae字节跳动VS Code fork IDE + WebNoL1 context compaction + rules engineering; L3 multi-agent; entering on price and the domestic model ecosystem

The tone-setting evidence for this track: GitHub officially stated in 2026 that "the model provides the raw intelligence, while the harness determines how effectively that intelligence is applied" (the harness shapes how effectively that intelligence is applied), and published the controlled-comparison methodology of its agentic harness; the official Terminal-Bench also made clear that its leaderboard "ranks systems, not models". The dual confirmation by the official leaderboard and top vendors elevates "Harness completeness determines product gaps" from an industry intuition to a citable engineering conclusion (methodological details in Chapter 5 of this whitepaper).

2.3. Competitive Focus

The competitive focus of the AI IDE track has clearly migrated:

  1. L1 Context Engineering is the first competitive dimension. A medium-sized repository holds hundreds of thousands of lines of code, and no context window is large enough to hold them all. The three-piece set of rules file hierarchy (loaded close by), semantic codebase indexing, and context compaction determines how complex a requirement a single task can carry.
  2. L6 Governance is an entry requirement, not a bonus item. The higher the autonomy, the larger the blast radius of an incident — the Amazon Q prompt-injection incident of 2025-08-11 (malicious instructions attempting to induce the agent to delete AWS resources) was the first large-scale public exposure of the structural weakness that "external content being processed can directly become instructions". Filesystem isolation and network isolation are both indispensable, credentials are invisible by default, and destructive commands require hard-coded interception.
  3. Positive evidence shows that governance and autonomy are positive-sum: Anthropic officially disclosed that its sandboxing cut internal permission prompts by 84% while improving safety — constraints make the speed of scaling possible.

2.4. Business Models: Three Billing Logics

Billing LogicRepresentativeAdvantageRisk
Subscription-quota modelClaude Code (with Claude subscription), Codex CLI (with ChatGPT subscription)Predictable costHitting limits at peak times; long tasks forced to interrupt
Credit / usage-based modelCursor (quota pool + usage-based), GitHub Copilot (AI Credits from 2026-06-01, 1 credit = USD 0.01)A single conversation and a longer agent session no longer cost the same; fairerCost uncertainty; requires watching the burn rate
Quota-refresh modelWindsurf (daily / weekly quota refresh from 2026-03)Easy budgetingPer-run ceiling constrains complex tasks

The three logics are converging toward a "subscription + usage" hybrid. The migration of billing models means buyers need new management capabilities: organization-level consumption-rate observation and budget alerts. The GitHub Copilot case also shows a product-line consolidation trend — its agentic harness, as the single shared component of the Copilot SDK, simultaneously drives multiple experiences including CLI, App, and code review; the positioning of "the Harness as a platform asset" is becoming ever clearer.


3. AI Agents Platform Track

3.1. Market Overview: A Three-Tier Supply Structure

The supply of AI Agents platforms has a clear three-tier structure, and the three tiers follow different Harness-completeness logics:

TierRepresentativeSupply LogicHarness Coverage
Code frameworksLangGraph, OpenAI Agents SDK, Claude Agent SDK, CrewAI, Microsoft Agent Framework, AutoGenAimed at engineering teams; open source or free; primarily covers L2 / L3L4 / L5 / L6 mostly left blank, built by users themselves or plugged in externally (e.g., LangSmith provides tracing and evaluation)
Low-code platformsDify, Coze (扣子), 腾讯元器Aimed at business teams; visual orchestration + managed runtimeL1 through L3 productized and packaged; L5 observability and L6 tenant governance tiered with the enterprise edition
Vertical autonomous agentsDevin (Cognition), ManusAimed at end-to-end task delivery; billed by workloadThe six-layer loop encapsulated inside the product; the most black-box

On growth signals (all vendor-side or third-party figures, to be cited with caution): Devin's ARR, per third-party compilation, grew from USD 37 million (2025-05) to USD 492 million (2026-05), and the valuation figure rose from USD 10.2 billion to USD 26 billion — the source is third-party aggregation (low confidence), [To be verified], and this whitepaper does not use it as a basis for industry-size estimation; the operating milestones officially disclosed by Devin (processing 147 trillion tokens and driving 80 million virtual computers in the 8 months since launch) are likewise self-reported. The certain fact on the framework side is: core frameworks are generally open source and free (LangGraph, CrewAI, and AutoGen are all MIT), and revenue is shifting to the managed and observability layers (LangSmith free tier 5,000 traces/month, Developer tier USD 39/user/month).

3.2. Representative Platform Comparison Matrix

PlatformCategoryOpen Source / LicenseBilling FormGovernance & Compliance Highlights
LangGraphCode frameworkMIT open sourceFramework free, model API fees onlyCheckpointing capability the most complete among comparable frameworks
OpenAI Agents SDKCode frameworkOpen sourceFramework free; default tracing dashboard requires an OpenAI platform accountDeeply coupled to the model vendor's runtime
Claude Agent SDKCode frameworkClosed-source SDKNo license fee, billed by Claude API tokensBuilt-in compaction, sub-agent context isolation, hooks
CrewAICode frameworkMIT open sourceAMP platform free tier 50 workflow runs/month, Enterprise customEnterprise tier provides FedRAMP High, SSO, RBAC (official figure)
DifyLow-code platformOpen source (partial) + commercial editionFree + cloud subscriptionMature private deployment; USD 30 million Series Pre-A in 2026-03 (third-party relay)
Coze (扣子)Low-code platformDomestic edition closed source, overseas edition has an open-source version (conflicting figures)Free tier + token overageEnterprise subscription about RMB 20,000 to 200,000/year (third-party review figure)
DevinVertical autonomous agentNoCore USD 20/month + USD 2.25/ACU (1 ACU ≈ 15 minutes of active work); Team USD 500/month including 250 ACUsEnterprise tier provides VPC deployment, SAML/OIDC SSO, centralized control (official figure)
ManusVertical autonomous agentNo (L1/L4/L5 mechanisms closed source, unverifiable)Freemium + credit consumption (Starter from USD 20/month)Fewest publicly available governance details

3.3. Competitive Focus and Business Models

The competitive focus of this track shows two directions:

  1. Competition at the framework layer is about the "control plane". Once orchestration primitives (loops, sub-agents, tool calls) converged, differentiation shifted to state management (LangGraph's checkpointing), context governance (Claude Agent SDK's compaction), and observability (LangSmith). This is consistent with the rising weight of L3 and L4 in the six-layer model.
  2. Competition at the platform and vertical tiers is about L6 and cost metering. The core question in enterprise procurement has shifted from "can we orchestrate" to "how do we manage permissions, how do we export audit logs, how do we cap costs". Devin bills by ACU (active work time), Manus by credits, Dify / Coze by runs and tokens — the three forms arrive at the same destination: translating unpredictable token consumption into budgetable units of workload, which is essentially the productization of L5 observability capability.

The structural fact of the business models: the framework layer does not make money; what makes money is the runtime, observability, and governance. This distribution corroborates the judgment of this whitepaper — value is moving from the model-access layer to the Harness layer.

Same-Day Increment (2026-09-12; corrected 2026-09-13): The OpenAI Agents API entered public beta on 2026-09-10 (grade A, per the OpenAI official Changelog; the 2026-09-12 snapshot had recorded it as 09-11 per media reports), putting the orchestration capability of the existing Agents SDK under management — together with AWS Bedrock AgentCore GA (2026-06) this forms two moves in the same direction, marking that the competitive focus of this track is shifting from "are the orchestration primitives good to use" to "is the managed runtime governable". For details, see the research library: 03-市场研究/02-AI-Agents组/21-openai-agents-api.md.


4. AI Image and Content Generation Track

4.1. Market Overview

The image track is the most commercially mature of the four. Take 美图公司 as an example (HKEX announcement figure, very high confidence): total revenue for H1 2026 was RMB 2.21 billion (+22.1% year on year), of which imaging and design product revenue was RMB 1.77 billion (+30.9%, 80% of the total); adjusted net profit attributable to the parent was RMB 650 million (+39.5%); MAU was 282 million, with over 18.44 million paid subscribers (+19.7%) and a subscription penetration rate of 6.5%; productivity-app MAU was 33 million (+43.5%), with 2.35 million paid subscribers and ARR of about RMB 620 million. 美图's observability design is especially noteworthy: it uses "the quarter-over-quarter growth in AI compute-point consumption" (over 46% in both Q1 and Q2) as a proxy indicator for value hit-rate — measuring value hit-rate by consumption depth is a rarely mature practice of L5 observability in the content track.

The structural change on the supply side is equally significant: Google is injecting top-tier image capability into a general-purpose API via Nano Banana (Gemini-family image model) (first generation 1K/2K single image USD 0.039, Pro series 1K/2K USD 0.134, compiled from screenshots of the official pricing page), while domestic 通义万相 is waging price competition with API pricing of RMB 0.2 per image (Alibaba Cloud official) — basic image generation is being rapidly commoditized, and profit is migrating to workflows, asset management, and industry solutions (the Harness layer).

4.2. Representative Platform Comparison Matrix

PlatformDeveloperFormPricing FormSix-Layer Harness Positioning
MidjourneyMidjourneyWeb / DiscordSubscription USD 10 to 120/monthThe archetype of strong model, weak Harness: L5 relies on human-eye subjective evaluation and community feedback; no official regression-set mechanism observed
Nano Banana (Gemini-family)GoogleAPI + in-product embeddingBilled by token / by imageDistributes capability via API; the Harness is borne by upper-layer applications
即梦 AI字节跳动SaaS + APIConsumer membership (two conflicting third-party figures)Model + creation tools + 抖音 distribution closed loop
可灵 AI快手SaaS + APIInternational site USD 6.99 to 127.99/month (official page); domestic site per secondary-source figureElement reference + series generation supporting consistency; commercial rights tiered
Leonardo.aiLeonardoWeb + APISubscription USD 12 to 60/month + PAYG API (USD 5 credit for new accounts, up to 10 concurrency)Three types of capacity constraints explicitly documented + Webhook + pricing calculator; highly engineered
RunwayRunwayWeb + APISubscription USD 12 to 76/month (annual); API USD 0.01/creditVideo / image hybrid workflows, Adobe integration
ComfyUIComfy Org (open-source community)Local node workflowsFree (GPL-3.0)60,000+ node ecosystem; workflow JSON is the state, versionable; all six layers programmable and self-buildable
WeShop / 美图设计室美图 and othersSaaSSubscription + compute pointsAgent Teams multi-agent collaboration + creative asset reuse (self-reported)

4.3. Competitive Focus and Business Models

Two structural observations on this track:

  1. The watershed of the degree of Harness-ification is at L3 and L4. 美图设计室 (multi-agent collaboration + creative asset reuse + compute-point metering), ComfyUI (workflows as JSON graphs, version-controllable, SDK-able), and Leonardo (explicit capacity constraints + Webhook + cost estimation) represent the third-generation form; Midjourney and 妙鸭 represent "strong model, weak Harness" — leading model quality but thin engineering carrying capacity, and 妙鸭 has already exited for this reason (see Section 4.4).
  2. Business models move from one-time payment to a subscription + consumption hybrid. Midjourney is pure subscription, Runway pure credits, Leonardo dual-track subscription + PAYG, 美图 subscription + compute points. Consumption-based pricing presupposes precise usage metering and alerting — another instance of an L5 capability becoming commercial infrastructure.

4.4. A Typical Failure Case: 妙鸭相机

妙鸭相机 is the only negative sample in this track with a complete lifecycle, and a textbook case of "strong model, weak Harness":

  • Explosion: launched 2023-07-17, topped the App Store overall chart in August 2023, daily actives broke 600,000, with over 4,000 users queued at peak to generate photos; the technical base was the "提香" model self-developed by 阿里大文娱.
  • Engineering shortfall: the L2 / L3 / L5 layers were almost blank — every interaction was an atomic operation (upload, wait, pick a template, produce the photo), with no intermediate state to intervene in and no parameters to adjust; the user's only quality-correction entry point was a "make it more like me" similarity fine-tune. L6 saw an early user-agreement trust incident (the official apology and revision).
  • Commercial shortfall: the digital avatar was a one-time RMB 9.9 (promotional price) buyout + 10 finished photos — typical traffic-driver pricing, lacking follow-on consumption scenarios and a subscription anchor; compared with similar identity products (可灵 caps subjects at 30 to 500 by membership tier) it lacked asset-operations depth.
  • Tech-route lock-in: the training-based identity pipeline (≥20 photos per user, full per-user training, hours of queuing) did not complete its migration after zero-shot identity-injection solutions (InstantID, PuLID) matured in 2024–2025, and the cost structure was locked in. From the Harness perspective: the training-based route made identity assets into "private state inside the product" — non-migratable, non-composable, non-auditable; the zero-shot route made identity assets into "context that can circulate across pipelines", gaining real asset liquidity.
  • Ending: the team formally disbanded at the end of September 2025 (the authoritative source for this fact is, to date, only the statement of a third-party tool site), and the product continues only at a minimal level of operation.

The failure proposition of 妙鸭 is worth writing into the project review of every content product: given that generation quality is already good enough, what determines a product's life or death is not the model but the Harness that carries it.


5. AI Comic-Drama and Novel Track

5.1. AI Comic-Drama: Overcapacity and Scarce Certainty

AI comic drama is the scene in the AIGC content industry where engineering has advanced fastest and where the contradiction that "model capability does not equal production usability" was exposed earliest. Industry data for H1 2026 (DataEye report and 中国网络视听协会 data, via secondary relay; the full original report was not cross-verified):

MetricValueNotes
Total micro-short-drama releases367,000 titles, over 74% AI content中国网络视听协会 data (relayed)
Douyin-native new AI dramas and comic dramas221,900 titles, 515.738 billion cumulative plays
Hit rateOnly 1,055 titles exceeded 100 million plays, about 0.47%DataEye
Break-even rateUnder 1.3%, roughly 1 of every 77 titles recoupsMeasured against a break-even line of 50 million plays
Release speedOn average one title every 36 secondsDataEye
Revenue per 10,000 playsFell from a peak of RMB 30 to 100 to RMB 5 to 10, a drop of over 90%DataEye
Market sizeExpected to break RMB 40 billion in 2026 (+138%); overseas expansion expected to exceed USD 4 billion中研普华 / DataEye relay, [To be verified]

Capacity is piling up at a rate of one title every 36 seconds, the hit rate is under 0.5%, and unit prices have dropped by 90% — the industry has moved from "racing for capacity" to "racing for certainty". And certainty is precisely what models cannot provide: the same character looks different in episode 1 and episode 47, the same scene has lighting jump between two shots, and plot state is lost when writing continues across sessions. Not one of these problems can be solved by swapping in a stronger model; they can only be solved by the Harness's state management (L4), orchestration (L3), and regression evaluation (L5).

5.2. AI Comic-Drama Representative Platform Comparison Matrix

PlatformDeveloperCompetitive LeverL4 (Memory / Assets) ImplementationL6 Governance
即梦 AI字节跳动Model + tools + 抖音 distribution closed loop; multi-shot narrativeStable retention of character features; asset-library figure not disclosedReal-person avatar authentication; on 2026-04-28 investigated by the cyberspace authority for failing to effectively implement AI-generated synthetic content labeling (publicly recorded penalty event)
可灵 AI快手Intelligent storyboard + element reference; overseas revenue about 70% (third-party figure)Series generation + extended consistencyCommercial rights unlocked by tier
PixVerse爱诗科技CLI / Skills compatible with coding agents such as Claude Code and Codex; covers 175 countriesCharacter-consistent characters + Team asset libraryTeam Plan RBAC + credit cap + usage analytics
海螺 AIMiniMaxH3-Context-IR context compression (about 100k tokens compressed to about 4k) + open weights@ reference system + In-Context RegenerationOpen weights are self-auditable; adaptation to multiple domestic chips completed
Vidu生数科技Reference-to-video "everything can be referenced"Subject library (characters / props / scenes), up to 7 reference imagesSaaS / MaaS tiering
白日梦 AI光魔科技Permanent character library, claims 95%+ cross-episode similarityThe character library is the core selling pointPublic governance information missing
ComfyUIOpen-source community60,000+ nodes, fully local offlineCharacter LoRA + style anchors; workflow JSON is the state, versionableGPL-3.0, labeling pipeline can be self-built (compliance self-build example)
豆包 / Seedance字节 SeedMaaS API outputNo native asset library; upper layers must build it themselvesReal-person face ban + real-person verification for digital avatars (one of the strictest publicly known compliance samples)

The layering rule of this track is clear: L1 through L3 (four-modality input, intelligent storyboard, multi-shot narrative, node workflows) have by 2026 become standard equipment on leading platforms, highly homogenized; L4 has not yet homogenized — from "up to 7 reference images" (Vidu) to "about 100k tokens compressed to about 4k tokens" (海螺 H3) to "workflow JSON is the state" (ComfyUI), the implementation paths are entirely different. Whoever builds L4 solidly (character anchoring, asset library, shot state machine) will be able to push AI comic drama from "can generate" to "can continuously produce a hundred episodes".

5.3. AI Novels: The Long-Horizon State-Management Race

AI novels are the content scene where L4 and L1 bear the greatest pressure in the six-layer model: the core difficulty of long serialized fiction is "not collapsing character settings or forgetting foreshadowing after hundreds of thousands of words" (the web-novel industry names this guaranteed-to-occur fault "eating the book"). The industry has converged on three L4 paradigms:

ParadigmMechanismRepresentative Implementation
A · Structured rolling summaryWhen context overflows, automatically compressed into a structured summary, retaining key plot threadsThe rolling-summary layer of Claude's long-form workflow
B · Keyword-triggered entry librarySettings broken into entries, injected into context only when relevantNovelAI Lorebook; AI Dungeon Story Cards
C · Versioned state snapshotAfter each chapter, state is extracted and archived as a versioned snapshot, revisitable on demandClaude Book's state/current/ symbolic link + per-chapter archiving; 灵蟹创作's "project constitution"

The platform landscape shows "Chinese and overseas platforms solving the same problem differently": 阅文妙笔 enters with "妙笔通鉴" in a "second brain" positioning (specializing in foreshadowing mining and detail retrieval); its disclosed business indicators are Miaobi daily actives more than doubling, daily-average token consumption up over 90%, and weekly author usage over 75% (vendor self-reported); 番茄小说 is known for the strongest L6 (mandatory AI declaration, expansion and continuation writing banned for guaranteed-minimum books, machine judgment of low-quality content); NovelAI stands in the subscription market on Lorebook and privacy-first (XSalsa20 client-side encryption, no training); Sudowrite supports English long-form with Story Bible + reading the first 20,000 words + 25-document chapter continuity. On the community side, the GitHub open-source project Claude-Code-Novel-Writer (v4.1, 2026-08-21) provides, through multi-agent orchestration of AGENTS.md + 7 roles + 5 skills, an engineering model of "authoritative source / derived index separation" (the manuscript is the authoritative source; state tracking is regenerable).

5.4. Governance Is a Prerequisite, Not an After-the-Fact Check

L6 in both content tracks has already shifted from a "bonus item" to a "precondition for launch":

  • Regulatory baseline: the 《人工智能生成合成内容标识办法》 (issued 2025-03-07, in effect 2025-09-01) requires videos to carry a prominent explicit label on the opening frame and an implicit label written into the file metadata; the mandatory national standard GB 45438—2025 quantifies it further: explicit label text height no less than 5% of the shortest edge of the frame, lasting no less than 2 seconds at normal playback speed, with metadata fields including Label / ContentProducer / ProduceID, etc.
  • Enforcement instance: on 2026-04-28, 即梦 AI was lawfully investigated by the cyberspace authority for failing to effectively implement the labeling rules — even for a top-tier big-company platform, the labeling pipeline can break.
  • Platform governance divergence: from 2025-09-23, 番茄 mandates authors to declare "whether AI is used"; AI continuation and expansion writing are limited to authors of non-guaranteed-minimum books (authorization tiered by contract type); 晋江文学城 permits only three AI use scenarios: text proofreading, creative-element assistance, and creative rough-outline assistance. Overseas platforms govern mainly around copyright ownership (whether work rights are claimed, whether user works are used for training) and privacy (encryption, no-training commitments).

The direct implication for engineering teams: compliance labeling must be built as an automatic node at the tail of the generation pipeline, a first-class attribute of the artifact alongside resolution and frame rate, and cannot rely on manual post-hoc addition.


6. Cross-Track Common Judgments

6.1. Competition Is Shifting from Model Capability to Harness Completeness

Aggregation of independent evidence across the four tracks:

TrackEvidence Against "Model-Capability Determinism"Shift of the Competitive Center of Gravity
AI IDEThe same model can differ by double-digit percentage points under different scaffolds; GitHub officially makes clear that the harness determines the efficiency of intelligence application; the Terminal-Bench leaderboard "ranks systems, not models"Length of model list → rules hierarchy, sandbox granularity, approval policy, observability
AI Agents PlatformsAfter orchestration primitives converged, framework differentiation shifted to state management and observability; frameworks are open source and free, revenue comes from managed services and governanceOrchestration capability → control plane (checkpointing, context governance, audit, cost metering)
AI ImageBasic generation capability rapidly commoditized (API pricing at the RMB 0.2-per-image level); 妙鸭, the "strong model, weak Harness" case, has exitedGeneration quality → workflows, asset management, usage metering
AI Comic-Drama and NovelsOvercapacity + 0.47% hit rate; consistency and continuity problems cannot be solved by swapping modelsPer-generation ceiling (L1 through L3) → continuous-production ceiling (L4)

The four tracks converge on the same conclusion in entirely different ways: once model capability crosses the "usable" threshold, the perceptible gap between products is determined mainly by the completeness of the Harness six layers. This and the evaluation findings of Chapter 5 of this whitepaper are two sides of the same coin — since an evaluation score is itself a composite product of "model × Harness", market competition naturally unfolds in the same composite way.

6.2. L4 and L5 Are Industry-Wide Shared Shortfalls

Cross-aggregating the six-layer maturity of the four tracks yields a highly consistent conclusion:

LayerCross-Track StatusEvidence
L1 Context EngineeringRelatively strongestRules files, retrieval, and compaction have mature practice in every track (the AI IDE three-piece set, Lorebook, reference-image injection, H3-Context-IR)
L2 Tooling & ExecutionStrong and rapidly convergingSandboxing, MCP, CLI-ification, and node-based execution primitives have become widespread
L3 Orchestration & ControlStrong, homogenizingMulti-agent orchestration, plan mode, and node DAGs are now standard equipment
L4 Memory & StateShared shortfallAI comic-drama asset libraries take various forms with no unified baseline; the "eating the book" failure is guaranteed in novels; 妙鸭's identity assets are non-migratable; long-term memory on low-code platforms is generally weak
L5 Evaluation & ObservabilityShared shortfallMost platforms have only "multi-candidate regeneration"-style weak evaluation; Chapter 5 of this whitepaper shows that in the AI SRE / DevOps / multi-agent directions even public benchmarks are absent; there is no objective yardstick for content quality
L6 Governance & SafetyPolarizedCompliance-driven (content tracks, enterprise editions) is strong; tool-type (Midjourney, some frameworks) is weak

The common root cause of the L4 shortfall is: the model itself is stateless, and every "consistency" problem is essentially a cross-call, cross-session, cross-cycle state-injection and persistence problem — character consistency (comic drama), setting consistency (novels), migratable identity assets (image), and recoverable sessions (IDE) are projections of the same kind of engineering problem onto different tracks. The common root cause of the L5 shortfall is: the output of content and open-domain tasks lacks an objective yardstick for judgment, and public benchmarks are absent. Therefore, L4 and L5 are both an industry-wide shortfall and, in the coming two years, the clearest differentiation window — whoever first turns "reproducibility and asset persistence" (L4) and "effect evaluation" (L5) into product capabilities rather than marketing claims will win, in their own track, a differentiated position like Vidu's subject library, 海螺's Context-IR, or 美图's compute-point observability.


7. Differences Between the Chinese and Overseas Markets

7.1. Ecosystem Structure

  • Overseas: the main axis is "model vendors + subscription + API economy". The allocation logic is product strength and global market coverage (PixVerse covers 175 countries; 可灵's roughly 70% overseas revenue share is its growth engine), with pricing mainly USD subscriptions and usage billing.
  • China: the main axis is "big-company channel closed loop + free strategy + revenue-sharing ecosystem". 即梦 connects to 抖音, 可灵 connects to 快手, with model + tools + distribution integrated; some domestic tools enter on a fully free strategy (e.g., 字节 Trae, and 腾讯元宝's public statement that it is "currently fully free to use, with no plans to charge yet"); on the content side a distinctive revenue-sharing system has formed (comic drama shared by "duration × unit price × type coefficient × copyright coefficient") along with guaranteed-minimum incentives (the 2026 full-year guaranteed-minimum budget for live-action short dramas at the 抖音 group exceeds RMB 1.5 billion).

7.2. Compliance Environment

  • China: compliance is an explicit hard constraint with enforcement instances already on record — the 《人工智能生成合成内容标识办法》 and GB 45438—2025 form a quantified baseline, and 即梦 being investigated on 2026-04-28 is the landmark event; content platforms additionally face filing thresholds (in January 2026 the filing threshold for key micro-short-dramas was raised from RMB 1 million to RMB 3 million) and subject-matter review.
  • Overseas: the transparency obligations of the EU AI Act apply from 2026-08-02 to labeling of AI-generated text directed at the public (an open-source license does not constitute an exemption); the United States is driven mainly by platform self-regulation and litigation. Overall, China follows a "pre-labeling + platform verification" model, the EU a "transparency obligation" model, and the United States an "after-the-fact accountability" model.

7.3. Deployment Form

  • China: Xinchuang (信创) and domestic-chip adaptation are real demands — on the first day of MiniMax H3's open-source release, adaptation was completed for 华为昇腾, 摩尔线程, 沐曦, 海光, 昆仑芯, 天数智芯, 壁仞, and others; private deployment carries high weight in enterprise procurement (Dify's private-deployment capability is its core selling point).
  • Overseas: cloud first, SaaS first; local deployment appears mainly in regulated industries and the open-source community (ComfyUI's fully local offline capability lets it satisfy both kinds of need at once).

7.4. Open-Source Weight Strategies

  • Chinese vendors are more aggressive about open-sourcing model weights: MiniMax H3 (open weights for a video model), 通义 Qwen family (coding and image), 智谱 GLM family (NovelAI's Xialong is fine-tuned from GLM-4.6) — open weights have become a lever for acquiring the developer ecosystem and domestic-chip adaptation.
  • Overseas top vendors handle it in tiers: OpenAI open-sourced Codex CLI under Apache 2.0 (CLI layer open source, model closed source); the framework layer is generally MIT (LangGraph, CrewAI, AutoGen); frontier model weights are essentially closed source. Overall, China shows an "open weights in exchange for ecosystem" strategy, and the overseas shows an "open-source tools and frameworks in exchange for standards, closed weights to protect profit" strategy.

7.5. Overview of Differences

DimensionChinese MarketOverseas Market
Ecosystem structureBig-company channel closed loop (model + tools + distribution integrated); free-strategy entry; revenue-sharing and guaranteed-minimum ecosystemModel vendors + subscription + API economy; global market coverage; USD subscriptions and usage billing
CompliancePre-labeling + platform verification (labeling measures, GB 45438—2025, filing threshold), with enforcement instances already on recordEU transparency obligations (EU AI Act, applicable 2026-08-02); the United States mainly platform self-regulation and after-the-fact accountability
Deployment formXinchuang (信创) and domestic-chip adaptation are real demands; private deployment carries high weightCloud first, SaaS first; local deployment concentrated in regulated industries and the open-source community
Open-source strategyOpen weights in exchange for ecosystem (MiniMax H3, Qwen family, GLM family)Open-source tools and frameworks in exchange for standards (Codex CLI, LangGraph, AGENTS.md); closed weights to protect profit
Governance focusContent labeling, filing, subject-matter review, training-data authorization (the 番茄 agreement uproar)Copyright ownership, privacy commitments (encryption, no training), youth and content policy
Content monetizationRevenue-sharing system + guaranteed-minimum incentives + ad placement (placement accounts for about 70% of production-chain cost)Subscriptions, tips, copyright licensing; distribution mainly via own channels and communities

It must be emphasized: the differences are differences of degree, not of essence. The two markets are fully aligned on the two structural judgments that "the center of gravity of competition is shifting from model capability to Harness completeness" (Section 6.1) and "L4 / L5 are shared shortfalls" (Section 6.2) — the differences show up in the speed of convergence and the constraints, not in the direction.


8. Three Possible Paths of Landscape Evolution

8.1. Path One: Platform Convergence

Big companies, leveraging their integrated advantage in models, channels, and compute, productize all six Harness layers, and small and medium players degrade into model suppliers or vertical plugins. Supporting signals: GitHub unifies multiple product lines with a single agentic harness; the "model + tools + distribution" closed loops of 即梦 / 可灵; 美图 integrating multiple product lines with a platform-level compute-point system. Constraint: L4 and L5 are precisely the layers hardest for big companies to go deep on — asset persistence and effect evaluation require long-term customer companionship, not traffic tactics, and vertical players still have a chance to hold their incumbent advantages.

8.2. Path Two: Vertical Specialization

Industry knowledge and the state layer become the moat, and general-purpose platforms are hard to replace. Supporting signals: the continuous-production capability of AI comic drama depends on the industry-specific asset library and storyboard state machine; the "project constitution" and foreshadowing ledger of web novels depend on domain semantics; in enterprise Agents procurement the weight of L6 (permissions, audit, privatization) keeps rising. 妙鸭's failure supports this path from the negative side — a "general-purpose hit product" without a state layer and operational depth has an extremely short lifecycle. Constraint: the vertical market size is limited, making it hard to amortize foundation-model R&D costs; in the end it mostly depends on big-company model APIs, and bargaining power is constrained by others.

8.3. Path Three: Open-Source Erosion

Open-source workflows and open-source weights squeeze closed-source SaaS from both ends: ComfyUI covers professional production with a 60,000+ node ecosystem; Codex CLI proves that closed-source vendors will proactively open-source the Harness layer to contest standards; Chinese vendors' open-weight strategy keeps lowering the bar for self-building. If open protocols such as MCP and open standards such as AGENTS.md keep penetrating (AGENTS.md has been adopted by 60,000+ open-source projects, hosted by the Agentic AI Foundation under the Linux Foundation), the six-layer capability will increasingly exist in the form of "assemblable open-source components", and closed-source platforms will be forced to narrow their profit surface toward "managed services and governance" (L5 / L6). Constraint: open-source solutions lack evaluation and governance by default, and enterprise self-building is costly — this in turn gives commercial products a space for survival built around L5 / L6.

8.4. Observation Indicators

The three paths are not mutually exclusive; the final landscape is more likely a mixture of all three. Judgments on where the landscape is heading can rest on the following observable indicators:

Observation IndicatorPoints to Path One (Convergence)Points to Path Two (Specialization)Points to Path Three (Open-Source Erosion)
Speed of open-source substitution of the six-layer capabilitySlowIrrelevantFast (L1 through L3 successively open-sourced)
Decision center of enterprise procurementPlatform brand and integrationDomain state layer and compliance capabilityComposability and exit cost
Openness of top vendors' Harness layerClosed loopPartially open (protocol-compatible)Fully open (standard output)
Productization progress of L4 / L5 capabilityBuilt into and bound to the platformDeeply bound to the vertical industryOpen-source components + commercial managed services

9. Conclusion

Based on systematic research across the four tracks, this chapter draws three conclusions:

  1. Competition is shifting from "model capability" to "Harness completeness". The AI IDE track is doubly confirmed by the official leaderboard and top vendors ("ranks systems, not models"; the harness determines the efficiency of intelligence application); in the Agents track frameworks are free and revenue flows to managed services and governance; in the image track basic generation is rapidly commoditized and profit flows to workflows and asset management; in both content tracks there is overcapacity and scarce certainty. The four tracks converge on the same structural change in different ways, and its mirror on the evaluation side is the finding of Chapter 5 of this whitepaper — a score has always been a composite product of "model × Harness".
  2. L4 (Memory & State) and L5 (Evaluation & Observability) are industry-wide shared shortfalls. The model is stateless, and every "consistency" is a state-management problem; content and open-domain tasks lack an objective yardstick, and public benchmarks are absent in multiple high-value directions. The complete-lifecycle failure of 妙鸭相机 (strong model, weak Harness, no state-asset liquidity, no evaluation loop) and the 0.47% hit rate of AI comic drama mark the commercial cost of these two layers' shortfalls from both the positive and negative sides; cases such as Vidu's subject library, 海螺's Context-IR, and 美图's compute-point observability mark the differentiated return of making up the shortfalls.
  3. China and overseas markets will take different shapes in the mixed evolution of the three paths: China is characterized by hard compliance constraints, channel closed loops, and an open-weight strategy; the overseas by a subscription and API economy, transparency obligations, and tool open-sourcing. Whatever the shape, the ultimate decisive variable of the landscape is the same — whoever can convert the model's uncertainty into engineering predictability holds the pricing power of the next competitive cycle.

Statement of Information Gaps

  1. Devin's ARR and valuation data (USD 37 million to USD 492 million; USD 10.2 billion to USD 26 billion) come from third-party aggregation (low confidence); all are marked [To be verified] and were not used in any industry-size estimation.
  2. The authoritative source for the disbanding of the 妙鸭相机 team (end of 2025-09) is, to date, only the statement of a third-party tool site; no tech-media coverage or business-registration corroboration has been located, [To be verified].
  3. The AI comic-drama market size (RMB 40 billion / USD 4 billion overseas) and hit-rate data come mainly from the DataEye report via secondary relay (百度百科, KOCPC, and others); the full original report was not obtained for cross-verification.
  4. Consumer pricing across platforms has multiple conflicting figures (即梦, 可灵 domestic site, 美图设计室, Coze enterprise edition, Vidu, and others); this whitepaper accepts only verifiable official-page figures, and the rest are marked [To be verified] or not cited.
  5. The release date of 可灵 3.0 and "up to 3 minutes", and 白日梦's "95%+ cross-episode similarity", etc., appear only in secondary sources, [To be verified].
  6. The 阅文妙笔 business indicators (daily actives +100%, token +90%, weekly usage over 75%) and the 美图设计室 Agent Teams results are both vendor self-reported figures.
  7. Most platforms in the AI novel and image tracks have no officially disclosed verifiable effectiveness indicators; third-party review data are mostly marketing pieces and were not cited.
  8. This whitepaper does not adopt any self-media data on big domestic companies' AI coding efficiency as hard data (no primary source for this kind of data was found in this project's search).

10. References

  1. 2025 Stack Overflow Developer Survey — Stack Overflow, 2025-07-30. https://survey.stackoverflow.co/2025/
  2. State of AI-assisted Software Development 2025 — DORA / Google Cloud, 2025-11-12. https://dora.dev/research/2025/dora-report/
  3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (randomized controlled trial) — METR, 2025-07-10. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  4. Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks — GitHub Blog, 2026. https://github.blog/ai-and-ml/github-copilot/evaluating-performance-and-efficiency-of-the-github-copilot-agentic-harness-across-models-and-tasks/
  5. Terminal-Bench official site and Leaderboard — Stanford / Laude Institute, from 2025 to 2026. https://www.tbench.ai/
  6. AGENTS.md official site — Agentic AI Foundation (under the Linux Foundation). https://agents.md/
  7. Sandboxing: a safer and more autonomous approach — Anthropic, 2025. https://www.anthropic.com/engineering/claude-code-sandboxing
  8. MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities — MiniMax official blog. https://www.minimax.io/blog/minimax-h3
  9. 美图公司 2026 interim results announcement — 香港交易所披露易 (hkexnews). https://www.hkexnews.hk/
  10. 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — issued 2025-03-07, in effect 2025-09-01. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  11. Mandatory national standard GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — 全国网络安全标准化技术委员会. https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
  12. DataEye《2026 上半年 AI 剧 / 漫剧数据报告》 — DataEye研究院 (reprinted via secondary sources such as KOCPC). https://en.kocpc.com.tw/archives/25203
  13. 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — 澎湃新闻, 2026. https://www.thepaper.cn/newsDetail_forward_33993493
  14. 中国青年报《"妙笔通鉴""漫剧助手"发布,AI 赋能网文创作和 IP 改编》 — 中国青年报, 2025. https://new.qq.com/rain/a/20251017A08J6J00
  15. GitHub — forsonny/Claude-Code-Novel-Writer (Multi-Agent Novel Writer v4.1) — 2026. https://github.com/forsonny/Claude-Code-Novel-Writer
  16. Apache Software Foundation《AI Agents and AGENTS.md Files (Draft)》 — ASF legal-discuss, 2026. https://www.apache.org/legal/generative-tooling-agents.html
  17. 通义万相 and Qwen-family models: 阿里云 official pricing page — 阿里云. https://help.aliyun.com/
  18. Leonardo.ai official pricing and API documentation — Leonardo. https://leonardo.ai/pricing
  19. Runway official pricing page — Runway. https://runwayml.com/pricing
  20. Comfy official site — Comfy Org. https://comfy.org/