产业格局:从模型能力竞争到 Harness 完整度竞争
1. 引言:四个赛道与一个共同变量
1.1. 研究范围与分析框架
图 1-1|Harness 六层能力模型:竞争重心从模型能力到完整度
数据来源:基于本文分析绘制的示意图。
本章基于对四个内容与工具赛道的系统调研,给出 AI Harness 视角下的产业格局全景:
- AI IDE:以编码场景为核心的智能体开发环境(Cursor、Claude Code、GitHub Copilot、Codex CLI、Trae 等);
- AI Agents 平台:智能体框架、低代码平台与垂直自主智能体(LangGraph、OpenAI Agents SDK、Claude Agent SDK、Dify、Coze、Devin、Manus 等);
- AI 图像与内容生成:文生图、图像编辑与设计工具(Midjourney、即梦、可灵、Nano Banana、Leonardo、ComfyUI 等);
- AI 漫剧与小说:长程内容生产(即梦 AI 漫剧、可灵、Vidu、白日梦;阅文妙笔、番茄、NovelAI、Sudowrite 等)。
分析框架沿用本白皮书第 3 章的六层能力模型:不把平台当作"工具"罗列功能,而是把每个平台当作一套已经落地的 Harness 实现来解剖——看它在哪一层做了真投入、在哪一层留了缺口、缺口由谁来补。所有厂商经营数据均标注 self-reported 或来源等级;口径冲突处并列呈现并标 。
1.2. 全章核心判断
四个赛道的产品形态、客户群体与商业模式差异巨大,但在 2025 至 2026 年呈现出同一个结构性变化:竞争的重心正从"接了哪个模型"转向"Harness 六层的完整度"。本章将论证:这一转变在每个赛道都有独立证据;L4(记忆与状态层,可复现性与资产持久化)与 L5(评估与观测层)是全行业共同的短板;而中国与海外市场将在合规、部署形态与开源策略三个维度上走出不同的收敛路径。
2. AI IDE 赛道
2.1. 市场全景:采用率饱和,信任度未跟上
AI IDE 是竞争最激烈、也最能验证"模型能力不等于工程可用性"的赛道。三组权威调研数据勾出市场的基本形态:
| 指标 | 数值 | 来源(注明口径) |
|---|---|---|
| 开发者采用率 | 84% 正在使用或计划使用 AI 工具(2024 年为 76%);51% 专业开发者每日使用 | Stack Overflow 2025 开发者调查,2025-07-30,49,000+ 份回答 |
| 信任度 | 46% 不信任 AI 输出准确性(2024 年为 31%);仅 3.1% 高度信任 | Stack Overflow 2025 |
| 智能体形态渗透 | 约 31% 已在使用 AI agent;37.9% 不打算用 | Stack Overflow 2025 |
| 主要挫败 | 66% 认为"AI 方案几乎对但不完全对";45.2% 认为调试 AI 代码更耗时 | Stack Overflow 2025 |
| 企业侧采用 | 90% 以上使用 AI(另有 95% 口径);80% 感知个体生产力提升;30% 有单次超 4 小时的长任务 | DORA 2025,2025-11-12 |
| 实测效率 | 资深开发者使用 AI 后实测慢 19%,自评快 20% | METR 随机对照试验,2025-07-10 |
这组数据的含义是:市场教育已经完成,信任建设尚未开始。采用率与不信任率同步上升,说明工具价值已被承认,而产出尚未获得工程上的可预期性。"66% 认为几乎对但不完全对"描述的不是能力问题,而是可验证性问题——恰恰是 Harness 的 L5 层要解决的。
2.2. 代表平台对比矩阵
| 平台 | 开发商 | 形态 | 开源 | Harness 六层的显著特征 |
|---|---|---|---|---|
| Cursor | Anysphere | VS Code 分支 IDE | 否 | L1 规则层级 + 代码库索引;L6 团队市场 + SSO + 审计日志 + 代码追踪 API |
| Claude Code | Anthropic | 终端 + IDE + Web | 否 | L2 细粒度沙箱 + 凭据掩码;L1 就近规则 + 技能渐进披露 + 压缩;L3 计划模式 + 子智能体 + 钩子 |
| GitHub Copilot | Microsoft / GitHub | 多 IDE 插件 + CLI + Web | 否 | L6 组织策略 + 内容排除 + 审计 + IP 赔付;L5 平台侧代码审查与用量看板;共享 agentic harness(Copilot SDK 单一组件)驱动多条产品线 |
| Codex CLI | OpenAI | 终端 + IDE 扩展 + Web | 是(Apache 2.0) | L2 OS 级沙箱与审批正交配置;L1 AGENTS.md 层级;L3 非交互执行与会话恢复 |
| Windsurf | Cognition | VS Code 分支 IDE | 否 | L4 记忆系统 + 检查点回退为差异化核心;L3 计划模式 + 并行会话 |
| Trae | 字节跳动 | VS Code 分支 IDE + Web | 否 | L1 上下文压缩 + 规则工程;L3 多智能体;以价格与国内模型生态切入 |
关于该赛道的定调性证据:GitHub 官方在 2026 年明确表述"模型提供原始智能,而 harness 决定这份智能被应用得多有效"(the harness shapes how effectively that intelligence is applied),并公开了其 agentic harness 的受控对比方法学;Terminal-Bench 官方亦明确其排行榜"排的是系统而非模型"。官方榜单与头部厂商的双重确认,使"Harness 完整度决定产品差距"从行业直觉升级为可引用的工程结论(方法论细节见本白皮书第 5 章)。
2.3. 竞争焦点
AI IDE 赛道的竞争焦点已明确迁移:
- L1 上下文工程是第一竞争维度。一个中型仓库几十万行代码,而上下文窗口再大也装不下。rules 文件层级(就近加载)、代码库语义索引、上下文压缩(compaction)三件套决定单次任务能承载多复杂的需求。
- L6 治理是准入项而非加分项。自主度越高,事故爆炸半径越大——2025-08-11 的 Amazon Q 提示注入事件(恶意指令试图诱导智能体删除 AWS 资源)是"被处理的外部内容可以直接成为指令"这一结构性弱点的首次大规模公开暴露。文件系统隔离与网络隔离缺一不可,凭据默认不可见,破坏性命令需硬编码拦截。
- 正面证据表明治理与自主性是正和:Anthropic 官方披露其沙箱化使内部权限提示减少 84%,同时提升安全性——约束让规模化速度成为可能。
2.4. 商业模式:三种计费逻辑
| 计费逻辑 | 代表 | 优点 | 风险 |
|---|---|---|---|
| 订阅额度制 | Claude Code(随 Claude 订阅)、Codex CLI(随 ChatGPT 订阅) | 成本可预测 | 高峰期撞限,长任务被迫中断 |
| 信用点 / 用量制 | Cursor(额度池 + 按用量)、GitHub Copilot(2026-06-01 起 AI Credits,1 credit = 0.01 美元) | 一次对话与一时代理会话不再同价,更公平 | 成本不确定性,需盯消耗速率 |
| 配额刷新制 | Windsurf(2026-03 起按日 / 周刷新配额) | 便于预算 | 复杂任务单次上限受约束 |
三种逻辑正向"订阅 + 用量"混合收敛。计费模式的迁移意味着采购方需要新的管理能力:组织级消耗速率观测与预算告警。GitHub Copilot 的案例还显示产品线整合趋势——其 agentic harness 作为 Copilot SDK 的单一共享组件,同时驱动 CLI、App、代码审查等多条体验,"Harness 即平台资产"的定位日益清晰。
3. AI Agents 平台赛道
3.1. 市场全景:三层供给结构
AI Agents 平台供给呈清晰的三层结构,且三层遵循不同的 Harness 完整度逻辑:
| 层 | 代表 | 供给逻辑 | Harness 覆盖 |
|---|---|---|---|
| 代码框架 | LangGraph、OpenAI Agents SDK、Claude Agent SDK、CrewAI、Microsoft Agent Framework、AutoGen | 面向工程团队,开源或免费,主要覆盖 L2 / L3 | L4 / L5 / L6 大多留白,由使用者自建或外接(如 LangSmith 提供追踪与评估) |
| 低代码平台 | Dify、Coze(扣子)、腾讯元器 | 面向业务团队,可视化编排 + 托管运行时 | L1 至 L3 产品化封装,L5 观测与 L6 租户治理随企业版分级 |
| 垂直自主智能体 | Devin(Cognition)、Manus | 面向端到端任务交付,按工作量计费 | 六层闭环封装于产品内,黑盒程度最高 |
增长信号方面(均为厂商侧或第三方口径,需谨慎引用):Devin 的 ARR 据第三方汇编从 3,700 万美元(2025-05)增至 4.92 亿美元(2026-05),估值口径从 102 亿美元升至 260 亿美元——来源为第三方汇总(低置信),,本白皮书不作为行业规模测算依据;Devin 官方披露的运营里程碑(上线 8 个月处理 147 万亿 token、驱动 8,000 万虚拟计算机)同样为 self-reported。框架侧的确定性事实是:核心框架普遍开源免费(LangGraph、CrewAI、AutoGen 均 MIT),收入转向托管与观测层(LangSmith 免费档 5,000 traces/月,Developer 档 39 美元/用户/月)。
3.2. 代表平台对比矩阵
| 平台 | 类别 | 开源 / 许可 | 计费形态 | 治理与合规亮点 |
|---|---|---|---|---|
| LangGraph | 代码框架 | MIT 开源 | 框架免费,仅付模型 API 费 | 检查点(checkpointing)能力为同类框架中最完整 |
| OpenAI Agents SDK | 代码框架 | 开源 | 框架免费;默认追踪看板需 OpenAI 平台账号 | 与模型厂商运行时深度耦合 |
| Claude Agent SDK | 代码框架 | 闭源 SDK | 无许可费,按 Claude API token 计费 | 内置 compaction、子智能体上下文隔离、hooks |
| CrewAI | 代码框架 | MIT 开源 | AMP 平台免费档 50 次工作流执行/月,Enterprise 定制 | 企业档提供 FedRAMP High、SSO、RBAC(官方口径) |
| Dify | 低代码平台 | 开源(部分) + 商业版 | 免费 + 云订阅 | 私有化部署成熟;2026-03 获 3,000 万美元 Series Pre-A(第三方转述) |
| Coze(扣子) | 低代码平台 | 国内版闭源、海外版有开源版本(口径冲突) | 免费档 + 按 token 超量 | 企业版订阅约 2 万至 20 万元/年(第三方评测口径) |
| Devin | 垂直自主智能体 | 否 | Core 20 美元/月 + 2.25 美元/ACU(1 ACU 约 15 分钟主动工作);Team 500 美元/月含 250 ACU | Enterprise 档提供 VPC 部署、SAML/OIDC SSO、集中管控(官方口径) |
| Manus | 垂直自主智能体 | 否(L1/L4/L5 机制闭源,不可验证) | Freemium + credit 消耗制(Starter 20 美元/月起) | 公开治理细节最少 |
3.3. 竞争焦点与商业模式
该赛道的竞争焦点呈现两个方向:
- 框架层的竞争在"控制平面"。当编排原语(循环、子智能体、工具调用)趋同后,差异化转向状态管理(LangGraph 的 checkpointing)、上下文治理(Claude Agent SDK 的 compaction)与可观测性(LangSmith)。这与六层模型中 L3 与 L4 的权重上升一致。
- 平台层与垂直层的竞争在 L6 与成本计量。企业采购的核心问题已从"能不能编排"变为"权限怎么管、审计怎么导出、成本怎么封顶"。Devin 按 ACU(主动工作时间)计费、Manus 按 credit 计费、Dify / Coze 按执行与 token 计费——三家形态殊途同归:把不可预测的 token 消耗翻译成可预算的工作量单位,这本质上是 L5 观测能力的产品化。
商业模式的结构性事实是:框架层不挣钱,挣钱的是运行时、观测与治理。这一分布印证了本白皮书的判断——价值正在从模型接入层向 Harness 层转移。
当日增量(2026-09-12;2026-09-13 修正):OpenAI Agents API 于 2026-09-10 进入公测(A级,OpenAI 官方 Changelog 口径;2026-09-12 快照曾依媒体报道记为 09-11),把既有 Agents SDK 的编排能力托管化——这与 AWS Bedrock AgentCore GA(2026-06)构成同一方向的两次落子,标志着该赛道的竞争焦点正从"编排原语是否好用"转向"托管运行时是否可治理"。详见调研库 03-市场研究/02-AI-Agents组/21-openai-agents-api.md。
4. AI 图像与内容生成赛道
4.1. 市场全景
图像赛道是四个赛道中商业化最成熟的。以美图公司为例(港股公告口径,极高置信):2026 上半年总收入 22.1 亿元(同比 +22.1%),其中影像与设计产品收入 17.7 亿元(+30.9%,占 80%);经调整归母净利润 6.5 亿元(+39.5%);MAU 2.82 亿,付费订阅用户超 1,844 万(+19.7%),订阅渗透率 6.5%;生产力应用 MAU 3,300 万(+43.5%),付费订阅 235 万,ARR 约 6.2 亿元。美图的可观测设计尤其值得注意:其以"AI 算力点消费环比增幅"(Q1、Q2 均超 46%)作为价值命中率代理指标——用消费深度衡量价值命中率,是 L5 观测在内容赛道的罕见成熟实践。
供给侧的结构变化同样显著:Google 以 Nano Banana(Gemini 系图像模型)把顶级图像能力注入通用 API(初代 1K/2K 单图 0.039 美元、Pro 系列 1K/2K 0.134 美元,官方定价页截图整理口径),国内通义万相以 0.2 元/张的 API 定价(阿里云官方)展开价格竞争——基础图像生成正在快速商品化,利润向工作流、资产管理与行业方案(Harness 层)迁移。
4.2. 代表平台对比矩阵
| 平台 | 开发商 | 形态 | 付费形态 | Harness 六层定位 |
|---|---|---|---|---|
| Midjourney | Midjourney | Web / Discord | 订阅 10 至 120 美元/月 | 强模型弱 Harness 的代表:L5 依赖人眼主观评测与社区反馈,未见官方回归集机制 |
| Nano Banana(Gemini 系) | API + 产品内嵌 | 按 token / 按图计费 | 以 API 分发能力,Harness 由上层应用承担 | |
| 即梦 AI | 字节跳动 | SaaS + API | C 端会员(两套第三方口径冲突) | 模型 + 创作工具 + 抖音分发闭环 |
| 可灵 AI | 快手 | SaaS + API | 国际站 6.99 至 127.99 美元/月(官方页);国内站为二次来源口径 | 元素引用 + 系列生成支撑一致性;商用权按档位 |
| Leonardo.ai | Leonardo | Web + API | 订阅 12 至 60 美元/月 + PAYG API(新账户赠 5 美元,最高 10 并发) | 三类容量约束显式文档化 + Webhook + 定价计算器,工程化程度高 |
| Runway | Runway | Web + API | 订阅 12 至 76 美元/月(年付);API 0.01 美元/credit | 视频 / 图像混合工作流,Adobe 集成 |
| ComfyUI | Comfy Org(开源社区) | 本地节点工作流 | 免费(GPL-3.0) | 60,000+ 节点生态;工作流 JSON 即状态、可版本化;六层全部可编程自建 |
| WeShop / 美图设计室 | 美图等 | SaaS | 订阅 + 算力点 | Agent Teams 多智能体协同 + 创作资产复用(self-reported) |
4.3. 竞争焦点与商业模式
该赛道的两个结构性观察:
- Harness 化程度的分水岭在 L3 与 L4。美图设计室(多 Agent 协同 + 创作资产复用 + 算力点计量)、ComfyUI(工作流即 JSON 图、可版本控制、可 SDK 化)、Leonardo(容量约束显式化 + Webhook + 成本预估)代表第三代形态;Midjourney、妙鸭代表"强模型弱 Harness"——模型质量领先但工程承载薄弱,且妙鸭已因此出局(详见 4.4 节)。
- 商业模式从单次付费转向订阅 + 消耗混合。Midjourney 纯订阅、Runway 纯 credit、Leonardo 订阅 + PAYG 双轨、美图订阅 + 算力点。消耗制的前提是精准的用量计量与告警——又一个 L5 能力成为商业基础设施的实例。
4.4. 典型失败样本:妙鸭相机
妙鸭相机是本赛道唯一具备完整生命周期的负面样本,也是"强模型弱 Harness"的教科书案例:
- 爆发:2023-07-17 上线,2023 年 8 月登顶 App Store 总榜,日活突破 60 万,高峰期 4,000 余人排队出片;技术底座为阿里大文娱自研"提香"模型。
- 工程短板:L2 / L3 / L5 三层几乎空白——每次交互都是原子操作(上传、等待、选模板、出片),无中间状态可干预、无参数可调节;用户唯一的质量修正入口是"更像我一点"的相似度微调。L6 在早期发生用户协议信任事故(官方道歉并修订)。
- 商业短板:数字分身一次性 9.9 元(推广价)买断 + 10 张成片,是典型引流品定价,缺乏后续消耗场景与订阅锚点;对比同类身份类产品(可灵按会员档位限制 30 至 500 个主体)缺少资产运营纵深。
- 技术路线固化:训练式身份管线(每用户 ≥20 张照片、完整 per-user 训练、数小时排队)在 2024 至 2025 年零样本身份注入方案(InstantID、PuLID)成熟后未完成迁移,成本结构被固化。从 Harness 视角看:训练式路线把身份资产做成了"产品内的私有状态",不可迁移、不可组合、不可审计;零样本路线把身份资产做成"可跨管线流通的上下文",获得真正的资产流动性。
- 结局:团队于 2025 年 9 月底正式解散(该事实的权威信源目前仅有第三方工具站表述),产品仅维持最低限度运营。
妙鸭的失败命题值得写进所有内容类产品的立项评审:在生成质量已经足够好的前提下,决定产品生死的不是模型,而是承载模型的 Harness。
5. AI 漫剧与小说赛道
5.1. AI 漫剧:产能过剩与确定性稀缺
AI 漫剧是 AIGC 内容产业中工程化程度提升最快、也最早暴露"模型能力不等于生产可用"矛盾的场景。2026 年上半年的行业数据(DataEye 报告与中国网络视听协会数据,经二手转引,原始报告全文未交叉验证):
| 指标 | 数值 | 说明 |
|---|---|---|
| 微短剧上线总量 | 36.7 万部,AI 内容占比超 74% | 中国网络视听协会数据(转引) |
| 抖音端原生新增 AI 剧及漫剧 | 22.19 万部,累计播放 5,157.38 亿次 | — |
| 爆款率 | 播放量破亿仅 1,055 部,约 0.47% | DataEye |
| 回本率 | 不足 1.3%,约每 77 部 1 部回本 | 按 5,000 万播放盈亏平衡线测算 |
| 上线速度 | 平均每 36 秒一部 | DataEye |
| 万次播放收入 | 从峰值 30 至 100 元跌至 5 至 10 元,跌幅超九成 | DataEye |
| 市场规模 | 2026 年预计突破 400 亿元(+138%);出海预计超 40 亿美元 | 中研普华 / DataEye 转引, |
产能以每 36 秒一部的速度堆积、爆款率不足 0.5%、单价跌去九成——行业从"抢产能"进入"抢确定性"。而确定性恰恰是模型给不了的:同一个角色在第 1 集和第 47 集长得不一样、同一场景在两个镜头之间光影跳变、剧情状态跨会话续写时丢失。这些问题没有一个能靠换更强的模型解决,只能靠 Harness 的状态管理(L4)、编排(L3)与回归评估(L5)解决。
5.2. AI 漫剧代表平台对比矩阵
| 平台 | 开发商 | 竞争抓手 | L4(记忆 / 资产)实现 | L6 治理 |
|---|---|---|---|---|
| 即梦 AI | 字节跳动 | 模型 + 工具 + 抖音分发闭环;多镜头叙事 | 角色特征稳定保持;资产库口径未公开 | 真人分身认证;2026-04-28 因未有效落实 AI 生成合成内容标识被网信部门查处(公开处罚事件) |
| 可灵 AI | 快手 | 智能分镜 + 元素引用;海外收入占比约 70%(第三方口径) | 系列生成 + 一致性延长 | 商用权按档位解锁 |
| PixVerse | 爱诗科技 | CLI / Skills 兼容 Claude Code、Codex 等编码智能体;覆盖 175 个国家 | Character 一致性角色 + Team 资产库 | Team Plan RBAC + 积分上限 + 用量分析 |
| 海螺 AI | MiniMax | H3-Context-IR 上下文压缩(约 100k token 压缩至约 4k)+ 开源权重 | @ 引用系统 + In-Context Regeneration | 开源权重可自审;完成多款国产芯片适配 |
| Vidu | 生数科技 | 参考生视频"万物可参" | 主体库(角色 / 道具 / 场景),最多 7 张参考图 | SaaS / MaaS 分层 |
| 白日梦 AI | 光魔科技 | 永久角色库,宣称跨集相似度 95%+ | 角色库为核心卖点 | 公开治理信息缺失 |
| ComfyUI | 开源社区 | 60,000+ 节点、完全本地离线 | 角色 LoRA + 风格锚点;工作流 JSON 即状态、可版本化 | GPL-3.0,标识管线可自建(合规自建范例) |
| 豆包 / Seedance | 字节 Seed | MaaS API 输出 | 无原生资产库,需上层自建 | 真人人脸禁用 + 数字分身真人校验(公开的合规最严样本之一) |
该赛道的分层规律清晰:L1 至 L3(四模态输入、智能分镜、多镜头叙事、节点工作流)在 2026 年已成为头部平台标配、高度同质化;L4 尚未同质化——从"最多 7 张参考图"(Vidu)到"约 100k token 压缩至约 4k token"(海螺 H3)再到"工作流 JSON 即状态"(ComfyUI),实现路径完全不同。谁把 L4 做扎实(角色锚定、资产库、镜头状态机),谁就能把 AI 漫剧从"能生成"推进到"能连续生产一百集"。
5.3. AI 小说:长程状态管理竞赛
AI 小说是六层模型中 L4 与 L1 压力最大的内容场景:长篇连载的核心难题是"几十万字后仍不崩人设、不忘伏笔"(网文行业以"吃书"命名这一必现故障)。行业收敛出三种 L4 范式:
| 范式 | 机制 | 代表实现 |
|---|---|---|
| A · 结构化滚动摘要 | 上下文超长时自动压缩为结构化摘要,保留关键线索 | Claude 长文工作流的滚动摘要层 |
| B · 关键词触发条目库 | 设定拆为条目,仅在相关时注入上下文 | NovelAI Lorebook;AI Dungeon Story Cards |
| C · 版本化状态快照 | 每章完成后抽取状态并存档为版本化快照,按需回溯 | Claude Book 的 state/current/ 符号链接 + 逐章归档;灵蟹创作的"项目宪法" |
平台格局呈现"中文平台与海外平台同题异解":阅文妙笔以"妙笔通鉴"切入"第二大脑"定位(专做伏笔挖掘与细节检索),其披露的业务指标为妙笔日活增长超一倍、日均 token 消耗增长超 90%、作家周使用率超 75%(厂商 self-reported);番茄小说以最强 L6 著称(强制 AI 申报、保底书禁用扩写续写、低质内容机器判定);NovelAI 以 Lorebook 与隐私优先(XSalsa20 客户端加密、不训练)立足订阅市场;Sudowrite 以 Story Bible + 读前 20,000 词 + 25 文档章节连续性支撑英文长篇。社区侧,GitHub 开源项目 Claude-Code-Novel-Writer(v4.1,2026-08-21)以 AGENTS.md + 7 角色 + 5 技能的多智能体编排,给出了"权威源 / 派生索引分离"(手稿为权威源,状态跟踪可重生成)的工程范本。
5.4. 治理是前置条件而非事后检查
内容双赛道的 L6 已经从"加分项"变为"上线前置条件":
- 法规基线:《人工智能生成合成内容标识办法》(2025-03-07 印发,2025-09-01 施行)要求视频在起始画面添加显著显式标识、在文件元数据中写入隐式标识;强制性国标 GB 45438—2025 进一步量化:显式标识文字高度不低于画面最短边长度的 5%、正常播放速度下持续不少于 2 秒,元数据字段含 Label / ContentProducer / ProduceID 等。
- 执法实例:2026-04-28,即梦 AI 因未有效落实标识规定被网信部门依法查处——即便一线大厂平台,标识管线也可能断裂。
- 平台治理分化:番茄 2025-09-23 起强制作者申报"是否使用 AI";AI 续写、扩写仅限非保底书作者使用(按合同类型分级授权);晋江文学城仅允许文字校对、创意要素辅助、创意粗纲辅助三种 AI 使用场景。海外平台则主要围绕版权归属(是否主张作品权利、是否用用户作品训练)与隐私(加密、不训练承诺)展开治理。
对工程团队的直接含义:合规标识必须做成生成管线末端的自动节点,与分辨率、帧率同为产物的一等属性,不能依赖人工后期补加。
6. 跨赛道共性判断
6.1. 竞争正从模型能力转向 Harness 完整度
四个赛道的独立证据汇总:
| 赛道 | "模型能力决定论"失效的证据 | 竞争重心的迁移 |
|---|---|---|
| AI IDE | 同一模型在不同脚手架下分差可达两位数百分点;GitHub 官方明确 harness 决定智能的应用效率;Terminal-Bench 排行榜"排系统不排模型" | 模型列表长度 → rules 层级、沙箱粒度、审批策略、可观测性 |
| AI Agents 平台 | 编排原语趋同后,框架差异化转向状态管理与可观测;框架开源免费、收入来自托管与治理 | 编排能力 → 控制平面(checkpointing、上下文治理、审计、成本计量) |
| AI 图像 | 基础生成能力快速商品化(0.2 元/张级 API 定价);"强模型弱 Harness"的妙鸭出局 | 生成质量 → 工作流、资产管理、用量计量 |
| AI 漫剧与小说 | 产能过剩 + 爆款率 0.47%;一致性、连续性问题无法靠换模型解决 | 单次生成上限(L1 至 L3)→ 连续生产上限(L4) |
四个赛道以完全不同的方式收敛到同一个结论:当模型能力越过"可用"阈值后,产品之间的可感知差距主要由 Harness 六层的完成度决定。这与本白皮书第 5 章的评测发现互为表里——既然评测分数本身是"模型 × Harness"的复合产物,市场竞争自然也以同样的复合方式展开。
6.2. L4 与 L5 是全行业共同短板
把四个赛道的六层成熟度做横向归并,可以得出一个高度一致的结论:
| 层 | 跨赛道状态 | 证据 |
|---|---|---|
| L1 上下文工程 | 相对最强 | rules 文件、检索、压缩在各赛道均有成熟实践(AI IDE 三件套、Lorebook、参考图注入、H3-Context-IR) |
| L2 工具与执行 | 强且快速趋同 | 沙箱、MCP、CLI 化、节点化执行原语普及 |
| L3 编排与控制 | 强、同质化中 | 多智能体编排、计划模式、节点 DAG 已成标配 |
| L4 记忆与状态 | 共同短板 | AI 漫剧的资产库形态各异且无统一基准;小说的"吃书"是必现故障;妙鸭的身份资产不可迁移;低代码平台的长期记忆普遍薄弱 |
| L5 评估与观测 | 共同短板 | 多数平台仅有"多候选重生成"式弱评估;本白皮书第 5 章证明 AI SRE / DevOps / 多智能体方向连公开基准都缺位;内容质量无客观标尺 |
| L6 治理与安全 | 两极分化 | 合规驱动型(内容赛道、企业版)强,工具型(Midjourney、部分框架)弱 |
L4 短板的共性根因是:模型本身无状态,所有"一致性"问题本质都是跨调用、跨会话、跨周期的状态注入与持久化问题——角色一致(漫剧)、设定一致(小说)、身份资产可迁移(图像)、会话可恢复(IDE)是同一类工程问题在不同赛道的投影。L5 短板的共性根因则是:内容与开放域任务的产出缺乏客观判定标尺,且公共基准缺位。因此,L4 与 L5 既是全行业的短板,也是未来两年最明确的差异化机会窗口——谁先把"可复现性与资产持久化"(L4)和"效果评估"(L5)做成产品能力而非营销表述,谁就能在各自的赛道拿到 Vidu 主体库、海螺 Context-IR、美图算力点观测那样的差异化位置。
7. 中国与海外市场差异
7.1. 生态结构差异
- 海外:以"模型厂商 + 订阅制 + API 经济"为主轴。分配逻辑是产品力与全球市场覆盖(PixVerse 覆盖 175 个国家、可灵海外收入占比约七成为其增长引擎),定价以美元订阅与用量计费为主。
- 中国:以"大厂渠道闭环 + 免费策略 + 分账生态"为主轴。即梦对接抖音、可灵对接快手,模型 + 工具 + 分发一体;部分国内工具以完全免费策略切入(如字节 Trae、腾讯元宝"目前完全免费使用,暂无收费计划"的公开表态);内容侧形成独特的分账体系(漫剧按"时长 × 单价 × 类型系数 × 版权系数"分成)与保底激励(2026 年抖音集团全年真人短剧保底预算超 15 亿元)。
7.2. 合规环境差异
- 中国:合规是明确的硬约束且已有执法实例——《人工智能生成合成内容标识办法》与 GB 45438—2025 构成量化基线,即梦 2026-04-28 被查处是标志性事件;内容平台另有备案门槛(2026 年 1 月重点微短剧备案门槛由 100 万元提高至 300 万元)与题材审核。
- 海外:EU AI Act 的透明度义务自 2026-08-02 起适用于面向公众的 AI 生成文本标注(开源许可不构成豁免);美国以平台自律与诉讼驱动为主。总体上,中国是"事前标识 + 平台核验"模式,欧盟是"透明度义务"模式,美国是"事后追责"模式。
7.3. 部署形态差异
- 中国:信创与国产芯片适配是真实需求——MiniMax H3 开源发布首日即完成华为昇腾、摩尔线程、沐曦、海光、昆仑芯、天数智芯、壁仞等适配;私有化部署在企业级采购中权重高(Dify 的私有化能力是其核心卖点)。
- 海外:云优先、SaaS 优先,本地部署主要出现在受监管行业与开源社区(ComfyUI 的完全本地离线能力使其同时满足两类需求)。
7.4. 开源权重策略差异
- 中国厂商更积极地把模型权重开源:MiniMax H3(视频模型开源权重)、通义 Qwen 系(编码与图像)、智谱 GLM 系(NovelAI 的 Xialong 即基于 GLM-4.6 微调)——开源权重成为获取开发者生态与国产芯片适配的杠杆。
- 海外头部厂商分层处理:OpenAI 将 Codex CLI 以 Apache 2.0 开源(CLI 层开源、模型闭源);框架层普遍 MIT(LangGraph、CrewAI、AutoGen);前沿模型权重基本闭源。总体上,中国呈现"权重开源换生态"策略,海外呈现"工具与框架开源换标准、权重闭源保利润"策略。
7.5. 差异总览
| 维度 | 中国市场 | 海外市场 |
|---|---|---|
| 生态结构 | 大厂渠道闭环(模型 + 工具 + 分发一体);免费策略切入;分账与保底生态 | 模型厂商 + 订阅制 + API 经济;全球市场覆盖;美元订阅与用量计费 |
| 合规 | 事前标识 + 平台核验(标识办法、GB 45438—2025、备案门槛),已有执法实例 | 欧盟透明度义务(EU AI Act,2026-08-02 适用);美国以平台自律与事后追责为主 |
| 部署形态 | 信创与国产芯片适配为真实需求;私有化部署权重高 | 云优先、SaaS 优先;本地部署集中于受监管行业与开源社区 |
| 开源策略 | 权重开源换生态(MiniMax H3、Qwen 系、GLM 系) | 工具与框架开源换标准(Codex CLI、LangGraph、AGENTS.md),权重闭源保利润 |
| 治理关注点 | 内容标识、备案、题材审核、训练数据授权(番茄协议风波) | 版权归属、隐私承诺(加密、不训练)、青少年与内容政策 |
| 内容变现 | 分账体系 + 保底激励 + 广告投放(投放约占制作链路成本 70%) | 订阅、打赏、版权授权;分发以自有渠道与社区为主 |
需要强调:差异是程度差异而非本质差异。两个市场在"竞争重心从模型能力转向 Harness 完整度"(6.1 节)与"L4 / L5 共同短板"(6.2 节)两个结构性判断上完全一致——差异体现在收敛的速度与约束条件,而非方向。
8. 格局演变的三种可能路径
8.1. 路径一:平台收敛
大厂依托模型、渠道与算力的一体化优势,把 Harness 六层全部产品化,中小玩家退化为模型供应商或垂直插件。支持信号:GitHub 用单一 agentic harness 统一多条产品线;即梦 / 可灵的"模型 + 工具 + 分发"闭环;美图以平台级算力点体系整合多产品线。约束条件:L4 与 L5 恰恰是大厂最难做深的层——资产持久化与效果评估需要长期客户陪伴而非流量打法,垂直玩家仍有机会守住在位优势。
8.2. 路径二:垂直深耕
行业知识与状态层成为护城河,通用平台难以替代。支持信号:AI 漫剧的连续生产能力取决于行业特有的资产库与分镜状态机;网文的"项目宪法"与伏笔台账依赖领域语义;企业级 Agents 采购中 L6(权限、审计、私有化)权重持续上升。妙鸭的失败从反面支持该路径——没有状态层与运营纵深的"通用爆品"生命周期极短。约束条件:垂直市场规模有限,难以摊薄基础模型研发成本,最终多依赖大厂模型 API,议价能力受制于人。
8.3. 路径三:开源生态侵蚀
开源工作流与开源权重从两端挤压闭源 SaaS:ComfyUI 以 60,000+ 节点生态覆盖专业生产,Codex CLI 证明闭源厂商也会主动开源 Harness 层以争夺标准,中国厂商的开源权重策略持续降低自建门槛。若 MCP 等开放协议与 AGENTS.md 等开放标准继续渗透(AGENTS.md 已被 60,000+ 开源项目采用,由 Linux Foundation 旗下 Agentic AI Foundation 托管),六层能力将越来越多地以"可组装的开源组件"形态存在,闭源平台被迫向"托管与治理"(L5 / L6)收窄盈利面。约束条件:开源方案默认缺失评估与治理,企业自建成本高——这反过来又给了商业产品以 L5 / L6 为核心的生存空间。
8.4. 观察指标
三条路径并非互斥,最终格局更可能是三者的混合。对格局走向的判断可依据以下可观察指标:
| 观察指标 | 指向路径一(收敛) | 指向路径二(深耕) | 指向路径三(开源侵蚀) |
|---|---|---|---|
| 六层能力的开源替代速度 | 慢 | 无关 | 快(L1 至 L3 陆续开源化) |
| 企业采购的决策重心 | 平台品牌与集成 | 领域状态层与合规能力 | 可组合性与退出成本 |
| 头部厂商 Harness 层的开放程度 | 封闭闭环 | 部分开放(协议兼容) | 全面开放(标准输出) |
| L4 / L5 能力的产品化进度 | 平台内置且绑定 | 垂直行业深度绑定 | 开源组件 + 商业托管 |
9. 总结
本章基于四个赛道的系统调研,给出三点结论:
- 竞争正从"模型能力"转向"Harness 完整度"。AI IDE 赛道由官方榜单与头部厂商双重确认("排系统不排模型"、harness 决定智能的应用效率);Agents 赛道框架免费、收入流向托管与治理;图像赛道基础生成快速商品化、利润流向工作流与资产管理;内容双赛道产能过剩、确定性稀缺。四个赛道以不同方式收敛到同一结构性变化,其评测侧的镜像就是本白皮书第 5 章的发现——分数从来是"模型 × Harness"的复合产物。
- L4(记忆与状态)与 L5(评估与观测)是全行业共同短板。模型无状态,一切"一致性"都是状态管理问题;内容与开放域任务缺乏客观标尺,公共基准在多个高价值方向缺位。妙鸭相机的完整生命周期失败(强模型、弱 Harness、无状态资产流动性、无评估闭环)与 AI 漫剧 0.47% 的爆款率,从正反两面标定了这两层短板的商业代价;而 Vidu 主体库、海螺 Context-IR、美图算力点观测等案例则标定了补齐短板的差异化回报。
- 中国与海外市场将在三条路径的混合演化中走出不同形状:中国以合规硬约束、渠道闭环与开源权重策略为特征,海外以订阅与 API 经济、透明度义务与工具开源为特征。无论哪种形状,格局的最终决定变量是同一个——谁能把模型的不确定性转化为工程上的可预期性,谁就掌握下一个竞争周期的定价权。
信息缺口声明
- Devin 的 ARR 与估值数据(3,700 万至 4.92 亿美元、102 亿至 260 亿美元)来源为第三方汇总(低置信),全部标注 ,未用于任何行业规模测算。
- 妙鸭相机团队解散(2025-09 底)的权威信源目前仅有第三方工具站表述,未定位到科技媒体报道或工商信息佐证,。
- AI 漫剧市场规模(400 亿元 / 40 亿美元出海)与爆款率数据主要来自 DataEye 报告经二手转引(百度百科、KOCPC 等),未获取原始报告全文交叉验证。
- 各平台 C 端定价存在多套冲突口径(即梦、可灵国内站、美图设计室、Coze 企业版、Vidu 等),本白皮书仅采信官方页可核验口径,其余标注 或不予引用。
- 可灵 3.0 的发布时间与"最长 3 分钟"、白日梦"跨集相似度 95%+"等仅见二手来源,。
- 阅文妙笔业务指标(日活 +100%、token +90%、周使用率超 75%)与美图设计室 Agent Teams 效果均为厂商 self-reported 口径。
- AI 小说与图像赛道多数平台无官方披露的可验证效果指标;第三方测评数据多为营销软文,未予引用。
- 本白皮书未采用任何国内大厂 AI 编码提效的自媒体数据作为硬数据(该类数据在本工程检索中未找到一手出处)。
10. 参考资料
- 2025 Stack Overflow Developer Survey — Stack Overflow,2025-07-30。https://survey.stackoverflow.co/2025/
- State of AI-assisted Software Development 2025 — DORA / Google Cloud,2025-11-12。https://dora.dev/research/2025/dora-report/
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity(随机对照试验) — METR,2025-07-10。https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks — GitHub Blog,2026。https://github.blog/ai-and-ml/github-copilot/evaluating-performance-and-efficiency-of-the-github-copilot-agentic-harness-across-models-and-tasks/
- Terminal-Bench 官方站与 Leaderboard — Stanford / Laude Institute,2025 至 2026。https://www.tbench.ai/
- AGENTS.md 官方站 — Agentic AI Foundation(Linux Foundation 旗下)。https://agents.md/
- Sandboxing: a safer and more autonomous approach — Anthropic,2025。https://www.anthropic.com/engineering/claude-code-sandboxing
- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities — MiniMax 官方博客。https://www.minimax.io/blog/minimax-h3
- 美图公司 2026 年中期业绩公告 — 香港交易所披露易(hkexnews)。https://www.hkexnews.hk/
- 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — 2025-03-07 印发,2025-09-01 施行。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- 强制性国家标准 GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — 全国网络安全标准化技术委员会。https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
- DataEye《2026 上半年 AI 剧 / 漫剧数据报告》 — DataEye研究院(经 KOCPC 等二手转引)。https://en.kocpc.com.tw/archives/25203
- 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — 澎湃新闻,2026。https://www.thepaper.cn/newsDetail_forward_33993493
- 中国青年报《"妙笔通鉴""漫剧助手"发布,AI 赋能网文创作和 IP 改编》 — 中国青年报,2025。https://new.qq.com/rain/a/20251017A08J6J00
- GitHub — forsonny/Claude-Code-Novel-Writer(Multi-Agent Novel Writer v4.1) — 2026。https://github.com/forsonny/Claude-Code-Novel-Writer
- Apache Software Foundation《AI Agents and AGENTS.md Files (Draft)》 — ASF legal-discuss,2026。https://www.apache.org/legal/generative-tooling-agents.html
- 通义万相与 Qwen 系模型阿里云官方定价页 — 阿里云。https://help.aliyun.com/
- Leonardo.ai 官方定价与 API 文档 — Leonardo。https://leonardo.ai/pricing
- Runway 官方定价页 — Runway。https://runwayml.com/pricing
- Comfy 官网 — Comfy Org。https://comfy.org/
Industry Landscape: From Model-Capability Competition to Harness-Completeness Competition
1. Introduction: Four Tracks and One Common Variable
1.1. Research Scope and Analytical Framework
图 1-1|Harness 六层能力模型:竞争重心从模型能力到完整度
数据来源:基于本文分析绘制的示意图。
This chapter, based on systematic research across four content and tool tracks, presents a panoramic view of the industry landscape from the AI Harness perspective:
- AI IDE: agent development environments centered on coding scenarios (Cursor, Claude Code, GitHub Copilot, Codex CLI, Trae, and others);
- AI Agents Platforms: agent frameworks, low-code platforms, and vertical autonomous agents (LangGraph, OpenAI Agents SDK, Claude Agent SDK, Dify, Coze, Devin, Manus, and others);
- AI Image and Content Generation: text-to-image, image editing, and design tools (Midjourney, 即梦, 可灵, Nano Banana, Leonardo, ComfyUI, and others);
- AI Comic-Drama and Novels: long-horizon content production (即梦 AI comic drama, 可灵, Vidu, 白日梦; 阅文妙笔, 番茄, NovelAI, Sudowrite, and others).
The analytical framework follows the Six-Layer Capability Model from Chapter 3 of this whitepaper: instead of treating platforms as "tools" and listing their features, each platform is dissected as a Harness implementation that has already been deployed in production — seeing where it made genuine investment, where it left gaps, and who fills those gaps. All vendor operating data are annotated as self-reported or with a source grade; where measurement scopes conflict, they are presented side by side and marked [To be verified].
1.2. Core Judgment of the Whole Chapter
The four tracks differ enormously in product form, customer base, and business model, yet between 2025 and 2026 they exhibit the same structural shift: the center of gravity of competition is moving from "which model is plugged in" to "the completeness of the Harness six layers". This chapter will argue: this shift has independent evidence in every track; L4 (Memory & State Layer, reproducibility and asset persistence) and L5 (Evaluation & Observability Layer) are shortfalls common to the whole industry; and the Chinese and overseas markets will follow different convergence paths along three dimensions — compliance, deployment form, and open-source strategy.
2. AI IDE Track
2.1. Market Overview: Adoption Saturated, Trust Has Not Kept Pace
The AI IDE is the most fiercely contested track, and the one that best verifies that "model capability does not equal engineering usability". Three sets of authoritative survey data sketch out the basic shape of the market:
| Metric | Value | Source (scope noted) |
|---|---|---|
| Developer adoption | 84% are using or planning to use AI tools (76% in 2024); 51% of professional developers use them daily | Stack Overflow 2025 Developer Survey, 2025-07-30, 49,000+ responses |
| Trust | 46% do not trust the accuracy of AI output (31% in 2024); only 3.1% trust it highly | Stack Overflow 2025 |
| Agent-form penetration | About 31% already use AI agents; 37.9% do not plan to | Stack Overflow 2025 |
| Main frustrations | 66% think "AI solutions are almost right but not entirely right"; 45.2% think debugging AI code takes longer | Stack Overflow 2025 |
| Enterprise-side adoption | Over 90% use AI (a separate 95% figure exists); 80% perceive individual productivity gains; 30% have single tasks exceeding 4 hours | DORA 2025, 2025-11-12 |
| Measured efficiency | Experienced developers measured 19% slower when using AI, but self-assessed 20% faster | METR randomized controlled trial, 2025-07-10 |
The implication of these numbers is: market education is complete, but trust-building has not yet begun. The simultaneous rise in adoption and distrust shows that the value of the tools is acknowledged, while their output has not yet attained engineering predictability. "66% think it is almost right but not entirely right" describes not a capability problem but a verifiability problem — precisely what the L5 layer of the Harness is meant to solve.
2.2. Representative Platform Comparison Matrix
| Platform | Developer | Form | Open Source | Notable Six-Layer Harness Characteristics |
|---|---|---|---|---|
| Cursor | Anysphere | VS Code fork IDE | No | L1 rules hierarchy + codebase indexing; L6 team marketplace + SSO + audit logs + code-tracking API |
| Claude Code | Anthropic | Terminal + IDE + Web | No | L2 fine-grained sandboxing + credential masking; L1 proximity-based rules + progressive disclosure of skills + compaction; L3 plan mode + sub-agents + hooks |
| GitHub Copilot | Microsoft / GitHub | Multi-IDE plugins + CLI + Web | No | L6 organization policies + content exclusion + auditing + IP indemnification; L5 platform-side code review and usage dashboards; a shared agentic harness (the single Copilot SDK component) drives multiple product lines |
| Codex CLI | OpenAI | Terminal + IDE extension + Web | Yes (Apache 2.0) | L2 orthogonal configuration of OS-level sandboxing and approvals; L1 AGENTS.md hierarchy; L3 non-interactive execution and session recovery |
| Windsurf | Cognition | VS Code fork IDE | No | L4 memory system + checkpoint rollback as the core differentiator; L3 plan mode + parallel sessions |
| Trae | 字节跳动 | VS Code fork IDE + Web | No | L1 context compaction + rules engineering; L3 multi-agent; entering on price and the domestic model ecosystem |
The tone-setting evidence for this track: GitHub officially stated in 2026 that "the model provides the raw intelligence, while the harness determines how effectively that intelligence is applied" (the harness shapes how effectively that intelligence is applied), and published the controlled-comparison methodology of its agentic harness; the official Terminal-Bench also made clear that its leaderboard "ranks systems, not models". The dual confirmation by the official leaderboard and top vendors elevates "Harness completeness determines product gaps" from an industry intuition to a citable engineering conclusion (methodological details in Chapter 5 of this whitepaper).
2.3. Competitive Focus
The competitive focus of the AI IDE track has clearly migrated:
- L1 Context Engineering is the first competitive dimension. A medium-sized repository holds hundreds of thousands of lines of code, and no context window is large enough to hold them all. The three-piece set of rules file hierarchy (loaded close by), semantic codebase indexing, and context compaction determines how complex a requirement a single task can carry.
- L6 Governance is an entry requirement, not a bonus item. The higher the autonomy, the larger the blast radius of an incident — the Amazon Q prompt-injection incident of 2025-08-11 (malicious instructions attempting to induce the agent to delete AWS resources) was the first large-scale public exposure of the structural weakness that "external content being processed can directly become instructions". Filesystem isolation and network isolation are both indispensable, credentials are invisible by default, and destructive commands require hard-coded interception.
- Positive evidence shows that governance and autonomy are positive-sum: Anthropic officially disclosed that its sandboxing cut internal permission prompts by 84% while improving safety — constraints make the speed of scaling possible.
2.4. Business Models: Three Billing Logics
| Billing Logic | Representative | Advantage | Risk |
|---|---|---|---|
| Subscription-quota model | Claude Code (with Claude subscription), Codex CLI (with ChatGPT subscription) | Predictable cost | Hitting limits at peak times; long tasks forced to interrupt |
| Credit / usage-based model | Cursor (quota pool + usage-based), GitHub Copilot (AI Credits from 2026-06-01, 1 credit = USD 0.01) | A single conversation and a longer agent session no longer cost the same; fairer | Cost uncertainty; requires watching the burn rate |
| Quota-refresh model | Windsurf (daily / weekly quota refresh from 2026-03) | Easy budgeting | Per-run ceiling constrains complex tasks |
The three logics are converging toward a "subscription + usage" hybrid. The migration of billing models means buyers need new management capabilities: organization-level consumption-rate observation and budget alerts. The GitHub Copilot case also shows a product-line consolidation trend — its agentic harness, as the single shared component of the Copilot SDK, simultaneously drives multiple experiences including CLI, App, and code review; the positioning of "the Harness as a platform asset" is becoming ever clearer.
3. AI Agents Platform Track
3.1. Market Overview: A Three-Tier Supply Structure
The supply of AI Agents platforms has a clear three-tier structure, and the three tiers follow different Harness-completeness logics:
| Tier | Representative | Supply Logic | Harness Coverage |
|---|---|---|---|
| Code frameworks | LangGraph, OpenAI Agents SDK, Claude Agent SDK, CrewAI, Microsoft Agent Framework, AutoGen | Aimed at engineering teams; open source or free; primarily covers L2 / L3 | L4 / L5 / L6 mostly left blank, built by users themselves or plugged in externally (e.g., LangSmith provides tracing and evaluation) |
| Low-code platforms | Dify, Coze (扣子), 腾讯元器 | Aimed at business teams; visual orchestration + managed runtime | L1 through L3 productized and packaged; L5 observability and L6 tenant governance tiered with the enterprise edition |
| Vertical autonomous agents | Devin (Cognition), Manus | Aimed at end-to-end task delivery; billed by workload | The six-layer loop encapsulated inside the product; the most black-box |
On growth signals (all vendor-side or third-party figures, to be cited with caution): Devin's ARR, per third-party compilation, grew from USD 37 million (2025-05) to USD 492 million (2026-05), and the valuation figure rose from USD 10.2 billion to USD 26 billion — the source is third-party aggregation (low confidence), [To be verified], and this whitepaper does not use it as a basis for industry-size estimation; the operating milestones officially disclosed by Devin (processing 147 trillion tokens and driving 80 million virtual computers in the 8 months since launch) are likewise self-reported. The certain fact on the framework side is: core frameworks are generally open source and free (LangGraph, CrewAI, and AutoGen are all MIT), and revenue is shifting to the managed and observability layers (LangSmith free tier 5,000 traces/month, Developer tier USD 39/user/month).
3.2. Representative Platform Comparison Matrix
| Platform | Category | Open Source / License | Billing Form | Governance & Compliance Highlights |
|---|---|---|---|---|
| LangGraph | Code framework | MIT open source | Framework free, model API fees only | Checkpointing capability the most complete among comparable frameworks |
| OpenAI Agents SDK | Code framework | Open source | Framework free; default tracing dashboard requires an OpenAI platform account | Deeply coupled to the model vendor's runtime |
| Claude Agent SDK | Code framework | Closed-source SDK | No license fee, billed by Claude API tokens | Built-in compaction, sub-agent context isolation, hooks |
| CrewAI | Code framework | MIT open source | AMP platform free tier 50 workflow runs/month, Enterprise custom | Enterprise tier provides FedRAMP High, SSO, RBAC (official figure) |
| Dify | Low-code platform | Open source (partial) + commercial edition | Free + cloud subscription | Mature private deployment; USD 30 million Series Pre-A in 2026-03 (third-party relay) |
| Coze (扣子) | Low-code platform | Domestic edition closed source, overseas edition has an open-source version (conflicting figures) | Free tier + token overage | Enterprise subscription about RMB 20,000 to 200,000/year (third-party review figure) |
| Devin | Vertical autonomous agent | No | Core USD 20/month + USD 2.25/ACU (1 ACU ≈ 15 minutes of active work); Team USD 500/month including 250 ACUs | Enterprise tier provides VPC deployment, SAML/OIDC SSO, centralized control (official figure) |
| Manus | Vertical autonomous agent | No (L1/L4/L5 mechanisms closed source, unverifiable) | Freemium + credit consumption (Starter from USD 20/month) | Fewest publicly available governance details |
3.3. Competitive Focus and Business Models
The competitive focus of this track shows two directions:
- Competition at the framework layer is about the "control plane". Once orchestration primitives (loops, sub-agents, tool calls) converged, differentiation shifted to state management (LangGraph's checkpointing), context governance (Claude Agent SDK's compaction), and observability (LangSmith). This is consistent with the rising weight of L3 and L4 in the six-layer model.
- Competition at the platform and vertical tiers is about L6 and cost metering. The core question in enterprise procurement has shifted from "can we orchestrate" to "how do we manage permissions, how do we export audit logs, how do we cap costs". Devin bills by ACU (active work time), Manus by credits, Dify / Coze by runs and tokens — the three forms arrive at the same destination: translating unpredictable token consumption into budgetable units of workload, which is essentially the productization of L5 observability capability.
The structural fact of the business models: the framework layer does not make money; what makes money is the runtime, observability, and governance. This distribution corroborates the judgment of this whitepaper — value is moving from the model-access layer to the Harness layer.
Same-Day Increment (2026-09-12; corrected 2026-09-13): The OpenAI Agents API entered public beta on 2026-09-10 (grade A, per the OpenAI official Changelog; the 2026-09-12 snapshot had recorded it as 09-11 per media reports), putting the orchestration capability of the existing Agents SDK under management — together with AWS Bedrock AgentCore GA (2026-06) this forms two moves in the same direction, marking that the competitive focus of this track is shifting from "are the orchestration primitives good to use" to "is the managed runtime governable". For details, see the research library: 03-市场研究/02-AI-Agents组/21-openai-agents-api.md.
4. AI Image and Content Generation Track
4.1. Market Overview
The image track is the most commercially mature of the four. Take 美图公司 as an example (HKEX announcement figure, very high confidence): total revenue for H1 2026 was RMB 2.21 billion (+22.1% year on year), of which imaging and design product revenue was RMB 1.77 billion (+30.9%, 80% of the total); adjusted net profit attributable to the parent was RMB 650 million (+39.5%); MAU was 282 million, with over 18.44 million paid subscribers (+19.7%) and a subscription penetration rate of 6.5%; productivity-app MAU was 33 million (+43.5%), with 2.35 million paid subscribers and ARR of about RMB 620 million. 美图's observability design is especially noteworthy: it uses "the quarter-over-quarter growth in AI compute-point consumption" (over 46% in both Q1 and Q2) as a proxy indicator for value hit-rate — measuring value hit-rate by consumption depth is a rarely mature practice of L5 observability in the content track.
The structural change on the supply side is equally significant: Google is injecting top-tier image capability into a general-purpose API via Nano Banana (Gemini-family image model) (first generation 1K/2K single image USD 0.039, Pro series 1K/2K USD 0.134, compiled from screenshots of the official pricing page), while domestic 通义万相 is waging price competition with API pricing of RMB 0.2 per image (Alibaba Cloud official) — basic image generation is being rapidly commoditized, and profit is migrating to workflows, asset management, and industry solutions (the Harness layer).
4.2. Representative Platform Comparison Matrix
| Platform | Developer | Form | Pricing Form | Six-Layer Harness Positioning |
|---|---|---|---|---|
| Midjourney | Midjourney | Web / Discord | Subscription USD 10 to 120/month | The archetype of strong model, weak Harness: L5 relies on human-eye subjective evaluation and community feedback; no official regression-set mechanism observed |
| Nano Banana (Gemini-family) | API + in-product embedding | Billed by token / by image | Distributes capability via API; the Harness is borne by upper-layer applications | |
| 即梦 AI | 字节跳动 | SaaS + API | Consumer membership (two conflicting third-party figures) | Model + creation tools + 抖音 distribution closed loop |
| 可灵 AI | 快手 | SaaS + API | International site USD 6.99 to 127.99/month (official page); domestic site per secondary-source figure | Element reference + series generation supporting consistency; commercial rights tiered |
| Leonardo.ai | Leonardo | Web + API | Subscription USD 12 to 60/month + PAYG API (USD 5 credit for new accounts, up to 10 concurrency) | Three types of capacity constraints explicitly documented + Webhook + pricing calculator; highly engineered |
| Runway | Runway | Web + API | Subscription USD 12 to 76/month (annual); API USD 0.01/credit | Video / image hybrid workflows, Adobe integration |
| ComfyUI | Comfy Org (open-source community) | Local node workflows | Free (GPL-3.0) | 60,000+ node ecosystem; workflow JSON is the state, versionable; all six layers programmable and self-buildable |
| WeShop / 美图设计室 | 美图 and others | SaaS | Subscription + compute points | Agent Teams multi-agent collaboration + creative asset reuse (self-reported) |
4.3. Competitive Focus and Business Models
Two structural observations on this track:
- The watershed of the degree of Harness-ification is at L3 and L4. 美图设计室 (multi-agent collaboration + creative asset reuse + compute-point metering), ComfyUI (workflows as JSON graphs, version-controllable, SDK-able), and Leonardo (explicit capacity constraints + Webhook + cost estimation) represent the third-generation form; Midjourney and 妙鸭 represent "strong model, weak Harness" — leading model quality but thin engineering carrying capacity, and 妙鸭 has already exited for this reason (see Section 4.4).
- Business models move from one-time payment to a subscription + consumption hybrid. Midjourney is pure subscription, Runway pure credits, Leonardo dual-track subscription + PAYG, 美图 subscription + compute points. Consumption-based pricing presupposes precise usage metering and alerting — another instance of an L5 capability becoming commercial infrastructure.
4.4. A Typical Failure Case: 妙鸭相机
妙鸭相机 is the only negative sample in this track with a complete lifecycle, and a textbook case of "strong model, weak Harness":
- Explosion: launched 2023-07-17, topped the App Store overall chart in August 2023, daily actives broke 600,000, with over 4,000 users queued at peak to generate photos; the technical base was the "提香" model self-developed by 阿里大文娱.
- Engineering shortfall: the L2 / L3 / L5 layers were almost blank — every interaction was an atomic operation (upload, wait, pick a template, produce the photo), with no intermediate state to intervene in and no parameters to adjust; the user's only quality-correction entry point was a "make it more like me" similarity fine-tune. L6 saw an early user-agreement trust incident (the official apology and revision).
- Commercial shortfall: the digital avatar was a one-time RMB 9.9 (promotional price) buyout + 10 finished photos — typical traffic-driver pricing, lacking follow-on consumption scenarios and a subscription anchor; compared with similar identity products (可灵 caps subjects at 30 to 500 by membership tier) it lacked asset-operations depth.
- Tech-route lock-in: the training-based identity pipeline (≥20 photos per user, full per-user training, hours of queuing) did not complete its migration after zero-shot identity-injection solutions (InstantID, PuLID) matured in 2024–2025, and the cost structure was locked in. From the Harness perspective: the training-based route made identity assets into "private state inside the product" — non-migratable, non-composable, non-auditable; the zero-shot route made identity assets into "context that can circulate across pipelines", gaining real asset liquidity.
- Ending: the team formally disbanded at the end of September 2025 (the authoritative source for this fact is, to date, only the statement of a third-party tool site), and the product continues only at a minimal level of operation.
The failure proposition of 妙鸭 is worth writing into the project review of every content product: given that generation quality is already good enough, what determines a product's life or death is not the model but the Harness that carries it.
5. AI Comic-Drama and Novel Track
5.1. AI Comic-Drama: Overcapacity and Scarce Certainty
AI comic drama is the scene in the AIGC content industry where engineering has advanced fastest and where the contradiction that "model capability does not equal production usability" was exposed earliest. Industry data for H1 2026 (DataEye report and 中国网络视听协会 data, via secondary relay; the full original report was not cross-verified):
| Metric | Value | Notes |
|---|---|---|
| Total micro-short-drama releases | 367,000 titles, over 74% AI content | 中国网络视听协会 data (relayed) |
| Douyin-native new AI dramas and comic dramas | 221,900 titles, 515.738 billion cumulative plays | — |
| Hit rate | Only 1,055 titles exceeded 100 million plays, about 0.47% | DataEye |
| Break-even rate | Under 1.3%, roughly 1 of every 77 titles recoups | Measured against a break-even line of 50 million plays |
| Release speed | On average one title every 36 seconds | DataEye |
| Revenue per 10,000 plays | Fell from a peak of RMB 30 to 100 to RMB 5 to 10, a drop of over 90% | DataEye |
| Market size | Expected to break RMB 40 billion in 2026 (+138%); overseas expansion expected to exceed USD 4 billion | 中研普华 / DataEye relay, [To be verified] |
Capacity is piling up at a rate of one title every 36 seconds, the hit rate is under 0.5%, and unit prices have dropped by 90% — the industry has moved from "racing for capacity" to "racing for certainty". And certainty is precisely what models cannot provide: the same character looks different in episode 1 and episode 47, the same scene has lighting jump between two shots, and plot state is lost when writing continues across sessions. Not one of these problems can be solved by swapping in a stronger model; they can only be solved by the Harness's state management (L4), orchestration (L3), and regression evaluation (L5).
5.2. AI Comic-Drama Representative Platform Comparison Matrix
| Platform | Developer | Competitive Lever | L4 (Memory / Assets) Implementation | L6 Governance |
|---|---|---|---|---|
| 即梦 AI | 字节跳动 | Model + tools + 抖音 distribution closed loop; multi-shot narrative | Stable retention of character features; asset-library figure not disclosed | Real-person avatar authentication; on 2026-04-28 investigated by the cyberspace authority for failing to effectively implement AI-generated synthetic content labeling (publicly recorded penalty event) |
| 可灵 AI | 快手 | Intelligent storyboard + element reference; overseas revenue about 70% (third-party figure) | Series generation + extended consistency | Commercial rights unlocked by tier |
| PixVerse | 爱诗科技 | CLI / Skills compatible with coding agents such as Claude Code and Codex; covers 175 countries | Character-consistent characters + Team asset library | Team Plan RBAC + credit cap + usage analytics |
| 海螺 AI | MiniMax | H3-Context-IR context compression (about 100k tokens compressed to about 4k) + open weights | @ reference system + In-Context Regeneration | Open weights are self-auditable; adaptation to multiple domestic chips completed |
| Vidu | 生数科技 | Reference-to-video "everything can be referenced" | Subject library (characters / props / scenes), up to 7 reference images | SaaS / MaaS tiering |
| 白日梦 AI | 光魔科技 | Permanent character library, claims 95%+ cross-episode similarity | The character library is the core selling point | Public governance information missing |
| ComfyUI | Open-source community | 60,000+ nodes, fully local offline | Character LoRA + style anchors; workflow JSON is the state, versionable | GPL-3.0, labeling pipeline can be self-built (compliance self-build example) |
| 豆包 / Seedance | 字节 Seed | MaaS API output | No native asset library; upper layers must build it themselves | Real-person face ban + real-person verification for digital avatars (one of the strictest publicly known compliance samples) |
The layering rule of this track is clear: L1 through L3 (four-modality input, intelligent storyboard, multi-shot narrative, node workflows) have by 2026 become standard equipment on leading platforms, highly homogenized; L4 has not yet homogenized — from "up to 7 reference images" (Vidu) to "about 100k tokens compressed to about 4k tokens" (海螺 H3) to "workflow JSON is the state" (ComfyUI), the implementation paths are entirely different. Whoever builds L4 solidly (character anchoring, asset library, shot state machine) will be able to push AI comic drama from "can generate" to "can continuously produce a hundred episodes".
5.3. AI Novels: The Long-Horizon State-Management Race
AI novels are the content scene where L4 and L1 bear the greatest pressure in the six-layer model: the core difficulty of long serialized fiction is "not collapsing character settings or forgetting foreshadowing after hundreds of thousands of words" (the web-novel industry names this guaranteed-to-occur fault "eating the book"). The industry has converged on three L4 paradigms:
| Paradigm | Mechanism | Representative Implementation |
|---|---|---|
| A · Structured rolling summary | When context overflows, automatically compressed into a structured summary, retaining key plot threads | The rolling-summary layer of Claude's long-form workflow |
| B · Keyword-triggered entry library | Settings broken into entries, injected into context only when relevant | NovelAI Lorebook; AI Dungeon Story Cards |
| C · Versioned state snapshot | After each chapter, state is extracted and archived as a versioned snapshot, revisitable on demand | Claude Book's state/current/ symbolic link + per-chapter archiving; 灵蟹创作's "project constitution" |
The platform landscape shows "Chinese and overseas platforms solving the same problem differently": 阅文妙笔 enters with "妙笔通鉴" in a "second brain" positioning (specializing in foreshadowing mining and detail retrieval); its disclosed business indicators are Miaobi daily actives more than doubling, daily-average token consumption up over 90%, and weekly author usage over 75% (vendor self-reported); 番茄小说 is known for the strongest L6 (mandatory AI declaration, expansion and continuation writing banned for guaranteed-minimum books, machine judgment of low-quality content); NovelAI stands in the subscription market on Lorebook and privacy-first (XSalsa20 client-side encryption, no training); Sudowrite supports English long-form with Story Bible + reading the first 20,000 words + 25-document chapter continuity. On the community side, the GitHub open-source project Claude-Code-Novel-Writer (v4.1, 2026-08-21) provides, through multi-agent orchestration of AGENTS.md + 7 roles + 5 skills, an engineering model of "authoritative source / derived index separation" (the manuscript is the authoritative source; state tracking is regenerable).
5.4. Governance Is a Prerequisite, Not an After-the-Fact Check
L6 in both content tracks has already shifted from a "bonus item" to a "precondition for launch":
- Regulatory baseline: the 《人工智能生成合成内容标识办法》 (issued 2025-03-07, in effect 2025-09-01) requires videos to carry a prominent explicit label on the opening frame and an implicit label written into the file metadata; the mandatory national standard GB 45438—2025 quantifies it further: explicit label text height no less than 5% of the shortest edge of the frame, lasting no less than 2 seconds at normal playback speed, with metadata fields including Label / ContentProducer / ProduceID, etc.
- Enforcement instance: on 2026-04-28, 即梦 AI was lawfully investigated by the cyberspace authority for failing to effectively implement the labeling rules — even for a top-tier big-company platform, the labeling pipeline can break.
- Platform governance divergence: from 2025-09-23, 番茄 mandates authors to declare "whether AI is used"; AI continuation and expansion writing are limited to authors of non-guaranteed-minimum books (authorization tiered by contract type); 晋江文学城 permits only three AI use scenarios: text proofreading, creative-element assistance, and creative rough-outline assistance. Overseas platforms govern mainly around copyright ownership (whether work rights are claimed, whether user works are used for training) and privacy (encryption, no-training commitments).
The direct implication for engineering teams: compliance labeling must be built as an automatic node at the tail of the generation pipeline, a first-class attribute of the artifact alongside resolution and frame rate, and cannot rely on manual post-hoc addition.
6. Cross-Track Common Judgments
6.1. Competition Is Shifting from Model Capability to Harness Completeness
Aggregation of independent evidence across the four tracks:
| Track | Evidence Against "Model-Capability Determinism" | Shift of the Competitive Center of Gravity |
|---|---|---|
| AI IDE | The same model can differ by double-digit percentage points under different scaffolds; GitHub officially makes clear that the harness determines the efficiency of intelligence application; the Terminal-Bench leaderboard "ranks systems, not models" | Length of model list → rules hierarchy, sandbox granularity, approval policy, observability |
| AI Agents Platforms | After orchestration primitives converged, framework differentiation shifted to state management and observability; frameworks are open source and free, revenue comes from managed services and governance | Orchestration capability → control plane (checkpointing, context governance, audit, cost metering) |
| AI Image | Basic generation capability rapidly commoditized (API pricing at the RMB 0.2-per-image level); 妙鸭, the "strong model, weak Harness" case, has exited | Generation quality → workflows, asset management, usage metering |
| AI Comic-Drama and Novels | Overcapacity + 0.47% hit rate; consistency and continuity problems cannot be solved by swapping models | Per-generation ceiling (L1 through L3) → continuous-production ceiling (L4) |
The four tracks converge on the same conclusion in entirely different ways: once model capability crosses the "usable" threshold, the perceptible gap between products is determined mainly by the completeness of the Harness six layers. This and the evaluation findings of Chapter 5 of this whitepaper are two sides of the same coin — since an evaluation score is itself a composite product of "model × Harness", market competition naturally unfolds in the same composite way.
6.2. L4 and L5 Are Industry-Wide Shared Shortfalls
Cross-aggregating the six-layer maturity of the four tracks yields a highly consistent conclusion:
| Layer | Cross-Track Status | Evidence |
|---|---|---|
| L1 Context Engineering | Relatively strongest | Rules files, retrieval, and compaction have mature practice in every track (the AI IDE three-piece set, Lorebook, reference-image injection, H3-Context-IR) |
| L2 Tooling & Execution | Strong and rapidly converging | Sandboxing, MCP, CLI-ification, and node-based execution primitives have become widespread |
| L3 Orchestration & Control | Strong, homogenizing | Multi-agent orchestration, plan mode, and node DAGs are now standard equipment |
| L4 Memory & State | Shared shortfall | AI comic-drama asset libraries take various forms with no unified baseline; the "eating the book" failure is guaranteed in novels; 妙鸭's identity assets are non-migratable; long-term memory on low-code platforms is generally weak |
| L5 Evaluation & Observability | Shared shortfall | Most platforms have only "multi-candidate regeneration"-style weak evaluation; Chapter 5 of this whitepaper shows that in the AI SRE / DevOps / multi-agent directions even public benchmarks are absent; there is no objective yardstick for content quality |
| L6 Governance & Safety | Polarized | Compliance-driven (content tracks, enterprise editions) is strong; tool-type (Midjourney, some frameworks) is weak |
The common root cause of the L4 shortfall is: the model itself is stateless, and every "consistency" problem is essentially a cross-call, cross-session, cross-cycle state-injection and persistence problem — character consistency (comic drama), setting consistency (novels), migratable identity assets (image), and recoverable sessions (IDE) are projections of the same kind of engineering problem onto different tracks. The common root cause of the L5 shortfall is: the output of content and open-domain tasks lacks an objective yardstick for judgment, and public benchmarks are absent. Therefore, L4 and L5 are both an industry-wide shortfall and, in the coming two years, the clearest differentiation window — whoever first turns "reproducibility and asset persistence" (L4) and "effect evaluation" (L5) into product capabilities rather than marketing claims will win, in their own track, a differentiated position like Vidu's subject library, 海螺's Context-IR, or 美图's compute-point observability.
7. Differences Between the Chinese and Overseas Markets
7.1. Ecosystem Structure
- Overseas: the main axis is "model vendors + subscription + API economy". The allocation logic is product strength and global market coverage (PixVerse covers 175 countries; 可灵's roughly 70% overseas revenue share is its growth engine), with pricing mainly USD subscriptions and usage billing.
- China: the main axis is "big-company channel closed loop + free strategy + revenue-sharing ecosystem". 即梦 connects to 抖音, 可灵 connects to 快手, with model + tools + distribution integrated; some domestic tools enter on a fully free strategy (e.g., 字节 Trae, and 腾讯元宝's public statement that it is "currently fully free to use, with no plans to charge yet"); on the content side a distinctive revenue-sharing system has formed (comic drama shared by "duration × unit price × type coefficient × copyright coefficient") along with guaranteed-minimum incentives (the 2026 full-year guaranteed-minimum budget for live-action short dramas at the 抖音 group exceeds RMB 1.5 billion).
7.2. Compliance Environment
- China: compliance is an explicit hard constraint with enforcement instances already on record — the 《人工智能生成合成内容标识办法》 and GB 45438—2025 form a quantified baseline, and 即梦 being investigated on 2026-04-28 is the landmark event; content platforms additionally face filing thresholds (in January 2026 the filing threshold for key micro-short-dramas was raised from RMB 1 million to RMB 3 million) and subject-matter review.
- Overseas: the transparency obligations of the EU AI Act apply from 2026-08-02 to labeling of AI-generated text directed at the public (an open-source license does not constitute an exemption); the United States is driven mainly by platform self-regulation and litigation. Overall, China follows a "pre-labeling + platform verification" model, the EU a "transparency obligation" model, and the United States an "after-the-fact accountability" model.
7.3. Deployment Form
- China: Xinchuang (信创) and domestic-chip adaptation are real demands — on the first day of MiniMax H3's open-source release, adaptation was completed for 华为昇腾, 摩尔线程, 沐曦, 海光, 昆仑芯, 天数智芯, 壁仞, and others; private deployment carries high weight in enterprise procurement (Dify's private-deployment capability is its core selling point).
- Overseas: cloud first, SaaS first; local deployment appears mainly in regulated industries and the open-source community (ComfyUI's fully local offline capability lets it satisfy both kinds of need at once).
7.4. Open-Source Weight Strategies
- Chinese vendors are more aggressive about open-sourcing model weights: MiniMax H3 (open weights for a video model), 通义 Qwen family (coding and image), 智谱 GLM family (NovelAI's Xialong is fine-tuned from GLM-4.6) — open weights have become a lever for acquiring the developer ecosystem and domestic-chip adaptation.
- Overseas top vendors handle it in tiers: OpenAI open-sourced Codex CLI under Apache 2.0 (CLI layer open source, model closed source); the framework layer is generally MIT (LangGraph, CrewAI, AutoGen); frontier model weights are essentially closed source. Overall, China shows an "open weights in exchange for ecosystem" strategy, and the overseas shows an "open-source tools and frameworks in exchange for standards, closed weights to protect profit" strategy.
7.5. Overview of Differences
| Dimension | Chinese Market | Overseas Market |
|---|---|---|
| Ecosystem structure | Big-company channel closed loop (model + tools + distribution integrated); free-strategy entry; revenue-sharing and guaranteed-minimum ecosystem | Model vendors + subscription + API economy; global market coverage; USD subscriptions and usage billing |
| Compliance | Pre-labeling + platform verification (labeling measures, GB 45438—2025, filing threshold), with enforcement instances already on record | EU transparency obligations (EU AI Act, applicable 2026-08-02); the United States mainly platform self-regulation and after-the-fact accountability |
| Deployment form | Xinchuang (信创) and domestic-chip adaptation are real demands; private deployment carries high weight | Cloud first, SaaS first; local deployment concentrated in regulated industries and the open-source community |
| Open-source strategy | Open weights in exchange for ecosystem (MiniMax H3, Qwen family, GLM family) | Open-source tools and frameworks in exchange for standards (Codex CLI, LangGraph, AGENTS.md); closed weights to protect profit |
| Governance focus | Content labeling, filing, subject-matter review, training-data authorization (the 番茄 agreement uproar) | Copyright ownership, privacy commitments (encryption, no training), youth and content policy |
| Content monetization | Revenue-sharing system + guaranteed-minimum incentives + ad placement (placement accounts for about 70% of production-chain cost) | Subscriptions, tips, copyright licensing; distribution mainly via own channels and communities |
It must be emphasized: the differences are differences of degree, not of essence. The two markets are fully aligned on the two structural judgments that "the center of gravity of competition is shifting from model capability to Harness completeness" (Section 6.1) and "L4 / L5 are shared shortfalls" (Section 6.2) — the differences show up in the speed of convergence and the constraints, not in the direction.
8. Three Possible Paths of Landscape Evolution
8.1. Path One: Platform Convergence
Big companies, leveraging their integrated advantage in models, channels, and compute, productize all six Harness layers, and small and medium players degrade into model suppliers or vertical plugins. Supporting signals: GitHub unifies multiple product lines with a single agentic harness; the "model + tools + distribution" closed loops of 即梦 / 可灵; 美图 integrating multiple product lines with a platform-level compute-point system. Constraint: L4 and L5 are precisely the layers hardest for big companies to go deep on — asset persistence and effect evaluation require long-term customer companionship, not traffic tactics, and vertical players still have a chance to hold their incumbent advantages.
8.2. Path Two: Vertical Specialization
Industry knowledge and the state layer become the moat, and general-purpose platforms are hard to replace. Supporting signals: the continuous-production capability of AI comic drama depends on the industry-specific asset library and storyboard state machine; the "project constitution" and foreshadowing ledger of web novels depend on domain semantics; in enterprise Agents procurement the weight of L6 (permissions, audit, privatization) keeps rising. 妙鸭's failure supports this path from the negative side — a "general-purpose hit product" without a state layer and operational depth has an extremely short lifecycle. Constraint: the vertical market size is limited, making it hard to amortize foundation-model R&D costs; in the end it mostly depends on big-company model APIs, and bargaining power is constrained by others.
8.3. Path Three: Open-Source Erosion
Open-source workflows and open-source weights squeeze closed-source SaaS from both ends: ComfyUI covers professional production with a 60,000+ node ecosystem; Codex CLI proves that closed-source vendors will proactively open-source the Harness layer to contest standards; Chinese vendors' open-weight strategy keeps lowering the bar for self-building. If open protocols such as MCP and open standards such as AGENTS.md keep penetrating (AGENTS.md has been adopted by 60,000+ open-source projects, hosted by the Agentic AI Foundation under the Linux Foundation), the six-layer capability will increasingly exist in the form of "assemblable open-source components", and closed-source platforms will be forced to narrow their profit surface toward "managed services and governance" (L5 / L6). Constraint: open-source solutions lack evaluation and governance by default, and enterprise self-building is costly — this in turn gives commercial products a space for survival built around L5 / L6.
8.4. Observation Indicators
The three paths are not mutually exclusive; the final landscape is more likely a mixture of all three. Judgments on where the landscape is heading can rest on the following observable indicators:
| Observation Indicator | Points to Path One (Convergence) | Points to Path Two (Specialization) | Points to Path Three (Open-Source Erosion) |
|---|---|---|---|
| Speed of open-source substitution of the six-layer capability | Slow | Irrelevant | Fast (L1 through L3 successively open-sourced) |
| Decision center of enterprise procurement | Platform brand and integration | Domain state layer and compliance capability | Composability and exit cost |
| Openness of top vendors' Harness layer | Closed loop | Partially open (protocol-compatible) | Fully open (standard output) |
| Productization progress of L4 / L5 capability | Built into and bound to the platform | Deeply bound to the vertical industry | Open-source components + commercial managed services |
9. Conclusion
Based on systematic research across the four tracks, this chapter draws three conclusions:
- Competition is shifting from "model capability" to "Harness completeness". The AI IDE track is doubly confirmed by the official leaderboard and top vendors ("ranks systems, not models"; the harness determines the efficiency of intelligence application); in the Agents track frameworks are free and revenue flows to managed services and governance; in the image track basic generation is rapidly commoditized and profit flows to workflows and asset management; in both content tracks there is overcapacity and scarce certainty. The four tracks converge on the same structural change in different ways, and its mirror on the evaluation side is the finding of Chapter 5 of this whitepaper — a score has always been a composite product of "model × Harness".
- L4 (Memory & State) and L5 (Evaluation & Observability) are industry-wide shared shortfalls. The model is stateless, and every "consistency" is a state-management problem; content and open-domain tasks lack an objective yardstick, and public benchmarks are absent in multiple high-value directions. The complete-lifecycle failure of 妙鸭相机 (strong model, weak Harness, no state-asset liquidity, no evaluation loop) and the 0.47% hit rate of AI comic drama mark the commercial cost of these two layers' shortfalls from both the positive and negative sides; cases such as Vidu's subject library, 海螺's Context-IR, and 美图's compute-point observability mark the differentiated return of making up the shortfalls.
- China and overseas markets will take different shapes in the mixed evolution of the three paths: China is characterized by hard compliance constraints, channel closed loops, and an open-weight strategy; the overseas by a subscription and API economy, transparency obligations, and tool open-sourcing. Whatever the shape, the ultimate decisive variable of the landscape is the same — whoever can convert the model's uncertainty into engineering predictability holds the pricing power of the next competitive cycle.
Statement of Information Gaps
- Devin's ARR and valuation data (USD 37 million to USD 492 million; USD 10.2 billion to USD 26 billion) come from third-party aggregation (low confidence); all are marked
[To be verified]and were not used in any industry-size estimation. - The authoritative source for the disbanding of the 妙鸭相机 team (end of 2025-09) is, to date, only the statement of a third-party tool site; no tech-media coverage or business-registration corroboration has been located,
[To be verified]. - The AI comic-drama market size (RMB 40 billion / USD 4 billion overseas) and hit-rate data come mainly from the DataEye report via secondary relay (百度百科, KOCPC, and others); the full original report was not obtained for cross-verification.
- Consumer pricing across platforms has multiple conflicting figures (即梦, 可灵 domestic site, 美图设计室, Coze enterprise edition, Vidu, and others); this whitepaper accepts only verifiable official-page figures, and the rest are marked
[To be verified]or not cited. - The release date of 可灵 3.0 and "up to 3 minutes", and 白日梦's "95%+ cross-episode similarity", etc., appear only in secondary sources,
[To be verified]. - The 阅文妙笔 business indicators (daily actives +100%, token +90%, weekly usage over 75%) and the 美图设计室 Agent Teams results are both vendor self-reported figures.
- Most platforms in the AI novel and image tracks have no officially disclosed verifiable effectiveness indicators; third-party review data are mostly marketing pieces and were not cited.
- This whitepaper does not adopt any self-media data on big domestic companies' AI coding efficiency as hard data (no primary source for this kind of data was found in this project's search).
10. References
- 2025 Stack Overflow Developer Survey — Stack Overflow, 2025-07-30. https://survey.stackoverflow.co/2025/
- State of AI-assisted Software Development 2025 — DORA / Google Cloud, 2025-11-12. https://dora.dev/research/2025/dora-report/
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (randomized controlled trial) — METR, 2025-07-10. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks — GitHub Blog, 2026. https://github.blog/ai-and-ml/github-copilot/evaluating-performance-and-efficiency-of-the-github-copilot-agentic-harness-across-models-and-tasks/
- Terminal-Bench official site and Leaderboard — Stanford / Laude Institute, from 2025 to 2026. https://www.tbench.ai/
- AGENTS.md official site — Agentic AI Foundation (under the Linux Foundation). https://agents.md/
- Sandboxing: a safer and more autonomous approach — Anthropic, 2025. https://www.anthropic.com/engineering/claude-code-sandboxing
- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities — MiniMax official blog. https://www.minimax.io/blog/minimax-h3
- 美图公司 2026 interim results announcement — 香港交易所披露易 (hkexnews). https://www.hkexnews.hk/
- 国家网信办等四部门《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — issued 2025-03-07, in effect 2025-09-01. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- Mandatory national standard GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》 — 全国网络安全标准化技术委员会. https://www.tc260.org.cn/upload/2025-03-15/1742009439794081593.pdf
- DataEye《2026 上半年 AI 剧 / 漫剧数据报告》 — DataEye研究院 (reprinted via secondary sources such as KOCPC). https://en.kocpc.com.tw/archives/25203
- 澎湃新闻《1914 元制作 1 集?漫剧仍困在隐性成本中》 — 澎湃新闻, 2026. https://www.thepaper.cn/newsDetail_forward_33993493
- 中国青年报《"妙笔通鉴""漫剧助手"发布,AI 赋能网文创作和 IP 改编》 — 中国青年报, 2025. https://new.qq.com/rain/a/20251017A08J6J00
- GitHub — forsonny/Claude-Code-Novel-Writer (Multi-Agent Novel Writer v4.1) — 2026. https://github.com/forsonny/Claude-Code-Novel-Writer
- Apache Software Foundation《AI Agents and AGENTS.md Files (Draft)》 — ASF legal-discuss, 2026. https://www.apache.org/legal/generative-tooling-agents.html
- 通义万相 and Qwen-family models: 阿里云 official pricing page — 阿里云. https://help.aliyun.com/
- Leonardo.ai official pricing and API documentation — Leonardo. https://leonardo.ai/pricing
- Runway official pricing page — Runway. https://runwayml.com/pricing
- Comfy official site — Comfy Org. https://comfy.org/