引言


1. 为什么需要这份白皮书

1.1. 一份面向发布的综述,而非又一份内部调研

2025 年至 2026 年,围绕大模型与智能体的讨论出现了一个值得注意的错位:厂商在谈论模型有多强,工程团队在谈论为什么落地仍然这么难。双方各自有大量材料,但彼此之间缺少一份能够把两侧证据放在同一张桌子上、用同一套术语讨论的综述文件。

本白皮书就是为填补这个空缺而写。它有三个自我约束:

  1. 基于既有调研,不做新的检索。全部素材来自本工程此前完成的 176 篇调研文档与检索报告(覆盖概述、八大行业组、七个市场组),撰写时不引入未经这些材料支撑的新数据。
  • 有观点,但必须给依据。本白皮书允许给出判断,但每一条判断都标注其事实依据与来源等级;依据不足之处明确标注,不用修辞掩盖证据缺口。
  • 区分事实与判断。凡属“已被公开来源确认的事实”,给出来源;凡属“本工程的分析性结论”,明确说明其为判断,供读者自行采信或反驳。
  • 因此,这份白皮书不是厂商宣传材料,也不预设任何产品立场。它更像一份证据清单加一套分析框架:读者可以不同意框架,但应当能够核验事实。

    1.2. 本白皮书试图回答的五个问题

    序号问题对应篇章
    1AI Harness 是什么?为什么它需要独立成一个概念、独立成一层工程能力?02-定义
    2当代 Harness 的六层能力模型是什么样?各层的关键组件与失败模式是什么?02-定义、03-架构
    3同一套六层模型,在不同行业的侧重为什么不同?差异背后的结构性原因是什么?04-实践
    4“Harness 已成为评测变量”意味着什么?企业应当如何建立自己的评测能力?05-基准
    5监管与标准正在如何回应?在中国境内运营需要守住的合规基线是什么?07-治理与风险

    五个问题依次递进:先定义对象,再看结构,再看行业差异,再看评测与监管这两个“外部检验”。如果读者时间有限,建议先读 02-定义的第 3 章(为什么需要独立成层),它是整份白皮书论证链的枢纽。

    1.3. 三类读者的预期收获

    读者角色读完本白皮书应能获得什么
    决策者一套判断 AI 投入是否有效的证据框架;理解为什么“再等下一代模型”不是对信任缺口的有效回应
    架构师六层能力模型的完整定义与边界判据;对自建 Harness 能力进行自查与补强的依据
    工程师各层的典型失败模式清单与工程对策;把每一次智能体失败固化为机制的方法论

    2. AI 从“能对话”到“能交付”的转折

    2.1. 转折已经发生的三重标志

    本工程在概述模块(详见 01-概述 / 02-发展历史)确认了三重标志,说明 AI 已经从“能对话”进入“能交付”的阶段:

    标志一:能力侧越过可用阈值。 SWE-bench Verified 上,Claude Opus 4.6 于 2026-02-05 取得 80.8%,成为该榜首个突破 80% 的模型;Terminal-Bench 2.0 上,Claude Sonnet 4.5 于 2025-09-29 取得 51.0%,成为该榜首个突破 50% 的模型。榜单回答的问题已经从“能不能做”变成“在什么约束下做、做错了谁负责”。

    标志二:采用侧越过普及阈值。 DORA 2025 报告显示 90% 以上的开发者已在工作中使用 AI;Stack Overflow 2025 开发者调研显示 84% 的受访者正在使用或计划使用 AI 工具。采用已经饱和,不再是变量。

    标志三:认知侧出现系统性裂缝。 头部实验室在 2026 年初几乎同时承认:制约系统表现的瓶颈已经不在模型内部。OpenAI 的表述是“我们最困难的挑战现在集中在设计环境、反馈回路与控制系统”;Anthropic 的表述是“Harness 设计是前沿智能体编码性能的关键”;Google DeepMind 的 Philipp Schmid 则给出了最简版本——“大多数智能体失败已不再是模型失败,而是上下文失败”。

    三重标志共同指向一个结论:讨论的焦点应当从“哪个模型更强”转向“模型被放置在什么工程环境中”。这个工程环境,就是本白皮书的主题——AI Harness。

    2.2. 三代演进:工程关注点持续外移

    图 2-1|AI 工程三代演进:瓶颈持续外移

    AI 工程三代演进:瓶颈持续外移 依据:124 篇调研文档与检索报告 · 示意:基于本文分析绘制 层层叠加 · 瓶颈外移 层层叠加 · 瓶颈外移 第一代 · 提示词工程时代(Prompt-Centric) 瓶颈位置:模型内部 约 2020—2023 · 核心工程对象:Prompt 本代瓶颈:模型不会用工具 第二代 · 工具与编排时代(Tool & Orchestration-Centric) 瓶颈位置:模型外 · 工具集成 约 2023—2025 · 核心工程对象:Tool + Orchestration 本代瓶颈:工具太多接不过来 第三代 · 运行时与评估时代(本图重点) 瓶颈位置:上下文与治理 2025—至今 · 核心工程对象:Context + Sandbox + Eval + Governance 本代瓶颈:上下文与治理决定成败 2026 年初被独立命名为 Harness:实践已运行一年半以上 每一代都是层层叠加而非替代:提示词工程仍然必要,工具调用被标准化为 MCP。 瓶颈持续外移 结构解读:三代是层层叠加而非替代,瓶颈从模型内部持续外移到工程环境。 焦点:采用率与信任度之间的缺口,正是 AI Harness 的市场空间与工程使命。

    数据来源:基于本文分析绘制的示意图。

    本工程统一采用三代划分(详见 01-概述 / 02-发展历史第 7 章):

    代际名称时间区间核心工程对象
    第一代提示词工程时代(Prompt-Centric)约 2020—2023Prompt
    第二代工具与编排时代(Tool & Orchestration-Centric)约 2023—2025Tool + Orchestration
    第三代运行时与评估时代(Runtime & Evaluation-Centric)2025—至今Context + Sandbox + Eval + Governance

    需要强调的是,三代之间是层层叠加而非相互替代:提示词工程今天仍然必要,只是不再充分;工具调用没有被淘汰,而是被标准化为 MCP。每一代解决的是上一代暴露出来的模型外部问题。这条“瓶颈持续外移”的演进规律,是理解为什么 Harness 在 2026 年被独立命名的钥匙。

    本工程在发展历史模块(详见 01-概述 / 02-发展历史第 8 章)还提炼出四条演进规律,它们构成本白皮书后续判断的方法论基础:

    1. 瓶颈持续外移:从“模型不会用工具”(第一代)到“工具太多接不过来”(第二代)到“上下文与治理决定成败”(第三代),每一代解决的都是上一代暴露的模型外部问题。
    2. 每一代都把上一代的核心手段沉淀为基础设施:提示词工程沉淀为 L1 的一部分,工具调用标准化为 MCP——新的一代不是替代,而是把前一代的手工活变成默认能力。
    3. 标准化总在事实标准出现之后:MCP 2024 年 11 月发布、2025 年 12 月才捐入 Linux Foundation;AGENTS.md 先被 60,000+ 项目采用,才成为捐赠标的。先有广泛采用的事实标准,后有中立治理的标准。
    4. 命名滞后于实践约 12—18 个月:上下文工程的实践始于 2024 年 8 月,2025 年 6 月才被命名;Harness 的产品实践始于 2025 年初,2026 年 2 月才被命名。当某个概念被正式命名时,其实践通常已经跑了至少一年。

    第四条规律对本白皮书的时间点判断尤为重要:Harness 的命名出现在 2026 年初,但这不意味着它是“刚刚出现的新事物”——恰恰相反,它意味着相关实践已经积累了一年半以上、到了必须被独立命名与独立治理的程度。

    2.3. 百万行代码实验:“能交付”的实证含义

    “能交付”不是修辞。OpenAI 在《Harness engineering》(2026-02-11)中披露的百万行代码实验给出了目前最完整的实证(A 级来源):

    指标数值
    代码总量约 100 万行(应用逻辑、基础设施、工具、文档、内部开发工具)
    PR 数 / 工程师数约 1,500 个 PR / 3 名工程师(后扩至 7 名)
    人均吞吐3.5 PR/天
    时间成本约为手写的 1/10
    人类手写代码0 行
    单任务最长运行时长超过 6 小时

    同样值得记录的是实验的起点:仓库的初始脚手架(包括最初的 AGENTS.md)也是由智能体自己生成的。这说明被交付的不仅是代码,还包括支撑代码持续生产的工程环境本身——后者正是 Harness。

    2.4. Harness 在此时被独立命名的三个必要条件

    一个自然的问题是:为什么 Harness 直到 2026 年初才被正式命名?本工程在概述模块(详见 01-概述 / 01-介绍第 7 章)给出了三个必要条件,它们在 2025 年底至 2026 年初同时满足:

    必要条件内容关键证据
    模型足够强强到失败不再主要源于模型智能不足,而源于环境设计不当两家头部实验室的“瓶颈外移”判断(A 级)
    工具有了标准MCP 让工具生态从 N×M 集成问题变为 N+MMCP 2024-11-25 发布,2025-12-09 捐入 AAIF(A 级)
    治理有了范本沙箱、权限、审计的实践被头部厂商跑通并量化Anthropic 沙箱:权限提示减少 84%(A 级)

    三个条件缺一则命名不成立:模型不够强时,瓶颈在模型内部,工程环境的重要性显示不出来;工具没有标准时,Harness 的建设成本高到无法成为通用实践;治理没有量化收益时,“约束即能力”只是口号。因此,“Harness 在 2026 年被命名”这个事实本身,就是三项工程条件同时成熟的证据。


    3. 行业真问题:三个未同步提升的工程能力

    3.1. 证据总览

    模型能力快速提升的同时,有三项工程能力没有同步提升。本白皮书将其概括为:落地成功率存疑、可复现性不足、可审计性缺位。支撑证据汇总如下:

    证据数据来源等级
    METR 随机对照试验(2025-07-10)资深开源开发者使用 AI 工具后实测慢 19%,同批开发者自评快 20%A
    DORA 2025 报告开发者自评生产力提升 80%;AI 采用与软件交付吞吐正相关,但与交付稳定性负相关A / B(负相关结论 )
    Stack Overflow 2025 开发者调研46% 的受访者不信任 AI 输出准确性;高度信任者仅 3.1%A
    Stack Overflow 2025 开发者调研66% 的受访者把“AI 方案几乎对但不完全对”列为最大挫败;45.2% 表示调试 AI 生成的代码比自己写更耗时A
    Terminal-Bench 方法论声明“每一个结果都是模型加智能体框架的组合……排行榜排的是系统,不是模型”A(项目)
    评测口径漂移同一模型家族在 SWE-bench Verified 可达 80.8%,在 SWE-bench Pro 上为 55.53%A / B
    IBM 泄露成本报告 202513% 的组织报告 AI 模型或应用遭泄露;被攻陷组织中 97% 未部署 AI 访问控制A
    中国监管动作GB/Z 185—2026 七部分闭环;《标识办法》2025-09-01 施行A

    三个问题分述如下。

    3.2. 落地成功率:体感与实测的背离

    METR 的随机对照试验(2025-07-10,A 级)是这一背离最直接的证据:资深开源开发者在真实开源项目中使用 AI 工具后,完成任务的实测时间反而延长了 19%,而同一批开发者的自评是“快了 20%”。两个数字方向相反,相差近 40 个百分点。

    本工程的分析(判断):背离的根源在于自评度量的是“写代码这一段”的体感速度,而实测度量的是“从接任务到合入主干”的端到端吞吐。AI 压缩了前者,却放大了后者中的验证成本、审查成本与处理“看起来对但实际错”的伪产物的成本。在缺乏客观度量的团队里,这类背离会长期被体感叙事掩盖。

    本工程给出的对策(判断,详见 02-行业赋能 / 03-软件工程组第 4 章):

    1. 统一度量到端到端:以部署频率、变更前置时间、变更失败率、服务恢复时间四项 DORA 指标为锚,禁止以“代码行数”“补全采纳率”作为效果结论。
    2. 把验证成本显性化:建立“验证时间 / 生成时间”比值指标,比值持续大于 1 说明验证环节已成为瓶颈。
    3. 以随机对照而非前后对比定结论:前后对比会被季节、需求难度、人员变动污染;随机对照是唯一能给出因果结论的设计。
    4. 对自评数据标注来源:引用自评数字时必须标注“自评”属性,不得与实测数字混用——这正是本节两个数字能够并列呈现的前提。

    3.3. 可复现性:同样的模型,不同的分数

    Terminal-Bench(Stanford + Laude Institute)在其方法论中明确声明:排行榜上排的是系统而非模型——每个成绩都是“模型 + 智能体框架(agent harness)”的组合成绩。这意味着Harness 已经成为评测中的显式变量

    公开数据中可以找到多个同模型不同分的对照案例。以下数字来自 B/C 级来源,全部保留 标注:

    对照项数据
    同为 GPT-5.3-Codex,不同智能体框架77.3% 对 75.1%,差 2.2 个百分点,纯由框架差异造成
    LangChain 仅改 Harness(同模型、同 API)52.8% 提升至 66.5%,排名从 30 名外升至前 5
    Vercel 将工具数量从 15 个削减到 2 个准确率 80% 升至 100%,Token 消耗下降 37%,速度提升 3.5 倍

    即便这些具体数字需要二次核实,其方向已被 A 级来源确认:Harness 的设计差异足以造成数个百分点甚至十几百分点的表现差异,这一量级常常超过换一个模型带来的差异。对企业而言,这同时是风险(榜单数字不可直接采信)与机会(不换模型也能通过工程手段提升表现)。

    此外,评测口径本身也存在漂移问题:任务分布不同的榜单之间不可比(Verified 的 80.8% 与 Pro 的 55.53% 反映的是任务难度差异,不是模型能力波动);推理力度、工具可用性、步数上限等配置差异对结果的影响,可能大于模型本身差异——例如 GPT-5.2 在 ARC-AGI-2 上 52.5% 的成绩明确标注为 xhigh 推理配置,直接用该数字与默认配置下的其他模型比较即为口径错误(B 级)。

    由此,本工程在软件工程组文档中确立了“引用数字必须带四要素”的强制约定:模型名与版本、榜单名、评测日期、推理与工具配置。四者缺任何一项,数字都不具备可比性。可比性层级为:同一榜同配置可比;同一榜不同配置可疑;跨榜不可比。详见 05-基准。

    3.4. 可审计性:采用已饱和,信任未跟上

    Stack Overflow 2025 调研刻画了一个清晰的矛盾:采用率 84% 已接近饱和,不信任输出准确性的比例却从 31%(2024)升至 46%(2025),高度信任者仅 3.1%。同一份调研中,对 AI 持正面情绪的开发者比例也从 70% 以上回落至 60%。信任没有跟上采用,缺的不是更好的模型,而是让模型行为可被约束、可被复现、可被举证的工程层

    监管侧的反应印证了这一点。欧盟《人工智能法案》(Regulation (EU) 2024/1689)对高风险系统提出日志留存、技术文件与人工监督义务;中国《人工智能生成合成内容标识办法》自 2025 年 9 月 1 日施行,把标识义务下沉到元数据与导出脚本一级;GB/Z 185—2026《人工智能 智能体互联》系列国家指导性技术文件则从身份码、身份管理到外部工具调用建立了七部分闭环。合规要求正在从“原则”下沉为“工程细节”,而承接这些细节的位置,正是 Harness 的 L5 评估与观测层与 L6 治理与安全层。

    值得并列呈现的一组证据是:IBM《Cost of a Data Breach Report 2025》显示,13% 的组织报告其 AI 模型或应用遭到泄露,其中 97% 的被攻陷组织未部署 AI 访问控制,63% 没有 AI 治理政策或仍在制定中(A 级)。AI 系统自身的治理缺位,已经从“潜在风险”变成了“已发生的损失”。

    3.5. 信任缺口就是 Harness 的空间

    综合三节证据,本白皮书的核心问题意识可以压缩为一句话(判断):采用率与信任度之间的缺口,就是 AI Harness 的市场空间与工程使命。模型的概率性输出无法被消除,但可以被工程手段收窄到可接受的区间——这正是 Harness 的定义性职责,下一章将给出正式定义。


    4. 本白皮书的方法与边界

    4.1. 材料来源与检索模式

    信息截止:本白皮书及其依托的调研库,全库信息截止 2026-09-12;晚于该日期的事件一律不写入正文。当日增量(2026-09-12 快照)以「当日增量」小节形式标注并汇入各章,未获官方原文确认的条目一律以 [待核实] 标示,详见版本发布页(/spec/releases/v1.1/)4.1 节快照增量明细与各章增量小节。

    本白皮书为基于既有资料模式的产品。全部素材来自本工程已完成的调研文档与检索报告,包括:

    材料域覆盖内容对应目录
    概述模块AI Harness 定义、发展历史、架构演进、未来趋势、全局结论01-概述
    行业赋能模块八大行业组(AI Infra、具身智能、软件工程、硬件研发、知识协同、数据科学、创意产业、风险合规)共 45 个方向文档02-行业赋能
    市场研究模块AI IDE、AI Agents、AI 图像、AI 漫剧、AI 小说、AI Infra、具身智能七组平台剖析(共 93 个平台)03-市场研究

    撰写过程中不引入上述材料之外的新数据。如果原始材料将某个数字标注为 ,本白皮书继承同一标注,不将其升级为确定事实。

    4.2. 来源分级与标注约定

    本白皮书沿用全工程的来源分级:

    等级含义使用规则
    A厂商或机构官方一手来源可直接引用,须给出出处
    B权威二手来源需注明转述,指标建议二次核对
    C社区与自媒体解读仅作线索,具体数字一律标注

    正文中出现的三种标注符号含义: 表示该数字或事实需二次核实后方可作为决策依据;[待填写] 表示该位置应有数据但未获得;[存疑] 表示存在相互矛盾的说法,正文并列呈现。

    4.3. 本白皮书不做什么

    1. 不做产品排名。市场研究模块的结论用于说明格局,本白皮书不据此给任何产品排出名次。
    2. 不做投资建议。Harness 层自身的市场规模目前没有权威测算(属 [待填写]),现有生成式 AI 市场预测口径差异近一个数量级,本白皮书不采用此类数字作论据。
    3. 不做技术玄学。凡属于“尚无公开解法”的问题(如单智能体与多智能体孰优),如实记录其为开放问题,不给出推测性答案。
    4. 不代替任何标准文本。引用标准时仅作工程视角解读,合规判断以标准与法规原文为准。
    5. 不重复调研库细节。行业案例的完整数据、逐条来源与信息缺口,回溯至 02-行业赋能与 03-市场研究对应文档;本白皮书只保留支撑结论所需的部分。
    6. 不回避不利证据。凡是与白皮书立场相悖的证据(如 METR 实测减速、长上下文负收益、厂商自报数字可信度有限),均如实呈现,不选择性引用。

    4.4. 事实与判断的区分

    本白皮书在行文上遵守以下区分约定:

    • 事实:来自 A/B 级来源且给出出处的陈述,例如“Anthropic 沙箱化权限提示减少 84%(A 级)”。
    • 判断:本工程基于事实所作的分析性结论,例如“约束即能力”“瓶颈已经外移”。判断必须能够回溯到支撑它的事实;读者若不接受判断,可以核验其事实基础。
    • 类比:如 Test Harness、马具、计算机四层栈等,仅用于辅助理解,不作为论据本身。

    4.5. 本白皮书的三条主线判断

    方法与边界明确之后,可以预告全书的论证主线。以下三条均为本工程的判断,其完整论证分别在第 2、3 章与后续各篇展开:

    1. 瓶颈已经外移:制约智能体系统表现的瓶颈,已从模型内部转移到模型外部的工程环境——这不是修辞,而是两家头部实验室在 2026 年初的同一判断。
    2. Harness 已成评测变量:评测排的是系统而非模型,Harness 的设计差异足以改变成绩排序——这同时解释了榜单数字为何不可直接采信,以及为何企业必须建立内部评测能力。
    3. 信任缺口即工程使命:采用率与信任度之间的缺口是工程问题而非模型问题,回应它的不是下一代模型,而是这一层此前长期缺位的工程承载层。

    5. 术语与阅读约定

    1. 统一定义:全文统一采用 AI Harness 的统一定义(见 02-定义第 1 章),任何章节不再另行定义或改写。
    2. 六层能力模型:L1 上下文工程层、L2 工具与执行层、L3 编排与控制层、L4 记忆与状态层、L5 评估与观测层、L6 治理与安全层,全文统一使用,层序不得调整。
    3. 三代演进:提示词工程时代(约 2020—2023)→ 工具与编排时代(约 2023—2025)→ 运行时与评估时代(2025—至今)。
    4. 缩写:专有名词、产品名、标准编号保留英文原文;缩写首次出现时给出全称。全部术语的中文释义与出处见 09-附录-术语表。
    5. 交叉引用:白皮书内部引用写“详见 0X-篇名”;引用调研库时写“详见 01-概述 / 0X-文件名”或“详见 02-行业赋能 / 0X-组名”。
    6. 日期冲突:个别事件在不同来源存在日期冲突(如 GB/Z 185—2026 的发布日期存在三种口径),本白皮书采用“2026 年上半年”等宽口径表述,并保留冲突说明。
    7. 规范引用:引用标准与法规时使用全称加编号的格式,例如《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号)、GB/T 45654—2025;条文解读仅为工程视角,合规判断以原文为准。
    8. 计量与符号:数值范围用 “~” 连接(如 1,000~2,000 tokens);百分比数字与 “%” 之间不空格;中英文之间加空格,中文标点使用全角。
    9. 图表引用:正文引用表格一律写作“如下表所示”或“见上表”,表格与正文声明中的数字必须一致——审核环节将逐项核对。

    6. 总结

    本章回答了“为什么需要这份白皮书”。三点结论:

    1. 转折已经发生:能力侧、采用侧、认知侧的三重标志表明,AI 已经从“能对话”进入“能交付”,而制约交付质量的是模型外部的工程环境。
    2. 真问题有三个:落地成功率存疑(自评与实测背离)、可复现性不足(Harness 已成评测变量、口径漂移)、可审计性缺位(采用 84% 对不信任 46%)。这三者共同构成信任缺口。
    3. 本白皮书的立场:信任缺口不会因为模型变强而自动关闭,它是工程问题,需要工程手段回应——这套手段的集合就是 AI Harness。

    7. 信息缺口声明

    1. OpenAI 弃用 SWE-bench Verified 作为主要评测基准的说法:本工程语料中未检索到一手来源确认,仅在任务派发环节被提及。本白皮书不将其作为论据使用,如需引用须另行取证,标 。
    2. 3.3 节全部同模型不同分的对照数字(77.3% 对 75.1%、52.8% 对 66.5%、Vercel 80% 对 100%):继承原始材料的 B/C 级来源属性,标 。
    3. DORA 2025 “AI 采用与交付稳定性负相关”结论:来自二手转述,标 。
    4. Harness 层自身市场规模:无权威测算,属 [待填写],本白皮书不给出任何数字。
    5. GB/Z 185—2026 发布日期:存在 2026-05-22 / 2026-06-26 / 2026-07-09 三种口径,本白皮书采用“2026 年上半年”。
    6. “Agent Harness”术语的首创出处:未找到一手首创文献,本白皮书不主张任何首创者归属。

    8. 参考资料

    1. Harness engineering: leveraging Codex in an agent-first world — OpenAI,2026-02-11。https://openai.com/index/harness-engineering/
    2. Harness design for long-running application development — Anthropic,2026。https://www.anthropic.com/engineering/harness-design-long-running-apps
    3. Effective context engineering for AI agents — Anthropic,2025。https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
    4. Effective harnesses for long-running agents — Anthropic,2025。https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
    5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity(随机对照试验) — METR,2025-07-10。https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
    6. 2025 Stack Overflow Developer Survey — Stack Overflow,2025-07-29。https://survey.stackoverflow.co/2025/
    7. DORA 2025 State of AI-assisted Software Development — Google Cloud / DORA,2025-11-12。https://dora.dev/research/2025/dora-report/
    8. Terminal-Bench 官方站 — Stanford / Laude Institute。https://www.tbench.ai/
    9. SWE-bench 官方站 — Princeton NLP 等。https://www.swebench.com/
    10. Linux Foundation Announces the Formation of the AAIF — Linux Foundation / AAIF,2025-12-09。https://aaif.io/
    11. 《人工智能 智能体互联》系列国家标准解读 — 中国产业经济信息网,2026。https://cinic.org.cn/xw/zcdt/1643418.html
    12. 人工智能生成合成内容标识办法 — 国家互联网信息办公室等四部门,2025。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
    13. Harness engineering for coding agent users — Birgitta Böckeler,martinfowler.com,2026。https://martinfowler.com/articles/harness-engineering.html
    14. My AI Adoption Journey — Mitchell Hashimoto,2026-02-05。https://mitchellh.com/writing/my-ai-adoption-journey

    Introduction

    1. Why This Whitepaper Is Needed

    1.1. A Review Written for Publication, Not Another Internal Study

    Between 2025 and 2026, the discussion around large models and agents revealed a notable misalignment: vendors talk about how capable the models are, while engineering teams talk about why deployment is still so hard. Each side has produced a great deal of material, but what is missing between them is a review document that can place both sides' evidence on the same table and discuss it in the same set of terms.

    This whitepaper was written to fill that gap. It is bound by three self-imposed constraints:

    1. Based on existing research; no new retrieval. All material comes from the 176 research documents and retrieval reports this project has previously completed (covering the overview, eight industry groups, and seven market groups); no new data unsupported by these materials is introduced during writing.
    2. It has viewpoints, but every one must be backed by evidence. This whitepaper permits judgments, but every judgment is annotated with its factual basis and source grade; where evidence is lacking, that is explicitly marked rather than masked by rhetoric.
    3. It distinguishes facts from judgments. Anything that is "a fact confirmed by public sources" is given a source; anything that is "an analytical conclusion of this project" is clearly stated as a judgment, for readers to accept or refute as they see fit.

    Therefore, this whitepaper is not vendor promotional material, nor does it presuppose any product stance. It is more like an evidence list plus an analytical framework: readers may disagree with the framework, but they should be able to verify the facts.

    1.2. The Five Questions This Whitepaper Attempts to Answer

    No.QuestionCorresponding Section
    1What is an AI Harness? Why does it need to stand on its own as a concept and as a distinct layer of engineering capability?02-Definition
    2What does the contemporary six-layer capability model of a Harness look like? What are each layer's key components and failure modes?02-Definition, 03-Architecture
    3Why does the same six-layer model receive different emphasis across industries? What structural reasons lie behind the differences?04-Practice
    4What does "the Harness has become an evaluation variable" mean? How should enterprises build their own evaluation capability?05-Benchmark
    5How are regulation and standards responding? What compliance baseline must be maintained for operations within China?07-Governance & Risk

    The five questions progress in sequence: first define the object, then examine its structure, then examine industry differences, and finally examine the two "external checks" of evaluation and regulation. If readers are short on time, it is recommended to first read Chapter 3 of 02-Definition (why it needs to stand on its own as a layer); it is the hub of the entire whitepaper's argument chain.

    1.3. Expected Takeaways for Three Types of Readers

    Reader RoleWhat They Should Gain from Reading This Whitepaper
    Decision makersAn evidence framework for judging whether AI investment is effective; an understanding of why "wait for the next generation of models" is not an effective response to the trust gap
    ArchitectsThe complete definition and boundary criteria of the six-layer capability model; the basis for self-auditing and strengthening their own Harness capabilities
    EngineersA list of typical failure modes for each layer and engineering countermeasures; a methodology for turning every agent failure into a hardened mechanism

    2. The Turning Point from "Being Able to Chat" to "Being Able to Deliver"

    2.1. Three Signs That the Turning Point Has Already Happened

    This project has confirmed three signs in the overview module (see 01-Overview / 02-Development History) showing that AI has moved from the stage of "being able to chat" into the stage of "being able to deliver":

    Sign one: capability has crossed the usability threshold. On SWE-bench Verified, Claude Opus 4.6 achieved 80.8% on 2026-02-05, becoming the first model on that leaderboard to break 80%; on Terminal-Bench 2.0, Claude Sonnet 4.5 achieved 51.0% on 2025-09-29, becoming the first model on that leaderboard to break 50%. The question these leaderboards answer has shifted from "can it be done" to "under what constraints can it be done, and who is responsible when it goes wrong".

    Sign two: adoption has crossed the diffusion threshold. The DORA 2025 report shows that more than 90% of developers already use AI at work; the Stack Overflow 2025 Developer Survey shows that 84% of respondents are using or planning to use AI tools. Adoption has saturated; it is no longer a variable.

    Sign three: a systemic crack has appeared on the cognition side. At the beginning of 2026, leading labs almost simultaneously acknowledged that the bottleneck constraining system performance is no longer inside the model. OpenAI put it as "our hardest challenges are now concentrated in designing environments, feedback loops, and control systems"; Anthropic put it as "harness design is the key to frontier agent coding performance"; and Google DeepMind's Philipp Schmid gave the most concise version — "most agent failures are no longer model failures, but context failures".

    The three signs together point to one conclusion: the focus of discussion should shift from "which model is stronger" to "what engineering environment the model is placed in". That engineering environment is the subject of this whitepaper — the AI Harness.

    2.2. Three Generations of Evolution: Engineering Focus Keeps Moving Outward

    图 2-1|AI 工程三代演进:瓶颈持续外移

    AI 工程三代演进:瓶颈持续外移 依据:124 篇调研文档与检索报告 · 示意:基于本文分析绘制 层层叠加 · 瓶颈外移 层层叠加 · 瓶颈外移 第一代 · 提示词工程时代(Prompt-Centric) 瓶颈位置:模型内部 约 2020—2023 · 核心工程对象:Prompt 本代瓶颈:模型不会用工具 第二代 · 工具与编排时代(Tool & Orchestration-Centric) 瓶颈位置:模型外 · 工具集成 约 2023—2025 · 核心工程对象:Tool + Orchestration 本代瓶颈:工具太多接不过来 第三代 · 运行时与评估时代(本图重点) 瓶颈位置:上下文与治理 2025—至今 · 核心工程对象:Context + Sandbox + Eval + Governance 本代瓶颈:上下文与治理决定成败 2026 年初被独立命名为 Harness:实践已运行一年半以上 每一代都是层层叠加而非替代:提示词工程仍然必要,工具调用被标准化为 MCP。 瓶颈持续外移 结构解读:三代是层层叠加而非替代,瓶颈从模型内部持续外移到工程环境。 焦点:采用率与信任度之间的缺口,正是 AI Harness 的市场空间与工程使命。

    数据来源:基于本文分析绘制的示意图。

    This project uniformly adopts a three-generation division (see 01-Overview / Chapter 7 of 02-Development History):

    GenerationNameTime RangeCore Engineering Object
    First generationPrompt-Centric eraApprox. 2020—2023Prompt
    Second generationTool & Orchestration-Centric eraApprox. 2023—2025Tool + Orchestration
    Third generationRuntime & Evaluation-Centric era2025—presentContext + Sandbox + Eval + Governance

    It is important to emphasize that the three generations stack on top of each other rather than replace one another: prompt engineering is still necessary today, just no longer sufficient; tool calling has not been eliminated, but has been standardized into MCP. Each generation solves the model-external problems that the previous generation exposed. This evolutionary pattern of "the bottleneck continually moving outward" is the key to understanding why the Harness was given its own name in 2026.

    In the development-history module (see 01-Overview / Chapter 8 of 02-Development History), this project also distilled four evolutionary laws, which form the methodological basis for the whitepaper's subsequent judgments:

    1. The bottleneck keeps moving outward: from "the model does not know how to use tools" (first generation) to "too many tools to integrate" (second generation) to "context and governance determine success or failure" (third generation); each generation solves the model-external problems that the previous generation exposed.
    2. Each generation sediments the previous one's core means into infrastructure: prompt engineering was sedimented into part of L1, and tool calling was standardized into MCP — a new generation does not replace, but turns the previous generation's manual work into default capability.
    3. Standardization always follows de facto standards: MCP was released in November 2024 but not donated to the Linux Foundation until December 2025; AGENTS.md was first adopted by 60,000+ projects before becoming the donation target. Broad adoption as a de facto standard comes first, and a neutrally governed standard comes later.
    4. Naming lags practice by roughly 12—18 months: the practice of context engineering began in August 2024 but was not named until June 2025; the product practice of the Harness began in early 2025 but was not named until February 2026. By the time a concept is formally named, its practice has usually already been running for at least a year.

    The fourth law matters especially for this whitepaper's judgment about timing: the Harness was named in early 2026, but that does not mean it is a "something that has just appeared" — quite the opposite, it means the related practice has already accumulated for more than a year and a half, reaching the point where it must be independently named and independently governed.

    2.3. The Million-Lines-of-Code Experiment: The Empirical Meaning of "Being Able to Deliver"

    "Being able to deliver" is not rhetoric. The million-lines-of-code experiment disclosed by OpenAI in Harness engineering (2026-02-11) provides the fullest empirical evidence to date (a Grade A source):

    MetricValue
    Total codeApprox. 1 million lines (application logic, infrastructure, tools, documentation, internal development tooling)
    Number of PRs / engineersApprox. 1,500 PRs / 3 engineers (later expanded to 7)
    Per-person throughput3.5 PRs/day
    Time costApprox. 1/10 of hand-written
    Human-written code0 lines
    Longest single-task runtimeOver 6 hours

    Equally worth recording is the starting point of the experiment: the repository's initial scaffolding (including the original AGENTS.md) was also generated by the agent itself. This shows that what is delivered is not only the code, but also the engineering environment itself that supports the code being continuously produced — and that environment is exactly the Harness.

    2.4. Three Necessary Conditions for the Harness to Be Independently Named at This Time

    A natural question is: why was the Harness not formally named until early 2026? In the overview module (see 01-Overview / Chapter 7 of 01-Introduction), this project gives three necessary conditions, all of which were satisfied simultaneously between late 2025 and early 2026:

    Necessary ConditionContentKey Evidence
    Models are strong enoughStrong enough that failure no longer mainly stems from insufficient model intelligence, but from improper environment designThe "bottleneck has moved outward" judgment of two leading labs (Grade A)
    Tools have a standardMCP turns the tool ecosystem from an N×M integration problem into N+MMCP released 2024-11-25, donated to AAIF 2025-12-09 (Grade A)
    Governance has a model to followThe practices of sandboxing, permissions, and auditing have been proven out and quantified by leading vendorsAnthropic sandbox: permission prompts reduced by 84% (Grade A)

    If any one of the three conditions is missing, the naming does not hold: when models are not strong enough, the bottleneck is inside the model and the importance of the engineering environment cannot be shown; when tools have no standard, the cost of building a Harness is too high for it to become common practice; when governance has no quantified benefit, "constraints are capability" is just a slogan. Therefore, the very fact that "the Harness was named in 2026" is itself evidence that the three engineering conditions matured at the same time.


    3. Real Industry Problems: Three Engineering Capabilities That Have Not Improved in Step

    3.1. Overview of the Evidence

    While model capability has improved rapidly, three engineering capabilities have not improved in step. This whitepaper summarizes them as: deployment success rates are in doubt, reproducibility is insufficient, and auditability is lacking. The supporting evidence is summarized below:

    EvidenceDataSource Grade
    METR randomized controlled trial (2025-07-10)Experienced open-source developers measured 19% slower with AI tools; the same developers self-rated 20% fasterA
    DORA 2025 reportDevelopers self-rated an 80% productivity improvement; AI adoption is positively correlated with software delivery throughput, but negatively correlated with delivery stabilityA / B (the negative-correlation conclusion [to verify])
    Stack Overflow 2025 Developer Survey46% of respondents do not trust the accuracy of AI output; only 3.1% trust it highlyA
    Stack Overflow 2025 Developer Survey66% of respondents named "AI solutions that are almost right but not quite" as their greatest frustration; 45.2% said debugging AI-generated code takes more time than writing it themselvesA
    Terminal-Bench methodology statement"Every result is a combination of a model plus an agent harness… the leaderboard ranks systems, not models"A (project)
    Evaluation-caliber driftThe same model family can reach 80.8% on SWE-bench Verified but 55.53% on SWE-bench ProA / B
    IBM Cost of a Data Breach Report 202513% of organizations report a breach of an AI model or application; of the compromised organizations, 97% had not deployed AI access controlsA
    Chinese regulatory actionsGB/Z 185—2026 seven-part closed loop; the Labeling Measures took effect 2025-09-01A

    The three problems are described in turn below.

    3.2. Deployment Success Rate: The Divergence Between Perception and Measurement

    METR's randomized controlled trial (2025-07-10, Grade A) is the most direct evidence of this divergence: after experienced open-source developers used AI tools on real open-source projects, their measured time to complete tasks actually increased by 19%, while the same developers self-rated as "20% faster". The two numbers point in opposite directions, differing by nearly 40 percentage points.

    This project's analysis (judgment): the root of the divergence is that self-assessment measures the perceived speed of "the code-writing segment", while measurement captures the end-to-end throughput "from receiving the task to merging into the main branch". AI compresses the former but amplifies the latter's verification cost, review cost, and the cost of handling "looks right but is actually wrong" artifacts. In teams that lack objective measurement, this kind of divergence is long masked by perceptual narratives.

    The countermeasures this project provides (judgment; see 02-Industry Enablement / Chapter 4 of the 03-Software Engineering Group):

    1. Unify measurement to end-to-end: anchor on the four DORA metrics of deployment frequency, change lead time, change failure rate, and service restoration time, and forbid using "lines of code" or "completion adoption rate" as conclusions about effectiveness.
    2. Make verification cost explicit: establish a "verification time / generation time" ratio metric; a ratio persistently greater than 1 means the verification stage has become the bottleneck.
    3. Base conclusions on randomized controlled trials rather than before-after comparisons: before-after comparisons can be contaminated by seasonality, requirement difficulty, and staff turnover; a randomized controlled trial is the only design that yields causal conclusions.
    4. Annotate the source of self-reported data: when citing self-reported numbers, they must be marked as "self-reported" and never mixed with measured numbers — this is precisely the premise that allows this section's two numbers to be presented side by side.

    3.3. Reproducibility: The Same Model, Different Scores

    Terminal-Bench (Stanford + Laude Institute) explicitly states in its methodology that what the leaderboard ranks are systems, not models — every score is the combined result of "a model + an agent harness". This means the Harness has already become an explicit variable in evaluation.

    Multiple controlled cases of the same model with different scores can be found in public data. The numbers below come from Grade B/C sources, and all retain the [to verify] annotation:

    ComparisonData
    Same GPT-5.3-Codex, different agent harnesses77.3% vs. 75.1%, a 2.2-percentage-point difference caused purely by harness differences
    LangChain changing only the Harness (same model, same API)52.8% improved to 66.5%, ranking rising from outside the top 30 into the top 5
    Vercel cutting tools from 15 to 2Accuracy rose from 80% to 100%, token consumption fell 37%, and speed improved 3.5×

    Even if these specific numbers need a second verification, their direction has been confirmed by Grade A sources: differences in Harness design are enough to cause performance differences of several or even a dozen-plus percentage points, a magnitude that often exceeds the difference from switching models. For enterprises, this is simultaneously a risk (leaderboard numbers cannot be taken at face value) and an opportunity (performance can be improved through engineering means without switching models).

    In addition, evaluation caliber itself has a drift problem: leaderboards with different task distributions are not comparable (the 80.8% on Verified and the 55.53% on Pro reflect differences in task difficulty, not fluctuations in model capability); configuration differences such as reasoning effort, tool availability, and step limits can affect results more than model differences themselves — for example, GPT-5.2's 52.5% on ARC-AGI-2 is explicitly annotated as using the xhigh reasoning configuration, so directly comparing that number with other models under default configurations is a caliber error (Grade B, [to verify]).

    Accordingly, this project established a mandatory convention in the software-engineering-group documents that "citing a number must carry four elements": model name and version, leaderboard name, evaluation date, and reasoning and tool configuration. If any one of the four is missing, the number lacks comparability. The comparability hierarchy is: same leaderboard with same configuration is comparable; same leaderboard with different configuration is questionable; cross-leaderboard is not comparable. See 05-Benchmark.

    3.4. Auditability: Adoption Has Saturated, but Trust Has Not Kept Up

    The Stack Overflow 2025 Survey portrays a clear contradiction: adoption at 84% is close to saturation, yet the share that does not trust output accuracy has risen from 31% (2024) to 46% (2025), with only 3.1% trusting it highly. In the same survey, the share of developers holding positive sentiment toward AI also fell from above 70% to 60%. Trust has not kept up with adoption, and what is missing is not a better model, but an engineering layer that makes model behavior constrainable, reproducible, and provable.

    The regulatory response confirms this. The EU AI Act (Regulation (EU) 2024/1689) imposes obligations of log retention, technical documentation, and human oversight on high-risk systems; China's Measures for the Labeling of AI-Generated and Synthesized Content took effect on September 1, 2025, pushing labeling obligations down to the level of metadata and export scripts; and GB/Z 185—2026, the Artificial Intelligence — Agent Interconnection series of national guiding technical documents, establishes a seven-part closed loop from identity codes and identity management to external tool invocation. Compliance requirements are moving from "principles" down to "engineering details", and the place that shoulders these details is precisely the Harness's L5 evaluation & observation layer and L6 governance & security layer.

    A set of evidence worth presenting side by side is: IBM's Cost of a Data Breach Report 2025 shows that 13% of organizations report a breach of their AI models or applications, of which 97% of the compromised organizations had not deployed AI access controls, and 63% had no AI governance policy or were still developing one (Grade A). The governance gap of AI systems themselves has turned from a "potential risk" into "losses already incurred".

    3.5. The Trust Gap Is the Harness's Space

    Putting the evidence of the three sections together, this whitepaper's core problem consciousness can be compressed into one sentence (judgment): the gap between adoption rate and trust is the market space and engineering mission of the AI Harness. A model's probabilistic output cannot be eliminated, but it can be narrowed by engineering means to an acceptable range — this is precisely the Harness's defining responsibility, and the next chapter will give the formal definition.


    4. This Whitepaper's Method and Boundaries

    4.1. Material Sources and Retrieval Pattern

    Information cutoff: This whitepaper and the research library it relies on have a full-library information cutoff of 2026-09-12; events later than that date are uniformly excluded from the body text. The day's increments (2026-09-12 snapshot) are marked as "Same-Day Increment" subsections and merged into each chapter; entries not confirmed against the official original text are uniformly marked with [To be verified]. For details, see the snapshot increment details in Section 4.1 of the v1.1 release page (/spec/releases/v1.1/) and the increment subsections of each chapter.

    This whitepaper is a product of the based on existing materials mode. All material comes from the research documents and retrieval reports this project has already completed, including:

    Material DomainContent CoveredCorresponding Directory
    Overview moduleAI Harness definition, development history, architecture evolution, future trends, overall conclusions01-Overview
    Industry-enablement moduleDocuments across 45 directions within the eight industry groups (AI Infra, embodied intelligence, software engineering, hardware R&D, knowledge collaboration, data science, creative industries, risk & compliance)02-Industry Enablement
    Market-research modulePlatform analyses of seven groups: AI IDEs, AI Agents, AI imaging, AI comic dramas, AI novels, AI Infra, embodied intelligence (93 platforms in total)03-Market Research

    During writing, no new data beyond the above materials is introduced. If an original material marked a certain number as [to verify], this whitepaper inherits the same annotation and does not upgrade it to an established fact.

    4.2. Source Grading and Annotation Conventions

    This whitepaper follows the source grading used across the whole project:

    GradeMeaningUsage Rule
    AOfficial first-hand source from a vendor or institutionCan be cited directly; the source must be given
    BAuthoritative second-hand sourceMust be noted as a paraphrase; metrics advisable to double-check
    CCommunity and self-media interpretationUsed only as a lead; any specific number must be annotated [to verify]

    The meanings of the three annotation symbols that appear in the body text: [to verify] means the number or fact needs a second verification before it can serve as a basis for decisions; [to fill in] means data should exist in that position but has not been obtained; [in doubt] means there are contradictory accounts, and the body text presents them side by side.

    4.3. What This Whitepaper Does Not Do

    1. It does not rank products. The conclusions of the market-research module are used to illustrate the landscape; this whitepaper does not rank any product on that basis.
    2. It does not give investment advice. There is currently no authoritative estimate of the market size of the Harness layer itself (it falls under [to fill in]); existing generative-AI market forecasts differ by nearly an order of magnitude in caliber, so this whitepaper does not use such numbers as evidence.
    3. It does not engage in technical mysticism. For any question that has "no publicly available solution" (such as whether single agents or multi-agents are superior), it honestly records the question as open and gives no speculative answer.
    4. It does not replace any standard text. When standards are cited, they are interpreted only from an engineering perspective; compliance judgments follow the original text of the standards and regulations.
    5. It does not repeat the research-library details. The full data, item-by-item sources, and information gaps of industry cases trace back to the corresponding documents in 02-Industry Enablement and 03-Market Research; this whitepaper retains only what is needed to support its conclusions.
    6. It does not avoid adverse evidence. Any evidence that runs counter to the whitepaper's position (such as METR's measured slowdown, the negative returns of long context, and the limited credibility of vendor self-reported numbers) is all presented honestly, not selectively cited.

    4.4. The Distinction Between Facts and Judgments

    In its wording, this whitepaper observes the following distinction conventions:

    • Facts: statements from Grade A/B sources with the source given, such as "Anthropic's sandboxed permission prompts reduced them by 84% (Grade A)".
    • Judgments: analytical conclusions this project has drawn based on facts, such as "constraints are capability" and "the bottleneck has already moved outward". Judgments must be traceable back to the facts that support them; readers who do not accept a judgment can verify its factual basis.
    • Analogies: such as the Test Harness, horse tack, and the four-layer computer stack — used only to aid understanding, not as arguments in themselves.
    • 4.5. The Whitepaper's Three Main-Line Judgments

      With the method and boundaries clarified, the book's overall argument thread can be previewed. The three following claims are all judgments of this project, and their full arguments are developed respectively in Chapters 2, 3, and the subsequent sections:

      1. The bottleneck has already moved outward: the bottleneck constraining agent-system performance has shifted from inside the model to the engineering environment outside the model — this is not rhetoric, but the shared judgment of two leading labs in early 2026.
      2. The Harness has already become an evaluation variable: evaluation ranks systems rather than models, and differences in Harness design are enough to change the ranking of scores — this simultaneously explains why leaderboard numbers cannot be taken at face value, and why enterprises must build their own internal evaluation capability.
      3. The trust gap is an engineering mission: the gap between adoption rate and trust is an engineering problem rather than a model problem; what responds to it is not the next generation of models, but the engineering-bearing layer that had long been missing at this level.

      5. Terminology and Reading Conventions

      1. Unified definitions: the whole text uniformly adopts the unified definition of the AI Harness (see Chapter 1 of 02-Definition); no chapter redefines or rewrites it.
      2. The six-layer capability model: L1 context-engineering layer, L2 tool and execution layer, L3 orchestration and control layer, L4 memory and state layer, L5 evaluation and observation layer, and L6 governance and security layer — used uniformly throughout the text, and the layer order must not be adjusted.
      3. Three-generation evolution: the Prompt-Centric era (approx. 2020—2023) → the Tool & Orchestration-Centric era (approx. 2023—2025) → the Runtime & Evaluation-Centric era (2025—present).
      4. Abbreviations: proper nouns, product names, and standard numbers retain their original English; abbreviations give the full name at first appearance. The Chinese glosses and sources of all terms are in 09-Appendix-Glossary.
      5. Cross-references: references within the whitepaper are written as "see 0X-Chapter Title"; references to the research library are written as "see 01-Overview / 0X-File Name" or "see 02-Industry Enablement / 0X-Group Name".
      6. Date conflicts: a few events have conflicting dates across sources (for instance, the release date of GB/Z 185—2026 has three calibers); this whitepaper uses broad-caliber wording such as "the first half of 2026", and preserves the conflict note.
      7. Citing regulations: standards and regulations are cited with the full name plus the number, for example the Measures for the Labeling of AI-Generated and Synthesized Content (国信办通字〔2025〕2 号) and GB/T 45654—2025; interpretations of provisions are only from an engineering perspective, and compliance judgments follow the original text.
      8. Units and symbols: numeric ranges are joined with "~" (e.g., 1,000~2,000 tokens); percentage numbers have no space before "%"; spaces are added between Chinese and English, and Chinese punctuation uses full-width forms.
      9. Referencing tables and figures: references to tables in the body text are always written as "as shown in the table below" or "see the table above"; the numbers in tables and in body-text statements must agree — the review stage will verify them item by item.

      6. Summary

      This chapter answers "why this whitepaper is needed". Three conclusions:

      1. The turning point has already happened: the three signs on the capability, adoption, and cognition sides show that AI has moved from "being able to chat" into "being able to deliver", and what constrains delivery quality is the engineering environment outside the model.
      2. There are three real problems: deployment success rate is in doubt (self-assessment diverges from measurement), reproducibility is insufficient (the Harness has become an evaluation variable, and caliber drifts), and auditability is lacking (84% adoption vs. 46% distrust). Together these three constitute the trust gap.
      3. This whitepaper's position: the trust gap will not close on its own just because models become stronger; it is an engineering problem that needs engineering means to respond — and the set of those means is the AI Harness.

      7. Information-Gap Statement

      1. The claim that OpenAI has dropped SWE-bench Verified as its primary evaluation benchmark: no first-hand source confirming this was found in this project's corpus; it was only mentioned during task assignment. This whitepaper does not use it as evidence; if it is to be cited, it must be separately verified, and it is annotated [to verify].
      2. All of Section 3.3's comparison numbers of the same model with different scores (77.3% vs. 75.1%, 52.8% vs. 66.5%, Vercel 80% vs. 100%): they inherit the Grade B/C source attributes of the original materials, annotated [to verify].
      3. The DORA 2025 conclusion that "AI adoption is negatively correlated with delivery stability": it comes from a second-hand paraphrase, annotated [to verify].
      4. The market size of the Harness layer itself: there is no authoritative estimate, so it falls under [to fill in], and this whitepaper gives no number.
      5. The release date of GB/Z 185—2026: there are three calibers — 2026-05-22 / 2026-06-26 / 2026-07-09 — and this whitepaper uses "the first half of 2026".
      6. The original source that coined the term "Agent Harness": no first-hand origin literature has been found; this whitepaper claims no attribution to any originator.

      8. References

      1. Harness engineering: leveraging Codex in an agent-first world — OpenAI, 2026-02-11. https://openai.com/index/harness-engineering/
      2. Harness design for long-running application development — Anthropic, 2026. https://www.anthropic.com/engineering/harness-design-long-running-apps
      3. Effective context engineering for AI agents — Anthropic, 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
      4. Effective harnesses for long-running agents — Anthropic, 2025. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
      5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (randomized controlled trial) — METR, 2025-07-10. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
      6. 2025 Stack Overflow Developer Survey — Stack Overflow, 2025-07-29. https://survey.stackoverflow.co/2025/
      7. DORA 2025 State of AI-assisted Software Development — Google Cloud / DORA, 2025-11-12. https://dora.dev/research/2025/dora-report/
      8. Terminal-Bench official site — Stanford / Laude Institute. https://www.tbench.ai/
      9. SWE-bench official site — Princeton NLP et al. https://www.swebench.com/
      10. Linux Foundation Announces the Formation of the AAIF — Linux Foundation / AAIF, 2025-12-09. https://aaif.io/
      11. Artificial Intelligence — Agent Interconnection: interpretation of the series of national standards — China Industrial Economic Information Network, 2026. https://cinic.org.cn/xw/zcdt/1643418.html
      12. Measures for the Labeling of AI-Generated and Synthesized Content — Cyberspace Administration of China and three other departments, 2025. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
      13. Harness engineering for coding agent users — Birgitta Böckeler, martinfowler.com, 2026. https://martinfowler.com/articles/harness-engineering.html
      14. My AI Adoption Journey — Mitchell Hashimoto, 2026-02-05. https://mitchellh.com/writing/my-ai-adoption-journey