参考资料


1. 使用说明

图 1-1|参考资料来源三级可信度分级体系(A / B / C)

参考资料来源三级可信度分级体系(A / B / C) 共收录 74 条(R1–R74)· 示意:基于本文分析绘制 A 级 · 官方一手来源(51 条) 使用规则:可直接引用,须给出出处 官方工程博客 / 新闻公告 Anthropic / OpenAI / Google / Linux Foundation arXiv 原文 · 官方开源仓库 · 标准解读 MCP 规范 · GB/Z 185 国标 · Stack Overflow 报告 可信度递减 B 级 · 权威二手来源(19 条) 使用规则:需注明转述;指标建议二次核对 百科与二手综述 Wikipedia · Wikiwand · AI Wiki · Klu 社区平台与第三方时间线 MLOps Community · daily.dev · Pure AI · Taskade · GitHub 仓库 需二次核实 C 级 · 社区与自媒体解读(4 条) 使用规则:仅作线索;具体数字一律标 后方可入正文 数据解读博客 AgentMarketCap 等自媒体 云社区开发者文章 腾讯云 / 阿里云开发者社区;框架论述须与 A 级交叉印证 结构解读:来源按可信度分为三级——A 级可直接引用、B 级注明转述、C 级仅作线索, 下游文档引用时须按等级选择引用方式,低等级数字一律标注 。

数据来源:基于本文分析绘制的示意图。

1.1. 来源分级标准

本文件收录的全部资料按三级可信度分级,供下游文档引用时判断引用方式:

等级含义使用规则本文件收录数量
A厂商或机构官方一手来源:Anthropic、OpenAI、Google、Linux Foundation / AAIF、Stack Overflow、人民网、中国日报、中国产业经济信息网(标准归口单位解读)、arXiv 原文、官方开源仓库等可直接引用,须给出出处51 条
B权威二手来源:Wikipedia、Wikiwand、AI Wiki、MLOps Community、daily.dev、Pure AI、Taskade、百度百科、第三方时间线仓库等需注明转述;指标建议二次核对19 条
C社区与自媒体解读:AgentMarketCap、腾讯云开发者社区、阿里云开发者社区等仅作线索;具体数字一律标 后方可入正文4 条

合计 74 条(R1 ~ R74)。

1.2. 编号规则

  • 编号形如 R1 ~ R55,跨节连续;
  • 同一条目在多个分类中出现的,以主分类为准,不重复编号;
  • 标注 者为链接或日期未经一手验证,引用时需二次确认;
  • 标注 [存疑] 者存在来源间冲突,已在 02-发展历史.md 第 10 节并列呈现。

2. 官方工程博客

2.1. Anthropic

编号名称年份 / 日期等级链接
R1Effective context engineering for AI agents(Applied AI 团队)2025Ahttps://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
R2Effective harnesses for long-running agents2025Ahttps://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
R3Harness design for long-running application development2026Ahttps://www.anthropic.com/engineering/harness-design-long-running-apps
R4Equipping agents for the real world with Agent Skills2025-10-16Ahttps://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
R5Sandboxing: a safer and more autonomous approach2025Ahttps://www.anthropic.com/engineering/claude-code-sandboxing
R6Introducing Agent Skills(2025-12-18 转为开放标准)2025-10-16Ahttps://www.anthropic.com/news/skills
R7Introducing the Model Context Protocol2024-11-25Ahttps://www.anthropic.com/news/model-context-protocol
R8Claude 3.7 Sonnet and Claude Code2025-02-24Ahttps://www.anthropic.com/news/claude-3-7-sonnet
R9Donating the Model Context Protocol and establishing the AAIF2025-12Ahttps://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
R10How we built our multi-agent research system2025Ahttps://www.anthropic.com/engineering/multi-agent-research-system

2.2. OpenAI

编号名称年份 / 日期等级链接
R11Harness engineering: leveraging Codex in an agent-first world(Ryan Lopopolo)2026-02-11Ahttps://openai.com/index/harness-engineering/ ;中文版 https://openai.com/zh-Hans-CN/index/harness-engineering/
R12Function calling and other API updates2023-06-13Ahttps://openai.com/blog/function-calling-and-other-API-updates
R13New tools for building agents(Responses API + Agents SDK)2025-03-11Ahttps://openai.com/blog/new-tools-for-building-agents
R14Introducing Codex(Codex CLI)2025-04-16Ahttps://github.com/openai/codex
R15OpenAI co-founds the Agentic AI Foundation under the Linux Foundation2025-12Ahttps://openai.com/index/agentic-ai-foundation/

2.3. Google

编号名称年份 / 日期等级链接
R16Agent Development Kit: Making it easy to build multi-agent applications2025-04-09Ahttps://googledevelopers.blogspot.com/en/agent-development-kit-easy-to-build-multi-agent-applications/
R17A year of open collaboration: Celebrating the anniversary of A2A2026-04-16Ahttps://opensource.googleblog.com/

2.4. Linux Foundation 与 AAIF


3. 标准与规范

3.1. 协议规范

编号名称机构年份等级链接
R21Model Context Protocol 官方站与规范(含 2026-07-28 版)MCP / AAIF2024—2026Ahttps://modelcontextprotocol.io/https://modelcontextprotocol.io/specification/2026-07-28/
R22MCP Protocol Versions(五版规范演进表)MCP Ruby SDK2026Ahttps://ruby.sdk.modelcontextprotocol.io/protocol-versions/
R23mcpkit · ProtocolVersion 枚举mcpkit(docs.rs)2026Ahttps://docs.rs/mcpkit/latest/enum.ProtocolVersion.html

MCP 五版规范演进链(A 级,可全量引用)

版本关键变更
2024-11-05初始版:stdio + HTTP+SSE;三原语 Tools / Resources / Prompts
2025-03-26Streamable HTTP 取代 HTTP+SSE;OAuth 2.1;工具注解(readOnly / destructive / idempotent);音频内容;Completions
2025-06-18Elicitation;结构化工具输出;资源链接;保护资源元数据;MCP-Protocol-Version 头必需;移除 JSON-RPC batching
2025-11-25Tasks(异步状态跟踪);并行工具调用;服务端 agent 循环;sampling 中工具调用
2026-07-28协议核心无状态化;Extensions 框架(Tasks、MCP Apps);正式弃用策略(最短 12 个月窗口);新版弃用 Roots / Sampling / Logging

3.2. 中国国家标准

编号名称机构年份等级链接
R24《人工智能 智能体互联》系列国家标准(发布报道)人民网2026-07-09Ahttps://finance-app.people.cn/n1/2026/0709/c1004-40757059.html
R25《人工智能 智能体互联》系列国家标准解读中国产业经济信息网(归口单位解读)2026Ahttps://cinic.org.cn/xw/zcdt/1643418.html
R26GB/Z 185—2026 落地:长三角智能体身份码节点首批发放中国日报2026-09-04Ahttps://cn.chinadaily.com.cn/a/202609/04/WS6a9a6773e4b09a165c788098.html

GB/Z 185—2026 七部分结构(A 级)

  1. 第 1 部分 总体架构
  2. 第 2 部分 身份码(编码、分配与管理)——遵循 GB/T 26231,分层 OID 标识体系
  3. 第 3 部分 身份管理(注册、账户、凭证、鉴别)
  4. 第 4 部分 智能体描述(能力描述及注册、发布、变更)
  5. 第 5 部分 智能体发现(发现流程)
  6. 第 6 部分 智能体交互(点对点、群组、混合)
  7. 第 7 部分 外部工具调用(架构、流程、数据格式)

发布日期存在 2026-05-22 / 2026-06-26 / 2026-07-09 三种口径,本工程并列呈现,正文采用“2026 年上半年”。

3.3. 官方产品文档

编号名称机构年份等级链接
R27Claude Code 官方文档 · SandboxingAnthropic2026Ahttps://code.claude.com/docs/zh-TW/sandboxing
R28Claude Code 官方文档 · Choose a sandbox environmentAnthropic2026Ahttps://code.claude.com/docs/en/sandbox-environments

4. 学术论文

编号名称作者 / 机构年份等级链接
R29ReAct: Synergizing Reasoning and Acting in Language ModelsYao 等(Princeton / Google Brain)2022-10(arXiv:2210.03629,ICLR 2023)Ahttps://arxiv.org/abs/2210.03629
R30Toolformer: Language Models Can Teach Themselves to Use ToolsSchick 等(Meta AI)2023-02(arXiv:2302.04761,NeurIPS 2023)Ahttps://arxiv.org/abs/2302.04761
R31SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez、Yang 等(Princeton / Stanford)2023-10(arXiv:2310.06770,ICLR 2024 Oral)Ahttps://arxiv.org/abs/2310.06770
R32A Survey of Context Engineering for Large Language ModelsMei 等2025-07(arXiv:2507.13334)Bhttps://arxiv.org/abs/2507.13334
R33Context Engineering 2.02025-10(arXiv:2510.26493)Bhttps://arxiv.org/abs/2510.26493
R34Gorilla: Large Language Model Connected with Massive APIsUC Berkeley2023-05(arXiv:2305.15334)Ahttps://arxiv.org/abs/2305.15334

5. 方法与实践文献

编号名称作者 / 机构年份等级链接
R35Harness engineering for coding agent usersBirgitta Böckeler(Thoughtworks)/ martinfowler.com2026Ahttps://martinfowler.com/articles/harness-engineering.html
R36Humans and Agents in Software Engineering Loopsmartinfowler.com2026Ahttps://martinfowler.com/articles/exploring-gen-ai/humans-and-agents.html
R37My AI Adoption JourneyMitchell Hashimoto2026-02-05Ahttps://mitchellh.com/writing/my-ai-adoption-journey
R3812-Factor AgentsDex Horthy / HumanLayer2024—2025Ahttps://github.com/humanlayer/12-factor-agents
R3912-Factor Agents: Patterns of Reliable LLM Applications(MLOps Community 演讲)Dex Horthy2025-08-06Bhttps://home.mlops.community/public/videos/12-factor-agents-patterns-of-reliable-llm-applications-dexter-horthy-agents-in-production-2025-2025-08-06

6. 行业报告与调研

编号名称机构年份 / 日期等级链接
R402025 Stack Overflow Developer SurveyStack Overflow2025-07-29Ahttps://survey.stackoverflow.co/2025/
R41Stack Overflow 2025 Developer Survey 官方新闻稿Stack Overflow2025Ahttps://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/
R42DORA 2025 State of AI-assisted Software DevelopmentGoogle Cloud / DORA2025-09Ahttps://dora.dev/dora-report-2025
R43The State of AI 2025: Agents, Innovation, and TransformationMcKinsey2025-11-05Ahttps://www.mckinsey.com/
R44AI Adoption Stats & Trends(DORA / McKinsey / JetBrains 汇总)daily.dev2026-07-19Bhttps://daily.dev/agentic-ai-hub/ai-adoption-stats-trends/
R45Karpathy Puts Context at the Core of AI CodingPure AI2025-09-23Bhttps://pureai.com/articles/2025/09/23/karpathy-puts-context-at-the-core-of-ai-coding.aspx

Stack Overflow 2025 关键数字(A 级,可全量引用)

指标数值
样本量49,000+ 份回答 / 177 国 / 62 题 / 314 项技术(第 15 届)
采用率84%(2024 年 76%);专业开发者每日使用 51%
信任度46% 不信任 / 33% 信任 / 3.1% 高度信任
资深开发者2.6% 高度信任 / 20% 高度不信任
正面情绪60%(2023/2024 为 70%+)
agent 使用率约 31%(14.1% 每日 + 9% 每周 + 7.8% 月度或偶尔);37.9% 不打算用
生产力用过 agent 者中 69% 认为提升
最大挫败66% “AI 方案几乎对但不完全对”;45.2% 调试更耗时

7. 基准与评测

编号名称机构年份等级链接
R46SWE-bench 官方站与排行榜Princeton / 社区2023—2026Ahttps://www.swebench.com/https://swe-bench-live.github.io/
R47Terminal-Bench(官方站)Stanford / Laude Institute2025—2026Ahttps://www.tbench.ai/
R48Terminal-Bench 仓库(Harbor 框架)社区2025—2026Ahttps://github.com/harbor-framework/terminal-bench
R49SWE-bench Verified 进展时间线 2023—2026AgentMarketCap2026-04-09Chttps://agentmarketcap.ai/blog/2026/04/09/swe-bench-verified-progress-timeline-2023-2026
R50Terminal-Bench: The CLI Autonomy StandardAgentMarketCap2026-04-09Chttps://agentmarketcap.ai/blog/2026/04/09/terminal-bench-cli-autonomy-standard-coding-agents

使用警示:R49、R50 为 C 级来源,其中包含大量高冲击力的量化数字(如“仅改 Harness 提升 13.7 个百分点”)。这些数字对论证极有价值,但来源可靠性不足,引用时必须显式标注 [待核实]


8. 百科与二手综述

编号名称来源等级链接
R51Model Context ProtocolWikipediaBhttps://en.wikipedia.org/wiki/Model_Context_Protocol
R52Agent2Agent (A2A)WikipediaBhttps://en.wikipedia.org/wiki/Agent2Agent
R53Model Context Protocol(镜像)WikiwandBhttps://www.wikiwand.com/en/articles/Model_Context_Protocol
R54Model Context Protocol(规范演进与三角色)KluBhttp://klu.ai/glossary/model-context-protocol
R55Context engineeringAI WikiBhttps://aiwiki.ai/wiki/context_engineering
R56SWE-bench / SWE-bench VerifiedAI WikiBhttps://aiwiki.ai/wiki/swe_benchhttps://aiwiki.ai/wiki/swe_bench_verified
R57OpenAI CodexAI WikiBhttps://aiwiki.ai/wiki/codex
R58Claude CodeAI WikiBhttps://aiwiki.ai/wiki/Claude_Code
R59Google Agent Development KitAI WikiBhttps://aiwiki.ai/wiki/google_adk
R60Function callingAI WikiBhttps://aiwiki.ai/wiki/function_calling
R61Agentic AI FoundationAI WikiBhttps://aiwiki.ai/wiki/agentic_ai_foundation
R62The History of AI Agents: From SHRDLU to the Agent LoopTaskadeBhttps://taskade.com/blog/ai-agents-history
R63Anthropic Claude 发布时间线GitHub 公开仓库(第三方)Bhttps://github.com/jqueryscript/anthropic-claude-timeline
R64Harness 架构百度百科Bhttps://baike.baidu.com/item/Harness%E6%9E%B6%E6%9E%84/67704948
R73Agent Harness:2026 年 AI 工程的核心范式腾讯云开发者社区Chttps://developer.cloud.tencent.com/article/2698416
R74Anthropic 的 Harness 工程架构演进(中文综述)阿里云开发者社区Chttps://developer.aliyun.com/article/1724413

使用警示:R73、R74 为 C 级来源。其具体数字一律标 [待核实] 后方可入正文;其框架性论述(如“三代 Harness 架构演进”“scaffold 效应”“Harness 即数据集”)可作为叙事线索使用,但需与 A 级来源交叉印证。


9. 开源项目

编号名称机构 / 作者等级链接
R65Model Context Protocol 规范与 SDK(GitHub 组织)MCP / AAIFAhttps://github.com/modelcontextprotocol
R66OpenAI Codex CLIOpenAIAhttps://github.com/openai/codex
R67Claude Agent SDK / Claude CodeAnthropicAhttps://github.com/anthropics/claude-agent-sdk-python
R68Google Agent Development Kit (ADK)GoogleAhttps://github.com/google/adk-python
R69Agent2Agent (A2A) 协议Linux Foundation / GoogleAhttps://github.com/a2aproject/A2A
R7012-Factor AgentsHumanLayerAhttps://github.com/humanlayer/12-factor-agents
R71Terminal-Bench / Harbor 框架Stanford / Laude Institute / 社区Ahttps://github.com/harbor-framework/terminal-bench
R72@anthropic-ai/sandbox-runtimeAnthropicAhttps://www.npmjs.com/package/@anthropic-ai/sandbox-runtime

10. 信息缺口声明

10.1. 完全无结果(暂无权威信息)

  1. ISO/IEC 层面的智能体互联国际标准:未检索到 ISO/IEC 已发布或已立项的智能体互联国际标准编号。仅确认中国 GB/Z 185—2026 自称“全球首套系统性智能体互联标准体系”,该表述出自中国媒体,未获国际方交叉印证
  • "Agent Harness"术语的首创者与首次出现出处:未找到确切的一手首创文献。可确认的是 Anthropic 于 2025 年已用 "harness" 描述 Claude Agent SDK;Mitchell Hashimoto 于 2026-02-05 命名 "Harness Engineering";OpenAI 于 2026-02-11 将其推向主流。更早的溯源(如 LangChain《The Anatomy of an Agent Harness》博客)未检索到原文与确切发布日
  • Wikipedia "Test harness" 条目原文:未直接抓取,仅经二手转述。
  • DORA 2025 报告官方 URL:仅见 https://dora.dev/dora-report-2025 的转述,未验证可访问性(见 R42)。
  • McKinsey《The State of AI 2025》官方 URL:同上(见 R43)。
  • Terminal-Bench 官方站(tbench.ai)与 leaderboard 当前数据:未直接抓取,所有榜单数字均为第三方转述(见 R47)。
  • AGNTCY 项目的权威一手资料:仅见于 AAIF 相关综述中的提及,未深入核实。
  • Harness 层自身的市场规模权威测算:未检索到,属 [待填写]
  • 10.2. 存在冲突,需择一或并列标注

    1. GB/Z 185—2026 发布日期:中国日报记 2026-05-22;百度百科记 2026-06-26;人民网报道日 2026-07-09(用语“近日发布”)。本工程并列三口径,正文采用“2026 年上半年”。
    2. AAIF 成立日期:TechCrunch 与 Wikipedia 记 2025-12-09;GIGAZINE 记 2025-12-10。本工程采用 2025-12-09。
    3. Terminal-Bench 2.0 发布时间:一处记 2025 年末,另一处记 2026-01。本工程写作“2025 年末至 2026 年初”。
    4. Anthropic computer use 发布时间:Taskade 时间线记 2023-10,与 Anthropic 官方 2024-10-22(公开 beta)冲突。本工程采信官方 2024-10,标 [存疑]
    5. AutoGPT / BabyAGI 发布月份:2023-03 与 2023-04 两种说法。本工程写作“2023 年 3—4 月”。
    6. MCP “首个公开规范版本”:2024-11-05(Ruby SDK 记为 Initial protocol revision)与 2024-11-25(公开宣布与生态启动)可并存表述。

    10.3. 数字来源等级不足,须标 后方可引用

    1. 全部 SWE-bench Verified 2026 年 SOTA 数字(87.6% / 88.7% / 88.6% / 93.9%)
    2. 全部 Terminal-Bench 2.0 榜单数字(78.4% / 77.3% / 75.1% / 74.7% / 71.9%)
      1. LangChain 52.8% → 66.5%、Vercel 80% → 100%、Can Duruk 6.7% → 68.3% 三组“仅改 Harness”实验
    3. Claude Code 350,000 DAU / 100 万合并 PR / Anthropic 内部约 25% 代码提交
    4. Codex Auto-review 的 1/200 与 99%,以及"GPT-5.4 Thinking"这一型号名
      1. Codex CLI “2026 年初约 95% Rust”
    5. Anthropic 三 Agent 实验的 20 分钟 / $9 vs 6 小时 / $200
    6. DORA 2025 / McKinsey 2025 / JetBrains 2025 的全部数字(均来自汇总站转引)
      1. 上下文退化“性能下降超 45%”与 Chroma “18 个前沿模型全部退化”
    7. 中国厂商动向(DeepSeek 2026-05-20 组建 Harness 团队、小米 2026-06-11 MiMo Code V0.1.0、灵犀智涌 2026-08 ROSS)——均仅见于百度百科词条转述,强烈建议二次核实
    8. Claude 模型 2026 年各版本时间线(Opus 4.6 / 4.7 / 4.8、Sonnet 5、Opus 5 等)——仅见于第三方 GitHub 时间线仓库,需与 Anthropic 官方公告核对(见 R63)
    9. 市场规模预测:Grand View Research(2025 年约 222 亿美元 → 2033 年约 3,247 亿美元,CAGR 40.8%)与 Bloomberg Intelligence(2032 年约 2.3 万亿美元)——两者口径差异近一个数量级,仅宜作方向性引用

    10.4. 建议补充检索的方向

    1. 中国信通院、中国人工智能产业发展联盟(AIIA)关于智能体的团体标准 / 行业规范
    2. OWASP Agentic AI Top 10 等安全侧规范(本次未检索;L6 治理层可能需要)
    3. IEEE 关于 AI Agent 的标准立项情况
    4. 欧盟 AI Act 中与自主智能体相关的条款(如适用)
    5. Harness 层自身的市场规模与商业化形态测算

    References

    1. How to Use This Document

    图 1-1|参考资料来源三级可信度分级体系(A / B / C)

    参考资料来源三级可信度分级体系(A / B / C) 共收录 74 条(R1–R74)· 示意:基于本文分析绘制 A 级 · 官方一手来源(51 条) 使用规则:可直接引用,须给出出处 官方工程博客 / 新闻公告 Anthropic / OpenAI / Google / Linux Foundation arXiv 原文 · 官方开源仓库 · 标准解读 MCP 规范 · GB/Z 185 国标 · Stack Overflow 报告 可信度递减 B 级 · 权威二手来源(19 条) 使用规则:需注明转述;指标建议二次核对 百科与二手综述 Wikipedia · Wikiwand · AI Wiki · Klu 社区平台与第三方时间线 MLOps Community · daily.dev · Pure AI · Taskade · GitHub 仓库 需二次核实 C 级 · 社区与自媒体解读(4 条) 使用规则:仅作线索;具体数字一律标 后方可入正文 数据解读博客 AgentMarketCap 等自媒体 云社区开发者文章 腾讯云 / 阿里云开发者社区;框架论述须与 A 级交叉印证 结构解读:来源按可信度分为三级——A 级可直接引用、B 级注明转述、C 级仅作线索, 下游文档引用时须按等级选择引用方式,低等级数字一律标注 。

    数据来源:基于本文分析绘制的示意图。

    1.1. Source Grading Criteria

    All materials collected in this document are graded into three tiers of credibility, to help downstream documents decide how to cite them:

    GradeMeaningUsage rulesCount in this document
    AOfficial first-hand sources from vendors or institutions: Anthropic, OpenAI, Google, Linux Foundation / AAIF, Stack Overflow, People's Daily, China Daily, China Industry-Economy Information Network (interpretation by the standards-owning body), original arXiv papers, official open-source repositories, etc.Can be cited directly; the source must be attributed51 items
    BAuthoritative secondary sources: Wikipedia, Wikiwand, AI Wiki, MLOps Community, daily.dev, Pure AI, Taskade, Baidu Baike, third-party timeline repositories, etc.Must note that it is a retelling; figures are advised to be double-checked19 items
    CCommunity and self-media interpretations: AgentMarketCap, Tencent Cloud Developer Community, Alibaba Cloud Developer Community, etc.For leads only; any specific figures must be marked [To be verified] before entering the body text4 items

    74 items in total (R1–R74).

    1.2. Numbering Rules

    • IDs take the form R1R55 and are continuous across sections;
    • If the same entry appears in multiple categories, it is assigned to its primary category and is not renumbered;
    • Entries marked [To be verified] have links or dates not verified against first-hand sources and must be re-confirmed before citing;
    • Entries marked [Disputed] reflect conflicts between sources and are presented side by side in Section 10 of 02-发展历史.md.

    2. Official Engineering Blogs

    2.1. Anthropic

    IDTitleYear / DateGradeLink
    R1Effective context engineering for AI agents (Applied AI team)2025Ahttps://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
    R2Effective harnesses for long-running agents2025Ahttps://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
    R3Harness design for long-running application development2026Ahttps://www.anthropic.com/engineering/harness-design-long-running-apps
    R4Equipping agents for the real world with Agent Skills2025-10-16Ahttps://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
    R5Sandboxing: a safer and more autonomous approach2025Ahttps://www.anthropic.com/engineering/claude-code-sandboxing
    R6Introducing Agent Skills (converted to an open standard on 2025-12-18)2025-10-16Ahttps://www.anthropic.com/news/skills
    R7Introducing the Model Context Protocol2024-11-25Ahttps://www.anthropic.com/news/model-context-protocol
    R8Claude 3.7 Sonnet and Claude Code2025-02-24Ahttps://www.anthropic.com/news/claude-3-7-sonnet
    R9Donating the Model Context Protocol and establishing the AAIF2025-12Ahttps://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
    R10How we built our multi-agent research system2025Ahttps://www.anthropic.com/engineering/multi-agent-research-system

    2.2. OpenAI

    IDTitleYear / DateGradeLink
    R11Harness engineering: leveraging Codex in an agent-first world (Ryan Lopopolo)2026-02-11Ahttps://openai.com/index/harness-engineering/; Chinese version https://openai.com/zh-Hans-CN/index/harness-engineering/
    R12Function calling and other API updates2023-06-13Ahttps://openai.com/blog/function-calling-and-other-API-updates
    R13New tools for building agents (Responses API + Agents SDK)2025-03-11Ahttps://openai.com/blog/new-tools-for-building-agents
    R14Introducing Codex (Codex CLI)2025-04-16Ahttps://github.com/openai/codex
    R15OpenAI co-founds the Agentic AI Foundation under the Linux Foundation2025-12Ahttps://openai.com/index/agentic-ai-foundation/

    2.3. Google

    IDTitleYear / DateGradeLink
    R16Agent Development Kit: Making it easy to build multi-agent applications2025-04-09Ahttps://googledevelopers.blogspot.com/en/agent-development-kit-easy-to-build-multi-agent-applications/
    R17A year of open collaboration: Celebrating the anniversary of A2A2026-04-16Ahttps://opensource.googleblog.com/

    2.4. Linux Foundation and AAIF


    3. Standards and Specifications

    3.1. Protocol Specifications

    IDTitleOrganizationYearGradeLink
    R21Model Context Protocol official site and specification (including the 2026-07-28 edition)MCP / AAIF2024—2026Ahttps://modelcontextprotocol.io/; https://modelcontextprotocol.io/specification/2026-07-28/
    R22MCP Protocol Versions (the five-version specification evolution table)MCP Ruby SDK2026Ahttps://ruby.sdk.modelcontextprotocol.io/protocol-versions/
    R23mcpkit · ProtocolVersion enummcpkit (docs.rs)2026Ahttps://docs.rs/mcpkit/latest/enum.ProtocolVersion.html

    MCP five-version specification evolution chain (Grade A, can be cited in full)

    VersionKey changes
    2024-11-05Initial version: stdio + HTTP+SSE; three primitives Tools / Resources / Prompts
    2025-03-26Streamable HTTP replaces HTTP+SSE; OAuth 2.1; tool annotations (readOnly / destructive / idempotent); audio content; Completions
    2025-06-18Elicitation; structured tool output; resource links; protected resource metadata; MCP-Protocol-Version header required; JSON-RPC batching removed
    2025-11-25Tasks (asynchronous state tracking); parallel tool calls; server-side agent loop; tool calls during sampling
    2026-07-28Statelessness of protocol core; Extensions framework (Tasks, MCP Apps); formal deprecation policy (minimum 12-month window); the new version deprecates Roots / Sampling / Logging

    3.2. Chinese National Standards

    IDTitleOrganizationYearGradeLink
    R24"Artificial Intelligence: Intelligent Agent Interconnection" series national standard (release coverage)People’s Daily2026-07-09Ahttps://finance-app.people.cn/n1/2026/0709/c1004-40757059.html
    R25Interpretation of the "Artificial Intelligence: Intelligent Agent Interconnection" series national standardChina Industry-Economy Information Network (interpretation by the standards-owning body)2026Ahttps://cinic.org.cn/xw/zcdt/1643418.html
    R26GB/Z 185-2026 implementation: first batch of identity-code nodes issued in the Yangtze River DeltaChina Daily2026-09-04Ahttps://cn.chinadaily.com.cn/a/202609/04/WS6a9a6773e4b09a165c788098.html

    GB/Z 185-2026 seven-part structure (Grade A)

    1. Part 1: Overall architecture
    2. Part 2: Identity code (encoding, allocation and management) - follows GB/T 26231, layered OID identification system
    3. Part 3: Identity management (registration, account, credential, authentication)
    4. Part 4: Intelligent agent description (capability description and registration, publication, change)
    5. Part 5: Intelligent agent discovery (discovery process)
    6. Part 6: Intelligent agent interaction (point-to-point, group, hybrid)
    7. Part 7: External tool invocation (architecture, process, data format)

    Three different dates have been reported for the release date: 2026-05-22 / 2026-06-26 / 2026-07-09. This project presents all three side by side, and uses "first half of 2026" in the body text.

    3.3. Official Product Documentation

    IDTitleOrganizationYearGradeLink
    R27Claude Code official documentation · SandboxingAnthropic2026Ahttps://code.claude.com/docs/zh-TW/sandboxing
    R28Claude Code official documentation · Choose a sandbox environmentAnthropic2026Ahttps://code.claude.com/docs/en/sandbox-environments

    4. Academic Papers

    IDTitleAuthor / InstitutionYearGradeLink
    R29ReAct: Synergizing Reasoning and Acting in Language ModelsYao et al. (Princeton / Google Brain)2022-10 (arXiv:2210.03629, ICLR 2023)Ahttps://arxiv.org/abs/2210.03629
    R30Toolformer: Language Models Can Teach Themselves to Use ToolsSchick et al. (Meta AI)2023-02 (arXiv:2302.04761, NeurIPS 2023)Ahttps://arxiv.org/abs/2302.04761
    R31SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez, Yang et al. (Princeton / Stanford)2023-10 (arXiv:2310.06770, ICLR 2024 Oral)Ahttps://arxiv.org/abs/2310.06770
    R32A Survey of Context Engineering for Large Language ModelsMei et al.2025-07 (arXiv:2507.13334)Bhttps://arxiv.org/abs/2507.13334
    R33Context Engineering 2.02025-10 (arXiv:2510.26493)Bhttps://arxiv.org/abs/2510.26493
    R34Gorilla: Large Language Model Connected with Massive APIsUC Berkeley2023-05 (arXiv:2305.15334)Ahttps://arxiv.org/abs/2305.15334

    5. Methodology and Practice Literature

    IDTitleAuthor / InstitutionYearGradeLink
    R35Harness engineering for coding agent usersBirgitta Böckeler (Thoughtworks) / martinfowler.com2026Ahttps://martinfowler.com/articles/harness-engineering.html
    R36Humans and Agents in Software Engineering Loopsmartinfowler.com2026Ahttps://martinfowler.com/articles/exploring-gen-ai/humans-and-agents.html
    R37My AI Adoption JourneyMitchell Hashimoto2026-02-05Ahttps://mitchellh.com/writing/my-ai-adoption-journey
    R3812-Factor AgentsDex Horthy / HumanLayer2024—2025Ahttps://github.com/humanlayer/12-factor-agents
    R3912-Factor Agents: Patterns of Reliable LLM Applications (MLOps Community talk)Dex Horthy2025-08-06Bhttps://home.mlops.community/public/videos/12-factor-agents-patterns-of-reliable-llm-applications-dexter-horthy-agents-in-production-2025-2025-08-06

    6. Industry Reports and Research

    IDTitleOrganizationYear / DateGradeLink
    R402025 Stack Overflow Developer SurveyStack Overflow2025-07-29Ahttps://survey.stackoverflow.co/2025/
    R41Stack Overflow 2025 Developer Survey official press releaseStack Overflow2025Ahttps://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/
    R42DORA 2025 State of AI-assisted Software DevelopmentGoogle Cloud / DORA2025-09Ahttps://dora.dev/dora-report-2025
    R43The State of AI 2025: Agents, Innovation, and TransformationMcKinsey2025-11-05Ahttps://www.mckinsey.com/
    R44AI Adoption Stats & Trends (compiled from DORA / McKinsey / JetBrains)daily.dev2026-07-19Bhttps://daily.dev/agentic-ai-hub/ai-adoption-stats-trends/
    R45Karpathy Puts Context at the Core of AI CodingPure AI2025-09-23Bhttps://pureai.com/articles/2025/09/23/karpathy-puts-context-at-the-core-of-ai-coding.aspx

    Key figures from Stack Overflow 2025 (Grade A, can be cited in full)

    MetricValue
    Sample size49,000+ responses / 177 countries / 62 questions / 314 technologies (15th edition)
    Adoption rate84% (76% in 2024); 51% of professional developers use it daily
    Trust46% do not trust / 33% trust / 3.1% trust highly
    Senior developers2.6% trust highly / 20% distrust highly
    Positive sentiment60% (70%+ in 2023/2024)
    Agent usageabout 31% (14.1% daily + 9% weekly + 7.8% monthly or occasionally); 37.9% do not plan to use it
    Productivity69% of those who have used agents report an improvement
    Biggest frustration66% "the AI solution is almost right but not quite"; 45.2% find debugging takes longer

    7. Benchmarks and Evaluation

    IDTitleOrganizationYearGradeLink
    R46SWE-bench official site and leaderboardPrinceton / community2023—2026Ahttps://www.swebench.com/; https://swe-bench-live.github.io/
    R47Terminal-Bench (official site)Stanford / Laude Institute2025—2026Ahttps://www.tbench.ai/
    R48Terminal-Bench repository (Harbor framework)community2025—2026Ahttps://github.com/harbor-framework/terminal-bench
    R49SWE-bench Verified progress timeline 2023-2026AgentMarketCap2026-04-09Chttps://agentmarketcap.ai/blog/2026/04/09/swe-bench-verified-progress-timeline-2023-2026
    R50Terminal-Bench: The CLI Autonomy StandardAgentMarketCap2026-04-09Chttps://agentmarketcap.ai/blog/2026/04/09/terminal-bench-cli-autonomy-standard-coding-agents

    Usage warning: R49 and R50 are Grade C sources and contain many high-impact quantified figures (e.g. "changing only the Harness improved performance by 13.7 percentage points"). These figures are highly valuable to the argument, but their source reliability is insufficient, so they must be explicitly marked [To be verified] before citing.


    8. Encyclopedia and Secondary Surveys

    IDTitleSourceGradeLink
    R51Model Context ProtocolWikipediaBhttps://en.wikipedia.org/wiki/Model_Context_Protocol
    R52Agent2Agent (A2A)WikipediaBhttps://en.wikipedia.org/wiki/Agent2Agent
    R53Model Context Protocol (mirror)WikiwandBhttps://www.wikiwand.com/en/articles/Model_Context_Protocol
    R54Model Context Protocol (specification evolution and three roles)KluBhttp://klu.ai/glossary/model-context-protocol
    R55Context engineeringAI WikiBhttps://aiwiki.ai/wiki/context_engineering
    R56SWE-bench / SWE-bench VerifiedAI WikiBhttps://aiwiki.ai/wiki/swe_bench; https://aiwiki.ai/wiki/swe_bench_verified
    R57OpenAI CodexAI WikiBhttps://aiwiki.ai/wiki/codex
    R58Claude CodeAI WikiBhttps://aiwiki.ai/wiki/Claude_Code
    R59Google Agent Development KitAI WikiBhttps://aiwiki.ai/wiki/google_adk
    R60Function callingAI WikiBhttps://aiwiki.ai/wiki/function_calling
    R61Agentic AI FoundationAI WikiBhttps://aiwiki.ai/wiki/agentic_ai_foundation
    R62The History of AI Agents: From SHRDLU to the Agent LoopTaskadeBhttps://taskade.com/blog/ai-agents-history
    R63Anthropic Claude release timelineGitHub public repository (third party)Bhttps://github.com/jqueryscript/anthropic-claude-timeline
    R64Harness architectureBaidu BaikeBhttps://baike.baidu.com/item/Harness%E6%9E%B6%E6%9E%84/67704948
    R73Agent Harness: the core paradigm of AI engineering in 2026Tencent Cloud Developer CommunityChttps://developer.cloud.tencent.com/article/2698416
    R74Anthropic's Harness engineering architecture evolution (Chinese review)Alibaba Cloud Developer CommunityChttps://developer.aliyun.com/article/1724413

    Usage warning: R73 and R74 are Grade C sources. Any of their specific figures must be marked [To be verified] before entering the body text; their framework-level arguments (such as "the three-generation Harness architecture evolution", "the scaffold effect", "Harness is the dataset") can be used as narrative threads, but must be cross-verified against Grade A sources.


    9. Open-Source Projects

    IDTitleOrganization / AuthorGradeLink
    R65Model Context Protocol specification and SDK (GitHub organization)MCP / AAIFAhttps://github.com/modelcontextprotocol
    R66OpenAI Codex CLIOpenAIAhttps://github.com/openai/codex
    R67Claude Agent SDK / Claude CodeAnthropicAhttps://github.com/anthropics/claude-agent-sdk-python
    R68Google Agent Development Kit (ADK)GoogleAhttps://github.com/google/adk-python
    R69Agent2Agent (A2A) protocolLinux Foundation / GoogleAhttps://github.com/a2aproject/A2A
    R7012-Factor AgentsHumanLayerAhttps://github.com/humanlayer/12-factor-agents
    R71Terminal-Bench / Harbor frameworkStanford / Laude Institute / communityAhttps://github.com/harbor-framework/terminal-bench
    R72@anthropic-ai/sandbox-runtimeAnthropicAhttps://www.npmjs.com/package/@anthropic-ai/sandbox-runtime

    10. Information Gap Statement

    10.1. No Results (No Authoritative Information Available)

    1. International standard for intelligent agent interconnection at the ISO/IEC level: no ISO/IEC standard number for intelligent agent interconnection that has been published or is under development was found. It was only confirmed that China’s GB/Z 185-2026 claims to be the "world’s first systematic intelligent agent interconnection standard system", a statement that comes from Chinese media and has not been cross-verified by international parties.
    2. Originator and first occurrence of the term "Agent Harness": no definitive first-hand origin document was found. It can be confirmed that Anthropic used "harness" to describe Claude Agent SDK in 2025; Mitchell Hashimoto coined "Harness Engineering" on 2026-02-05; OpenAI brought it into the mainstream on 2026-02-11. For earlier origins (such as the LangChain blog "The Anatomy of an Agent Harness"), the original text and an exact publication date could not be found.
    3. Original text of the Wikipedia "Test harness" entry: not directly fetched, only retold through secondary sources.
    4. Official URL of the DORA 2025 report: only a retelling of https://dora.dev/dora-report-2025 was found; accessibility was not verified (see R42).
    5. Official URL of McKinsey’s "The State of AI 2025": same as above (see R43).
    6. Terminal-Bench official site (tbench.ai) and current leaderboard data: not directly fetched; all leaderboard numbers are third-party retellings (see R47).
    7. Authoritative first-hand materials on the AGNTCY project: found only as a mention in AAIF-related reviews; not examined in depth.
    8. Authoritative estimate of the market size of the Harness layer itself: not found; it is [To be filled].

    10.2. Conflicts Exist; Choose One or Present Side by Side

    1. Release date of GB/Z 185-2026: China Daily records 2026-05-22; Baidu Baike records 2026-06-26; People’s Daily reported it on 2026-07-09 (using the wording "recently released"). This project presents all three dates side by side and uses "first half of 2026" in the body text.
    2. Foundation date of the AAIF: TechCrunch and Wikipedia record 2025-12-09; GIGAZINE records 2025-12-10. This project uses 2025-12-09.
    3. Release time of Terminal-Bench 2.0: one source records late 2025, another records 2026-01. This project writes "late 2025 to early 2026".
    4. Release time of Anthropic computer use: the Taskade timeline records 2023-10, conflicting with Anthropic official 2024-10-22 (public beta). This project trusts the official 2024-10 and marks it [Disputed].
    5. Release months of AutoGPT / BabyAGI: two claims, 2023-03 and 2023-04. This project writes "March-April 2023".
    6. MCP "first public specification version": 2024-11-05 (recorded as the Initial protocol revision by the Ruby SDK) and 2024-11-25 (public announcement and ecosystem launch) can be described side by side.

    10.3. Figures of Insufficient Source Grade; Must Be Marked [To be verified] Before Citing

    1. All SWE-bench Verified 2026 SOTA figures (87.6% / 88.7% / 88.6% / 93.9%)
    2. All Terminal-Bench 2.0 leaderboard figures (78.4% / 77.3% / 75.1% / 74.7% / 71.9%)
    3. The three "changing only the Harness" experiments: LangChain 52.8%→66.5%, Vercel 80%→100%, Can Duruk 6.7%→68.3%
    4. Claude Code 350,000 DAU / 1 million merged PRs / roughly 25% of code commits inside Anthropic
    5. Codex Auto-review’s 1/200 and 99%, as well as the model name "GPT-5.4 Thinking"
    6. Codex CLI "about 95% Rust in early 2026"
    7. Anthropic’s three-agent experiment: 20 minutes / $9 vs. 6 hours / $200
    8. All figures from DORA 2025 / McKinsey 2025 / JetBrains 2025 (all relayed from aggregation sites)
    9. Context degradation "performance drops by more than 45%" and Chroma "all 18 frontier models degrade"
    10. Movements of Chinese vendors (DeepSeek formed a Harness team on 2026-05-20, Xiaomi’s MiMo Code V0.1.0 on 2026-06-11, Lingxi Zhiyong’s ROSS in 2026-08) - all found only in Baidu Baike entry retellings; a second verification is strongly recommended
    11. The 2026 version timeline of Claude models (Opus 4.6 / 4.7 / 4.8, Sonnet 5, Opus 5, etc.) - found only in a third-party GitHub timeline repository; needs to be checked against Anthropic official announcements (see R63)
    12. Market size forecasts: Grand View Research (about $22.2 billion in 2025→about $324.7 billion in 2033, CAGR 40.8%) and Bloomberg Intelligence (about $2.3 trillion in 2032) - the two estimates differ by nearly an order of magnitude, so they should only be cited as directional

    10.4. Suggested Directions for Additional Search

    1. Group standards / industry norms on intelligent agents from the China Academy of Information and Communications Technology (CAICT) and the China Alliance of Artificial Intelligence Industry (AIIA)
    2. Security-side standards such as the OWASP Agentic AI Top 10 (not searched this time; the L6 governance layer may need them)
    3. The status of IEEE standard initiatives on AI agents
    4. Provisions in the EU AI Act relevant to autonomous intelligent agents (where applicable)
    5. Estimate of the market size and commercialization forms of the Harness layer itself