All materials collected in this document are graded into three tiers of credibility, to help downstream documents decide how to cite them:
Grade
Meaning
Usage rules
Count in this document
A
Official first-hand sources from vendors or institutions: Anthropic, OpenAI, Google, Linux Foundation / AAIF, Stack Overflow, People's Daily, China Daily, China Industry-Economy Information Network (interpretation by the standards-owning body), original arXiv papers, official open-source repositories, etc.
Can be cited directly; the source must be attributed
51 items
B
Authoritative secondary sources: Wikipedia, Wikiwand, AI Wiki, MLOps Community, daily.dev, Pure AI, Taskade, Baidu Baike, third-party timeline repositories, etc.
Must note that it is a retelling; figures are advised to be double-checked
19 items
C
Community and self-media interpretations: AgentMarketCap, Tencent Cloud Developer Community, Alibaba Cloud Developer Community, etc.
For leads only; any specific figures must be marked [To be verified] before entering the body text
4 items
74 items in total (R1–R74).
1.2. Numbering Rules
IDs take the form R1–R55 and are continuous across sections;
If the same entry appears in multiple categories, it is assigned to its primary category and is not renumbered;
Entries marked [To be verified] have links or dates not verified against first-hand sources and must be re-confirmed before citing;
Entries marked [Disputed] reflect conflicts between sources and are presented side by side in Section 10 of 02-发展历史.md.
2. Official Engineering Blogs
2.1. Anthropic
ID
Title
Year / Date
Grade
Link
R1
Effective context engineering for AI agents (Applied AI team)
Part 2: Identity code (encoding, allocation and management) - follows GB/T 26231, layered OID identification system
Part 3: Identity management (registration, account, credential, authentication)
Part 4: Intelligent agent description (capability description and registration, publication, change)
Part 5: Intelligent agent discovery (discovery process)
Part 6: Intelligent agent interaction (point-to-point, group, hybrid)
Part 7: External tool invocation (architecture, process, data format)
Three different dates have been reported for the release date: 2026-05-22 / 2026-06-26 / 2026-07-09. This project presents all three side by side, and uses "first half of 2026" in the body text.
Usage warning: R49 and R50 are Grade C sources and contain many high-impact quantified figures (e.g. "changing only the Harness improved performance by 13.7 percentage points"). These figures are highly valuable to the argument, but their source reliability is insufficient, so they must be explicitly marked [To be verified] before citing.
Usage warning: R73 and R74 are Grade C sources. Any of their specific figures must be marked [To be verified] before entering the body text; their framework-level arguments (such as "the three-generation Harness architecture evolution", "the scaffold effect", "Harness is the dataset") can be used as narrative threads, but must be cross-verified against Grade A sources.
9. Open-Source Projects
ID
Title
Organization / Author
Grade
Link
R65
Model Context Protocol specification and SDK (GitHub organization)
10.1. No Results (No Authoritative Information Available)
International standard for intelligent agent interconnection at the ISO/IEC level: no ISO/IEC standard number for intelligent agent interconnection that has been published or is under development was found. It was only confirmed that China’s GB/Z 185-2026 claims to be the "world’s first systematic intelligent agent interconnection standard system", a statement that comes from Chinese media and has not been cross-verified by international parties.
Originator and first occurrence of the term "Agent Harness": no definitive first-hand origin document was found. It can be confirmed that Anthropic used "harness" to describe Claude Agent SDK in 2025; Mitchell Hashimoto coined "Harness Engineering" on 2026-02-05; OpenAI brought it into the mainstream on 2026-02-11. For earlier origins (such as the LangChain blog "The Anatomy of an Agent Harness"), the original text and an exact publication date could not be found.
Original text of the Wikipedia "Test harness" entry: not directly fetched, only retold through secondary sources.
Official URL of the DORA 2025 report: only a retelling of https://dora.dev/dora-report-2025 was found; accessibility was not verified (see R42).
Official URL of McKinsey’s "The State of AI 2025": same as above (see R43).
Terminal-Bench official site (tbench.ai) and current leaderboard data: not directly fetched; all leaderboard numbers are third-party retellings (see R47).
Authoritative first-hand materials on the AGNTCY project: found only as a mention in AAIF-related reviews; not examined in depth.
Authoritative estimate of the market size of the Harness layer itself: not found; it is [To be filled].
10.2. Conflicts Exist; Choose One or Present Side by Side
Release date of GB/Z 185-2026: China Daily records 2026-05-22; Baidu Baike records 2026-06-26; People’s Daily reported it on 2026-07-09 (using the wording "recently released"). This project presents all three dates side by side and uses "first half of 2026" in the body text.
Foundation date of the AAIF: TechCrunch and Wikipedia record 2025-12-09; GIGAZINE records 2025-12-10. This project uses 2025-12-09.
Release time of Terminal-Bench 2.0: one source records late 2025, another records 2026-01. This project writes "late 2025 to early 2026".
Release time of Anthropic computer use: the Taskade timeline records 2023-10, conflicting with Anthropic official 2024-10-22 (public beta). This project trusts the official 2024-10 and marks it [Disputed].
Release months of AutoGPT / BabyAGI: two claims, 2023-03 and 2023-04. This project writes "March-April 2023".
MCP "first public specification version": 2024-11-05 (recorded as the Initial protocol revision by the Ruby SDK) and 2024-11-25 (public announcement and ecosystem launch) can be described side by side.
10.3. Figures of Insufficient Source Grade; Must Be Marked [To be verified] Before Citing
All SWE-bench Verified 2026 SOTA figures (87.6% / 88.7% / 88.6% / 93.9%)
All figures from DORA 2025 / McKinsey 2025 / JetBrains 2025 (all relayed from aggregation sites)
Context degradation "performance drops by more than 45%" and Chroma "all 18 frontier models degrade"
Movements of Chinese vendors (DeepSeek formed a Harness team on 2026-05-20, Xiaomi’s MiMo Code V0.1.0 on 2026-06-11, Lingxi Zhiyong’s ROSS in 2026-08) - all found only in Baidu Baike entry retellings; a second verification is strongly recommended
The 2026 version timeline of Claude models (Opus 4.6 / 4.7 / 4.8, Sonnet 5, Opus 5, etc.) - found only in a third-party GitHub timeline repository; needs to be checked against Anthropic official announcements (see R63)
Market size forecasts: Grand View Research (about $22.2 billion in 2025→about $324.7 billion in 2033, CAGR 40.8%) and Bloomberg Intelligence (about $2.3 trillion in 2032) - the two estimates differ by nearly an order of magnitude, so they should only be cited as directional
10.4. Suggested Directions for Additional Search
Group standards / industry norms on intelligent agents from the China Academy of Information and Communications Technology (CAICT) and the China Alliance of Artificial Intelligence Industry (AIIA)
Security-side standards such as the OWASP Agentic AI Top 10 (not searched this time; the L6 governance layer may need them)
The status of IEEE standard initiatives on AI agents
Provisions in the EU AI Act relevant to autonomous intelligent agents (where applicable)
Estimate of the market size and commercialization forms of the Harness layer itself