AI4S · 科学发现


1. Introduction

1.1 Background

AI for Science(AI4S,人工智能驱动的科学研究)是智能体技术中科学回报最高、也最容易被叙事透支的方向。一方面,2018 年以来的蛋白质结构预测、2023 年以来的气象大模型与材料发现,标志着 AI 第一次在若干核心科学任务上达到或超越传统方法的水平;另一方面,自主「AI 科学家」系统被独立评估发现存在新颖性误判与数值幻觉,AI 气象模型被证实系统性低估极端事件——没有严格验证层的 AI4S,就是幻觉放大器

政策侧同样在加速。科技部与国家自然科学基金委于 2023-03 启动「人工智能驱动的科学研究」专项部署,围绕药物研发、基因研究、生物育种、新材料研发等重点领域,推进算法模型创新、科研数据开放共享与算力平台建设,并支持高性能计算中心与智算中心异构融合发展。中国气象局则于 2026-05 启动第二届人工智能气象预报模型示范计划,并发布《人工智能气象预报模型预报性能检验评估》等气象行业标准——「检验评估」成为业务准入的前提,这与 Harness 六层模型中 L5 的角色高度同构。

本方向与其他五个方向有一个根本差异,值得先行声明:这里的 ground truth 不是人类标注数据集,而是物理定律、实验验证与守恒律约束。一条 SQL 可以比对,一份综述可以核查引用,但一个由 AI 提出的晶体结构,只有物理定律与实验合成才能裁决它是否成立。

1.2 定义

科学发现方向的 AI Harness,是指围绕假说生成、文献证据装配、模拟与实验规划、结果验证与结论分级,为智能体提供领域上下文装配、模拟与实验工具契约、lab-in-the-loop 编排、假说与失败记录记忆、以及以物理定律和实验验证为核心评估依据的工程承载层。

它不替代领域科学家的判断,不替代实验本身,也不替代同行评审。它承担的是「让智能体的每一次科学主张都被可验证的证据约束住」的那一层。

边界上需要澄清三点:

  • AI4S Harness 不是「自动出论文的机器」。自主科学发现系统的独立评估已经表明,未经严格验证的自动产出会以体面的形式携带错误(详见 3.3 节)。
  • AI4S Harness 不做实验裁决。物理实验与临床验证的执行权在人,智能体负责假说、设计与解读。
  • AI4S Harness 的评估成本结构特殊:廉价验证(守恒律、收敛性、与既有模型对照)可以高频执行,昂贵验证(实验合成、临床试验)必须人在回路、按需触发。

1.3 在 AI Harness 体系中的定位——L5 是第一性问题

图 1-1|AI4S 六层能力模型:L5 第一性问题与 L3 lab-in-the-loop 编排

AI4S 六层能力模型:L5 第一性问题,L3 lab-in-the-loop 侧重评级基于正文六层能力模型分析 · 示意:基于本文分析绘制 L6 治理与安全层 侧重 ★★★★★ 科研诚信 · 伦理审查 · 生物安全 · 数据汇交 L5 评估与观测层 · 第一性问题(本图重点) 侧重 ★★★★★ ground truth = 物理定律 · 实验 · 守恒律 | 廉价高频 · 昂贵按需 L4 记忆与状态层 侧重 ★★★★ 假说库 · 引用库 · 实验与失败记录 L3 编排与控制层 · lab-in-the-loop(本图重点) 侧重 ★★★★★ 智能体做假说 · 设计,人做定义 · 执行 L2 工具与执行层 侧重 ★★★★★ 模拟实验 MCP · 实验跟踪 · 仪器接口 L1 上下文工程层 侧重 ★★★★ 文献证据 · 领域数据库 · 失败实验记录 结构解读:AI4S 的核心在 L5(ground truth 是物理定律/实验/守恒律)与 L3(lab-in-the-loop),其余四层为这两层服务。

数据来源:基于本文分析绘制的示意图。

科学发现方向在六层能力模型中的侧重点如下(该侧重分析基于公开案例事实,非标准):

侧重AI4S 方向的具体内容
L1 上下文工程层★★★★文献证据与领域知识、领域数据库、协议与数据字典、失败实验记录
L2 工具与执行层★★★★★模拟与实验 MCP(如暴露模拟启动与数据排序工具的领域服务器)、实验跟踪、机器人与仪器接口
L3 编排与控制层★★★★★lab-in-the-loop 架构:智能体做假说生成、实验设计与数据解读,人做问题定义、采纳判断与物理实验执行
L4 记忆与状态层★★★★假说库、文献引用库、实验记录与失败记录(可搜索的排错知识库)
L5 评估与观测层★★★★★ground truth = 物理定律 / 实验验证 / 守恒律;分层评估:廉价物理约束高频执行,昂贵实验验证按需触发
L6 治理与安全层★★★★★科研诚信(不得伪造数据与引用)、实验伦理审查、生物安全、模型可解释性、《科学数据管理办法》先汇交后验收

核心在 L5 与 L3。

L5 是第一性问题,因为本方向的一切价值最终都要通过「是否经得起物理定律与实验」来裁决。可验证的 ground truth 形态包括:

  • 物理约束:能量、质量、动量守恒;收敛性;与解析解或成熟物理模型(如气象领域的 ECMWF HRES / ENS)的对照。
  • 实验验证:独立实验室的合成确认(如 GNoME 预测材料的 736 种外部独立合成)、临床或统计显著性(如靶点评估中的生存分析)。
  • 可复现性:同一结果可被独立团队在不同装置上重算或重现(对应 ACM 四层可复现性定义)。

反例警示同样来自 L5 的缺位:Sakana 的 AI Scientist 被独立评估发现,7 篇手稿中有 4 篇将既有概念误判为新颖并产生幻觉的数值结果(据公开报道)。没有严格 L5,AI4S 的产出规模越大,幻觉存量越大。

L3 是核心,因为领域正在收敛到 lab-in-the-loop 架构:智能体负责假说生成、实验设计与数据解读,人类研究者负责提出问题、判断采纳哪些 AI 输出、执行物理实验。斯坦福虚拟 CSO 系统中「科学评审智能体指出配体-受体推断缺乏空间上下文,从而触发空间转录组学分析」的实例(详见 3.3 节),展示了智能体之间自我批评作为编排机制的价值。

瓶颈层:L5 的成本结构。 廉价验证(守恒律、收敛、与既有模型对照)可以自动化高频执行;昂贵验证(实验合成、湿实验、临床)无法高频进行且必须人在回路。如何设计「代理指标先行、真值校准定期」的分层评估体系,是本方向 Harness 落地的最大工程挑战。

1.4 价值与瓶颈

价值体现在三处:

  • 扩展假设空间:模型可在数亿候选(蛋白结构、晶体结构、药物-适应症配对)中系统性地提出假设,把人类从枚举劳动中解放出来。
  • 压缩认知劳动:FutureHouse 的 Robin 系统被机构称为可在一天内完成人类研究者原本需要数月的部分文献综述与假说生成工作;另有报道称其将发现周期的认知劳动从估计 872~937 人时压缩到 2 小时以内(两处数字口径不一,详见信息缺口声明)。
  • 把验证标准从口头变成自动:守恒律检查、收敛判定、引用可核查,这些原本依赖评审人素质的环节,可以被固化为每次产出必须通过的门禁。

瓶颈同样有三处:

  • 新颖性误判与数值幻觉:AI Scientist 的 4/7 问题率说明,生成能力越强,未经验证的产出越危险。
  • 分布外失效:AI 气象模型系统性低估破纪录极端事件,且极端程度越高差距越大(详见 3.2 节)——训练分布的尾部恰是科学上最有价值的地方。
  • 验证成本与信任断层:昂贵验证无法高频进行,导致「AI 提出、人类挑选」的环节成为瓶颈;GNoME 的「十倍」表述争议正是这一断层的体现(详见 3.3 节)。

2. 名词解释

术语英文 / 缩写释义
AI for ScienceAI4S以 AI 方法加速科学发现的研究范式,覆盖蛋白、材料、气象、数学、物理等领域
蛋白质结构预测Protein Structure Prediction从氨基酸序列预测蛋白质三维结构的任务,CASP 竞赛为其传统评测场
CASPCritical Assessment of protein Structure Prediction蛋白质结构预测的两年一届全球评测竞赛,AlphaFold2 于 CASP14 成名
GDTGlobal Distance Test蛋白质结构预测正确性的度量,满分 100;AlphaFold2 在 CASP14 中位 GDT 为 92.4
深度学习结构预测数据库AlphaFold DBDeepMind 发布的蛋白质结构数据库,覆盖 2 亿个以上蛋白质结构
结构 motifs 扩散生成RFdiffusion在 RoseTTAFold 基础上做去噪微调、从「读结构」变为「发明结构」的生成模型
气象大模型AI Weather Model以神经网络替代或补充数值天气预报的模型,代表包括盘古(Pangu-Weather)、GraphCast、GenCast、风乌、伏羲等
再分析数据Reanalysis, ERA5物理模型与数据同化(4D-Var)结合观测产出的历史大气数据集,是多数 AI 气象模型的训练数据来源
集合预报Ensemble Forecast以多成员扰动产出概率预报的方法,GenCast 以扩散模型产出集合预报
物理信息神经网络PINN, Physics-Informed Neural Network将物理方程作为约束项嵌入训练损失的神经网络方法,是 AI 与物理混合建模的前沿方向
守恒律Conservation Laws能量、质量、动量守恒等物理定律,是 AI4S 产出最核心的自动化验证依据
材料发现Materials Discovery搜索与设计具有目标性质的新材料,代表系统包括 GNoME 与 MatterGen
金属有机框架MOF, Metal-Organic Framework多孔晶体材料,常用于气体吸附等场景,是 AI 材料筛选的典型对象
巨正则蒙特卡洛GCMC, Grand-Canonical Monte Carlo模拟吸附等统计物理问题的计算方法
托卡马克Tokamak磁约束核聚变装置;2022 年起强化学习控制器成功运行真实等离子体
lab-in-the-loop实验室在回路智能体负责假说、设计与解读,人类负责问题定义、采纳判断与物理实验执行的协作架构
AI 科学家AI Scientist自主生成假说、运行实验、撰写论文并模拟评审的端到端系统,代表为 Sakana AI 的同名系统
新颖性误判Novelty Misjudgment将既有概念误判为新发现;独立评估发现 AI Scientist 的 7 篇手稿中 4 篇存在此问题
数值幻觉Numerical Hallucination模型生成未经执行或与执行结果不符的数值
同行评审Peer Review科学出版的质量审查机制;注意:会议研讨会(workshop)层面的评审不等同于主会录用
科学数据汇交Scientific Data Deposition依《科学数据管理办法》,政府预算资助项目形成的科学数据须先汇交、再验收的机制

3. 案例

3.1 案例一:蛋白质结构预测——从 CASP 到可编程设计

3.1.1 背景

蛋白质结构预测是 AI4S 的第一个「定理级」成功场景:任务定义清晰(从序列到结构)、评测体系成熟(CASP 竞赛与 GDT 度量)、ground truth 由实验测定(X 射线 / 冷冻电镜结构)。2018 年 AlphaFold 在 CASP13 首次展示深度学习在此任务的潜力。

3.1.2 方案

以下为 2020—2024 年的关键里程碑(多来源交叉一致,来源级别以 A/B 为主,见参考资料):

时间成果关键数据
2020-12AlphaFold2 在 CASP14中位 GDT 92.4 / 100;97 个目标中 88 个给出最佳预测;平均误差约 1.6 Å(约一个原子宽度)
2021RoseTTAFold(Baker 实验室)发表于 Science「三轨」神经网络(序列 / 距离 / 坐标同时处理);单台游戏电脑最快 10 分钟算出结构;开源
2022AlphaFold 蛋白质结构数据库发布覆盖 2 亿个以上蛋白质,几乎涵盖所有已知蛋白序列
2023-03ESMFold(Meta)发表于 Science无需序列比对;折叠 6.17 亿宏基因组蛋白质
2023-07RFdiffusion(Baker 实验室)发表于 Nature从「读结构」变为「发明结构」;数百个设计在实验室被制造出来
2023-09AlphaMissense 发表于 Science对人类蛋白质组全部 7,100 万可能错义变异打分,分类了其中 89%(此前仅约 0.1% 经实验室或临床确认)
2024-05AlphaFold 3(DeepMind 与 Isomorphic Labs)发表于 Nature从单体蛋白扩展到复合物(含 DNA、RNA、离子、小分子);相比既有方法,蛋白与其他分子类型的相互作用精度提升至少 50%

采纳规模方面,据公开报道,190 个国家的 200 万研究者使用过 AlphaFold 2 或 3(B 级来源)。

3.1.3 效果

本案例确立了一条对 Harness 设计最有参考价值的路径——评测先行、生成跟进、开源放大

  1. CASP 竞赛提供了物理级 ground truth:预测结果与实验测定结构直接比对,胜负无争议。这正是 L5 评估层的理想形态——先有可信判据,再有模型竞争,最后才有叙事收敛。
  2. 开源与数据库化放大了边际价值:RoseTTAFold 开源、AlphaFold DB 覆盖 2 亿结构,使下游研究者无需自行训练即可消费成果。对应到 Harness 语言:模型成果只有固化为可检索、可复用的工件(L4),才能转化为科学生产力
  3. 从预测到设计的跃迁依赖反馈闭环:RFdiffusion 的设计必须在实验室被制造并验证——生成(L2/L3)与实验验证(L5)的闭环才是完整的科学发现回路。

需要如实说明的边界:AlphaFold 3 的完整模型并未完全开源,学界对其复现与公平使用存在讨论;「精度提升至少 50%」为论文对照既有方法的结论,适用范围限于其评测的相互作用类型。这些边界不影响上述结论,但引用时应予保留。

3.2 案例二:AI 气象大模型——超越数值预报与系统性低估极端事件

3.2.1 背景

中期天气预报是 AI4S 中「AI 第一次在常规任务上挑战物理学家写了一个世纪的方程」的领域。2023 年起,盘古(Pangu-Weather)、GraphCast、FourCastNet、伏羲(FuXi)、风乌等模型相继发表,在精度与速度上被报道超过业务数值预报。

3.2.2 方案

关键事实编年(多来源交叉一致):

时间成果关键数据
2023-07盘古 Pangu-Weather(华为云)发表于 Nature基于 39 年大气记录训练;首个被报道在精度与速度上双双超过业务数值预报的机器学习模型;单 GPU 全球预报约 1.4 秒;速度较传统数值方法快 10,000 倍以上
2023-11GraphCast(DeepMind)发表于 Science0.25° 分辨率、10 天预报;在 1,380 个验证目标中的 90.3% 上超越 ECMWF HRES;单 TPU 不到 1 分钟
2023伏羲 FuXi(复旦);FourCastNet(NVIDIA)FuXi 以级联模型实现 15 天预报,精度首次达到 ECMWF 15 天集合预报的集合平均;FourCastNet 首个 0.25° 分辨率模型
2023-04 起风乌(上海人工智能实验室)首次实现超过 10 天有效预报;FengWu-GHR 分辨率提升至 0.09°(约 9 km),有效预报时长由 10.75 天提升至 11.25 天(传统模式每 10 年才提高约 1 天)
2024-12GenCast(DeepMind)发表于 Nature扩散模型产出概率集合预报;在 1,320 个变量 × 预见期组合中 97.2% 优于 ECMWF ENS,预见期大于 36 小时时达 99.8%;单 TPU v5 约 8 分钟产出 15 天预报
2025-02ECMWF 自有业务模型 AIFS Single 投入业务运行2025-07-01 集合版投入业务
2025-10中国气象局评估以「伏羲」「璞云」「风乌」「风清」「盘古」为代表的第一梯队形成,关键环流要素可用预报天数达 10 天;同时指出 AI 模型更适合作为传统数值预报的有效补充
2026-05第二届人工智能气象预报模型示范计划启动建立覆盖临近、短中期至次季节全时间尺度的标准化评估体系;发布预报性能检验评估等气象行业标准,为模型业务准入提供依据

批判性内容(必须与成功叙事并列呈现):据 Science Advances 相关研究,主要 AI 气象模型系统性低估破纪录的高温、寒潮和大风事件,且极端程度越高,差距越大;其中 AIFS 尤其低估风暴峰值风速,原因是其 MSE 训练目标倾向于平滑产生尖锐峰值的大气梯度。结论是:物理模型在分布的尾部——极端事件所在之处——仍然占优

更深层的结构性依赖同样必须声明:AI 气象模型依赖 ERA5 再分析数据训练与推理,而 ERA5 本质上是物理模型 + 数据同化(4D-Var)的输出。若传统物理模型与同化管线被弃用,AI 模型将失去训练与日常推理所需的严格结构化初始条件。纯 AI 模型本质上是高度复杂的模式匹配插值器,缺乏对守恒律的结构性理解,被推出训练分布边界时理论上可产生物理上不可能的大气状态。前沿方向是 PINN 与可微求解器 + 机器学习参数化方案的深度混合模型。

3.2.3 效果

本案例给 AI4S(乃至所有数据驱动科学方向)三条结构化结论:

  1. 基准超越不等于任务超越:GraphCast 在 90.3% 的验证目标上超越 HRES 是事实;但极端事件恰恰落在剩余的尾部,而尾部正是预报价值最高的地方。「90% 超越」与「系统性低估极端」是同一组事实的两面,必须同时陈述。
  2. 训练数据依赖决定了能力天花板:ERA5 依赖意味着 AI 气象模型在方法论上不是物理模型的替代者,而是消费者。中国气象局「有效补充」的定位、以及以检验评估标准作为业务准入门槛的做法,是 Harness L5 思维在行业治理层的直接体现。
  3. 行业标准化的出现标志着方向成熟:2026 年气象行业标准把「检验评估」变成准入条件,等价于把 L5 从建议变成了强制门禁——这是所有 AI4S 子领域可以参照的治理路径。

3.3 案例三:自主科学发现系统的成色检验——成功、争议与 lab-in-the-loop

3.3.1 背景

2024—2026 年,「AI 科学家」类系统进入密集发布期:Sakana AI 的 The AI Scientist、FutureHouse 的 Robin 与 Kosmos、Google 的 co-scientist、斯坦福虚拟 CSO、ChemAgents、Lila Sciences 等。该领域的宣传叙事与真实成色之间存在显著落差,独立评估与学界批评提供了宝贵的校准数据。

3.3.2 方案

先列事实,再列质疑。

正面事实

  • 斯坦福虚拟 CSO 的 B7-H3 案例展示了多智能体协作与自我批评的完整闭环:评估 B7-H3 作为肺癌治疗靶点时,统计遗传学智能体未发现显著种系关联但标记出调控活性;单细胞图谱智能体识别出 B7-H3 在癌症相关成纤维细胞中富集;科学评审智能体指出配体-受体推断缺乏空间上下文——该自我批评触发后续空间转录组学分析,确认 B7-H3 高表达区域周围存在免疫排斥微环境;临床试验智能体在肺腺癌基因组图谱上运行分层生存分析(疾病特异性生存 HR = 1.82,p = 0.031);最终推荐抗体偶联药物(ADC)模态,并在药理学家智能体未找到可成药口袋后下调小分子与 PROTAC 的优先级。
  • 数学领域:AlphaGeometry2 解出 84%(42/50)的 IMO 几何题达金牌水平;2026-02 DeepMind 的 Aletheia 自主解决 700 个 Erdős 猜想中的 4 个开放问题并生成完整研究论文。
  • 物理控制领域:2022-02 DeepMind 与 EPFL 的强化学习控制器在真实托卡马克上以 90 路测量输入、19 路电压指令、10 kHz 频率运行,维持了需要专家数周手工调优才能达到的等离子体位形,为史上首次。

独立质疑(必须与成功案例同等篇幅呈现)

  1. AI Scientist 的新颖性误判与数值幻觉:Sakana 的 AI Scientist 被独立评估发现,在 7 篇手稿中的 4 篇里将既有概念误判为新颖,并产生幻觉的数值结果。其改进版生成的一篇论文曾在 ICLR 2025 的一个研讨会(workshop)中达到同行评审接收阈值——这是会议研讨会层面,不等同于顶会主会录用,引用时必须注明。
  2. GNoME 的「十倍」表述争议:GNoME 发现 220 万个新晶体结构、其中 38.1 万判定为稳定,外部实验室独立合成确认了 736 种。但「AI 将有用材料目录扩充十倍」的隐含主张在精神上(若非形式上)已被撤回;围绕 GNoME 的诺奖式兴奋已沉淀为更克制的叙事:AI 擅长提出候选,而人类化学家仍要从中挑选值得去做的少数
  3. Robin 的 ripasudil 质疑:FutureHouse 的 Robin 提出将青光眼药物 ripasudil 用于干性年龄相关性黄斑变性,机制为 ROCK 抑制——但受到质疑:Robin 贡献的是「药物-适应症配对」,而非底层机制本身;这是有意义的贡献,但比最初表述所暗示的更为有限。
  4. 人类仍在复杂科学任务上领先:另有近期 Nature 研究指出,人类在复杂科学任务上仍优于最好的 AI 智能体,但差距正在缩小(该转述来源可信度有限,标注 )。
  5. 缺乏共享基准Nature 的一篇系统选型指南指出,该领域仍缺乏共享基准来区分真实进展与营销,且「AI scientist」标签因厂商而异、含义差异极大(转述来源,建议核对原文)。

3.3.3 效果

正反事实并置之后,可以提炼出三条对本方向 Harness 设计最核心的结论:

  1. lab-in-the-loop 是被正反两方共同验证的架构。斯坦福案例证明智能体间自我批评(评审智能体触发空间分析)可以提升结论质量;AI Scientist 与 Robin 的质疑则证明,缺少外部验证的端到端自主产出的贡献密度有限。收敛结论:智能体负责假说、设计与解读,人类负责问题定义、采纳判断与物理实验——这正是 L3 编排层在本方向的标准形态。
  2. 数值必须由执行产生,新颖性必须经检索比对。AI Scientist 的两类失败(数值幻觉、新颖性误判)分别对应两条可工程化的门禁:数值只能来自模拟 / 代码执行结果;「首次发现」类主张必须先过文献检索比对。这两条已写入 4.1 节红线。
  3. 贡献叙事要与验证等级挂钩。GNoME 的 38.1 万稳定候选是计算级结论,736 种独立合成才是实验级结论,两者相差近三个数量级。Harness 的输出规范应强制区分「计算预测 / 实验确认 / 独立复现」三档证据等级,禁止混报。

4. 实践标准

4.1 AGENTS.md 规范

标准来源声明:以下为本文提出的科学发现方向(AI4S)AGENTS.md 标准建议稿。截至目前,不存在由官方机构、行业协会或标准组织发布的 AI4S 方向 AGENTS.md 规范原文,AGENTS.md 属社区约定而非标准。本建议稿继承数据科学组级 AGENTS.md 全部条款,并针对 AI4S 方向收紧与扩展。

# AGENTS.md —— 科学发现(AI for Science)

> 继承数据科学组级 AGENTS.md 全部条款。本节为科学发现方向的收紧与扩展。
> 本文件为标准建议稿,业界尚无官方标准版本。

## 角色与边界

- 本 Agent 是**科学假说的起草者与验证的组织者**:可以综述文献、提出假说、设计模拟与实验、执行计算验证、解读结果。
- 不可以直接执行的操作:执行物理 / 化学 / 生物实验、进行临床或伦理裁决、单方面宣布「首次发现」、替人类决定采纳哪个假说。
- **ground truth 不是人类标注,而是物理定律、实验验证与守恒律**。任何与物理约束冲突的计算结果,默认为错误,除非有明确证据说明约束条件差异。
- 问题的定义权、假说的采纳权与实验的执行权归领域科学家,不归本 Agent。

## 环境假设

- 领域工具可用:模拟引擎(分子模拟、CFD、气象模型等)、领域数据库、文献检索接口。
- 模拟与实验工具经工具契约暴露(如 MCP 服务器);工具不直接承载计算,而是产出可由工作流引擎 / 调度系统承载的作业描述。
- 验证判据可用:守恒律容差、收敛阈值、基准算例、(如适用)实验对照协议。
- 文献库与引用数据库可检索,且检索结果带可点击来源。
- 失败实验记录库可用(可为空集,但需有沉淀路径)。
- **任一验证判据缺失时,先与领域科学家共同定义判据,再开始计算。无判据的模拟产出一律标记为未验证。**

## 上下文加载顺序(Context Budget)

1. 科学问题与验收判据(必须):假说、目标性质、判据类型与容差
2. 领域约束:守恒律、边界条件、适用范围(必须)
3. 文献证据:与假说直接相关的既有工作与引用(**新颖性判断前必须**)
4. 领域数据库条目:结构、性质、已验证记录
5. 同类既往模拟:参数、判据、结果与失败记录
6. 计算资源约束:精度、规模、时长与成本上限
7. 实验可行性与安全约束:涉及湿实验 / 生物 / 临床时的伦理与安全前提

- 文献证据必须携带可点击来源;无来源的领域知识不得作为新颖性或正确性的依据。

## 工具契约

| 工具 | 用途 | 模式 | 约束 |
|---|---|---|---|
| 文献检索 | 查询既有工作、引用核查 | 只读 | 结果必须带来源;新颖性比对结果留痕 |
| 领域数据库 | 查询结构、性质、既有记录 | 只读 | 不得修改公共数据库 |
| 模拟工具接口 | 发起模拟任务描述 | 受限写 | emit 作业描述给工作流引擎;不直接操控调度器 |
| 数值与符号计算 | 计算、推导、统计分析 | 受限执行 | 数值必须由执行产生并记录输入输出 |
| 验证判据执行器 | 守恒律检查、收敛判定、基准对照 | 自动 | 每次模拟产出必经;结果不可被跳过 |
| 实验记录本 | 读写实验设计与结果记录 | 受限写 | 实验执行状态由人更新,Agent 不得代填 |

- 工具参数必须做模式校验;格式错误在调用前拦截。
- 每次模拟记录:判据、参数、版本、运行时长、守恒律偏差与收敛残差。

## 任务执行流程(SOP)

1. **问题定义**:与科学家对齐科学问题、目标性质与验收判据;判据缺失则先补判据。
2. **文献综述**:检索并汇总既有工作,标注来源;识别可借鉴方法与已知失败。
3. **假说生成**:给出假说及其可检验推论;每个假说附验证判据草案。
4. **新颖性比对**:对「首次 / 新发现」类主张执行文献检索比对;比对结果留痕,未过比对不得声称新颖。
5. **计算设计**:选择方法与参数,声明适用范围与已知的近似。
6. **执行与物理验证**:经工作流引擎执行模拟;自动执行守恒律检查、收敛判定与基准对照。
7. **结果解读**:区分「计算预测 / 实验确认 / 独立复现」三档证据等级;列出与预期不符的反例。
8. **实验设计移交**:需要湿实验 / 物理 / 临床验证时,输出实验设计草案,交由人类执行;Agent 不得代执行。
9. **结论分级与移交评审**:结论按证据等级分级输出;所有引用可点击核查;进入人工评审与(如适用)同行评审。
10. **归档**:假说、判据、参数、结果、失败记录与文献比对记录归档;失败原因写入失败记录库。

## 验证与证据要求

- **每条数值必须由执行产生**:来自模拟、代码或统计计算的实际输出,禁止由模型凭记忆或估算生成。
- **物理约束优先于拟合优度**:结果违反守恒律或产生物理上不可能的状态时,判定为无效,不得以「整体误差小」辩护。
- **证据等级三档制**:计算预测(未经实验)→ 实验确认 → 独立复现;输出中必须显式标注当前档位。
- **新颖性主张必须先过文献比对**:比对关键词、命中结果与差异说明留痕。
- **引用必须可点击且逐条可核查**:无来源的引用禁止出现在任何产出物中。
- 主动报告反例:与假说不符的结果必须完整呈现,禁止只报告支持性结果。
- 规模结论与已验证结论分开陈述:如「候选 38 万」与「已合成确认 736」是两句话,不是一句话。

## 失败与升级策略

- 同类失败重试不超过 2 次;第 3 次改变策略或升级。
- **物理约束检查失败**:默认结果无效;先排查边界条件与参数,再考虑方法适用性;不得放宽判据。
- **文献比对发现已有工作**:更新假说定位,如实引用既有工作,禁止重复包装。
- **模拟与既有基准冲突**:报告冲突并升级给科学家,禁止选择性地忽略不利的基准结果。
- **工具参数格式错误**:停止重试,修正调用构造;多次出现上报工具契约缺陷。
- **涉及生物安全 / 伦理敏感的实验设计**:停止生成细节,升级给伦理与安全负责人。
- **成本或时长超预算**:暂停并报告,获批后继续。
- 升级时携带:科学问题、假说、判据、已执行验证、反例清单、建议下一步。

## 安全与合规红线

- 不得伪造、篡改或选择性报告数据。
- 不得伪造引用、编造文献或夸大既有工作的支持程度。
- 不得在未过文献比对的情况下使用「首次」「首创」类表述。
- 不得代替人类执行或声称已执行物理、化学、生物实验。
- 涉及病原体、危险化学品、临床与动物实验的设计,须先通过伦理与安全审查。
- 依《科学数据管理办法》,政府预算资助项目形成的科学数据须先汇交、再验收;不得私匿应汇交数据。
- 不得绕过同行评审流程直接对外发布科学结论。

## 禁止事项

- 禁止生成未经验证的数值结论。
- 禁止把计算预测表述为实验确认;禁止把单点确认表述为普遍规律。
- 禁止把会议研讨会层面的接收表述为主会录用或「顶会发表」。
- 禁止跳过新颖性比对直接宣称突破。
- 禁止在判据未定义时开始大规模模拟。
- 禁止删除或淡化失败记录。
- 禁止跨方向复制通用模板;AI4S 的验证判据与治理要求与数据分析、HPC 有实质差异。

## 输出格式

- 结论先行 → 证据等级(计算预测 / 实验确认 / 独立复现)→ 判据执行记录 → 反例与局限 → 建议。
- 数值带单位与判据上下文;范围用「~」连接;百分比数值与 % 之间无空格。
- 文献引用带可点击来源;「首次发现」类表述必须附比对记录编号。
- 候选规模与已验证数量分开陈述,用表格列出:证据等级、数量、验证方式、验证方。
- 失败记录使用固定格式:假说、判据、失败现象、根因分析、可复用教训。

## 评估与自检

- [ ] 验收判据在计算开始前已定义并双方确认
- [ ] 全部数值来自实际执行,输入输出留痕
- [ ] 守恒律检查与收敛判定已自动执行且通过,或失败已如实报告
- [ ] 新颖性主张已过文献比对,比对记录留痕
- [ ] 全部引用可点击且逐条核查
- [ ] 证据等级三档制已执行,候选与已验证分开陈述
- [ ] 反例与不利结果已完整呈现
- [ ] 需实验验证的部分已移交人类,未代执行
- [ ] 伦理与安全约束已确认(如适用)
- [ ] 假说、判据、结果与失败记录已归档

4.2 SKILL.md 规范

标准来源声明:以下为本文提出的科学发现方向 SKILL.md 标准建议稿,同样不存在官方标准原文。

---
name: ai4s-hypothesis-validation
description: 科学假说从生成到计算验证与实验移交的标准执行流程。适用于文献综述与新颖性比对、假说生成、模拟验证设计、守恒律与收敛性检查、证据分级与实验设计移交等任务。触发场景:任何以「提出并验证科学假设」为目标的智能体任务。
version: 1.0
created: 2026-09-12
---

# 科学假说验证标准流程

## 适用场景

- 文献综述与新颖性比对(某方向已知什么、还有什么空隙)。
- 科学假说的生成、形式化与可检验推论设计。
- 计算验证:模拟参数设计、守恒律与收敛性检查、基准对照。
- 实验设计草案编制与移交(湿实验 / 物理 / 临床)。
- 不适用场景:实验的实际执行与裁决、临床决策、伦理与安全审批。

## 前置条件

- 科学问题与目标性质已与领域科学家对齐。
- 验证判据(守恒律容差、收敛阈值、基准算例)已定义;未定义则先补。
- 文献检索接口与领域数据库可用且结果带来源。
- 模拟工具经工具契约可用;工作流引擎与资源预算已确认。
- 失败记录库可用(可为空集)。

## 输入

| 输入项 | 必需 | 说明 |
|---|---|---|
| 科学问题与目标性质 | 是 | 待发现或待验证的对象 |
| 验收判据 | 是 | 守恒律容差、收敛阈值、基准与对照 |
| 领域约束与适用范围 | 是 | 物理边界条件、方法近似声明 |
| 资源预算 | 是 | 计算规模、时长与成本上限 |
| 既有工作与失败记录 | 否 | 文献综述初稿、既往模拟记录 |

## 输出

| 输出项 | 必需 | 说明 |
|---|---|---|
| 假说书 | 是 | 假说、可检验推论、验证判据草案 |
| 新颖性比对记录 | 是 | 检索关键词、命中结果、差异说明 |
| 计算验证报告 | 是 | 参数、版本、守恒律偏差、收敛残差、基准对照 |
| 证据分级结论 | 是 | 计算预测 / 实验确认 / 独立复现三档 |
| 反例与局限 | 是 | 与预期不符的结果及解释 |
| 实验设计草案 | 否 | 交由人类执行的实验方案与安全前提 |

## 执行步骤

1. **对齐问题与判据**
   与科学家确认科学问题、目标性质与验收判据;判据缺失时先共同定义,禁止无判据开算。

2. **文献综述**
   检索既有工作并标注来源;汇总已知方法、已知结论与已知失败;形成新颖性空隙假设。

3. **假说生成**
   给出假说及其可检验推论;每个假说附验证判据草案与证伪条件。

4. **新颖性比对**
   对「首次 / 新发现」类主张执行文献检索比对:记录关键词、命中结果、差异说明;未过比对的主张降级为「改进 / 应用」。

5. **计算设计**
   选择方法与参数;声明近似与适用范围;设计对照算例与负对照;估算资源与时长。

6. **执行与自动验证**
   经工作流引擎提交模拟;自动执行守恒律检查、收敛判定与基准对照;检查不通过则结果无效并进入失败分析。

7. **结果解读与分级**
   区分计算预测 / 实验确认 / 独立复现三档;规模结论与已验证结论分开陈述;列出全部反例。

8. **实验设计移交**
   需要实验验证时输出实验设计草案,含安全与伦理前提;执行权归人类,Agent 不得代执行。

9. **人工评审**
   结论连同判据记录、比对记录与反例一并提交评审;评审意见回流入假说书。

10. **归档与失败沉淀**
    假说书、验证报告、比对记录、反例与失败原因归档;失败教训写入失败记录库。

## 质量标准(DoD)

判据与验证:

- [ ] 判据在计算开始前定义并经科学家确认
- [ ] 守恒律检查与收敛判定已自动执行并留痕
- [ ] 全部数值来自实际执行,禁止模型生成数值
- [ ] 与基准 / 解析解 / 既有模型的对照已完成

证据与引用:

- [ ] 新颖性主张已过文献比对,记录留痕
- [ ] 全部引用可点击且逐条核查
- [ ] 证据等级三档制已执行:计算预测 / 实验确认 / 独立复现
- [ ] 规模结论与已验证数量分开陈述

诚信与治理:

- [ ] 反例与不利结果完整呈现
- [ ] 实验设计已移交,未代执行任何实验
- [ ] 伦理与安全前提已确认(如适用)
- [ ] 假说、判据、结果与失败记录已归档

## 常见失败与处理

| 失败现象 | 根因 | 处理方式 |
|---|---|---|
| 结果违反守恒律 | 边界条件错误 / 方法不适用 / 时间步长过大 | 判定无效;排查条件与方法;禁止以整体误差小辩护 |
| 收敛残差不达标 | 判据过严或数值方法限制 | 报告残差与判据差距;升级给科学家;禁止放宽判据求完成 |
| 文献比对发现已有相同工作 | 综述不足 | 更新假说定位为改进 / 应用;如实引用既有工作 |
| 数值「过于完美」 | 数值由模型生成而非执行 | 回溯输入输出记录;补执行;报告诚信事件 |
| 候选规模被误读为已验证数量 | 证据等级未分开陈述 | 重写结论:计算预测 N 项、实验确认 M 项,分开成句 |
| 假说无法证伪 | 判据缺失或表述含糊 | 回到判据定义步骤;补证伪条件后再开算 |
| 实验团队拒绝移交方案 | 安全 / 伦理前提缺失 | 补齐安全与伦理评估;升级给负责人 |
| 模拟与既有基准冲突 | 方法近似或基准适用范围不同 | 如实报告冲突与两种解释;升级裁决 |

## 示例

**任务**:为某类多孔材料筛选吸附性能候选,并评估新颖性。

1. 对齐判据:目标性质为吸附容量;判据为能量守恒偏差小于 1e-6、与已发表基准材料偏差方向一致、统计显著性 p 小于 0.05。
2. 文献综述:检索既有筛选工作 24 篇,汇总已知方法与最佳已报道吸附容量。
3. 假说生成:提出「孔径分布与吸附容量存在非线性关系」的假说,附证伪条件。
4. 新颖性比对:以三个关键词组检索,命中 6 篇相近工作;差异说明留痕,主张降级为「在特定子族中的系统验证」。
5. 计算设计:蒙特卡洛模拟,声明力场近似;2,000 候选,预算内分批。
6. 执行与自动验证:经工作流引擎分批执行;守恒律检查全部通过;2 批收敛失败,进入失败分析。
7. 结果解读:计算预测前 20 候选(证据等级:计算预测);与既有实验对照的 40 个已知材料偏差在判据内。
8. 实验设计移交:输出前 5 候选的合成与测定草案,含安全前提;执行归人类。
9. 人工评审:结论、判据记录、比对记录与反例一并提交。
10. 归档:全部工件归档;2 批收敛失败的原因写入失败记录库。

4.3 落地检查清单

4.3.1 上下文层(L1)

  • [ ] 文献证据库可检索,结果带可点击来源
  • [ ] 领域数据库条目(结构、性质、已验证记录)纳入上下文
  • [ ] 验收判据模板(守恒律容差、收敛阈值、基准算例)可用
  • [ ] 失败实验记录可被检索并注入上下文
  • [ ] 方法的近似与适用范围随上下文显式声明

4.3.2 工具与执行层(L2)

  • [ ] 模拟与实验规划工具经工具契约暴露,emit 作业描述而非直接操控调度器
  • [ ] 工具参数有模式校验,畸形参数在调用前拦截
  • [ ] 数值与符号计算有输入输出留痕
  • [ ] 每次模拟记录判据、参数、版本、时长与约束偏差

4.3.3 编排与控制层(L3)

  • [ ] lab-in-the-loop 架构落地:假说采纳与实验执行由人决策
  • [ ] 智能体间自我批评机制可用(评审智能体对结论提出质疑并触发补充分析)
  • [ ] 新颖性比对是生成「首次发现」类主张前的强制卡点
  • [ ] 实验设计移交有明确的状态流转(草案 → 人类执行 → 结果回流)

4.3.4 记忆与状态层(L4)

  • [ ] 假说库与文献引用库可持续积累
  • [ ] 失败记录库已建立,失败原因与教训可检索
  • [ ] 模拟参数与版本可追溯、可复算
  • [ ] 评审意见与新版本假说的对应关系留痕

4.3.5 评估与观测层(L5)

  • [ ] 物理级自动验证(守恒律、收敛、基准对照)为每次模拟产出的强制门禁
  • [ ] 证据等级三档制(计算预测 / 实验确认 / 独立复现)在全部输出中执行
  • [ ] 分层评估:廉价物理约束高频执行,昂贵实验验证人在回路、按需触发
  • [ ] 规模结论与已验证数量分开陈述
  • [ ] 反例与不利结果有完整呈现渠道

4.3.6 治理与安全层(L6)

  • [ ] 数值必须来自执行、引用必须可核查,两条诚信红线有技术强制
  • [ ] 「首次」类表述有比对记录编号约束
  • [ ] 涉及生物安全、危险化学品、临床与动物实验的设计有伦理与安全审查卡点
  • [ ] 科学数据汇交义务(先汇交、再验收)已纳入流程
  • [ ] 未验证结论不得对外发布,同行评审流程不可绕过

5. 总结

科学发现方向对 AI Harness 的核心诉求可以概括为一句话:ground truth 不是人类标注,而是物理定律、实验验证与守恒律约束——L5 因此成为第一性问题。

本方向的证据链呈现出罕见的对称性:

  1. 成功的案例都站在强验证之上。AlphaFold2 的说服力来自 CASP 的实验测定结构比对;GraphCast / GenCast 的说服力来自与 ECMWF HRES / ENS 的对照;斯坦福虚拟 CSO 的说服力来自评审智能体触发的空间分析与统计显著性。凡是站得住的 AI4S 成果,背后都有一套可执行、可复现的验证体系。
  2. 翻车的案例都缺了同一段验证。AI Scientist 的 7 篇手稿中 4 篇新颖性误判与数值幻觉;GNoME 的「十倍」表述争议(38.1 万稳定候选对 736 种独立合成,相差近三个数量级);Robin 的 ripasudil 被指出贡献的是「药物-适应症配对」而非底层机制;AI 气象模型系统性低估极端事件且极端程度越高差距越大——失败形式各异,根因同一:生成能力超过了验证能力
  3. 人类仍在复杂任务上领先,但行业已找到正确的协作形态:lab-in-the-loop。智能体做假说、设计与解读,人做问题定义、采纳判断与物理实验。这既是对「人类 vs AI」叙事的校准,也是 L3 编排层在本方向的标准答案。

因此本方向的重心落在 L5 与 L3:L5 把物理定律与实验验证固化为自动门禁与证据分级制度(计算预测 / 实验确认 / 独立复现三档,规模结论与已验证结论分开陈述);L3 把人机分工固化为 lab-in-the-loop 的编排机制。至于其余四层,都是为这两层服务的。

一句话概括本方向的主张:让每一个科学主张都带上证据等级,让每一条物理约束都成为自动门禁,让每一次「首次发现」都先过文献比对,让每一次实验都留在人类手中。

信息缺口声明

  1. 不存在科学发现方向 AGENTS.md / SKILL.md 的官方或行业公认标准原文。4.1 与 4.2 节均为本文提出的标准建议稿。
  2. 「AI Harness × 数据科学」及 AI4S 的权威定义、市场规模、采用率与专项资助金额、立项数量——无可靠公开来源,本文未给出任何此类数字。
  3. **「人类在复杂科学任务上仍优于最好的 AI 智能体」(Nature 研究)**来自二手转述,标注 ,建议核对原文后引用。
  4. ***Nature*「Which AI scientist suits your lab?」选型指南的内容**(缺乏共享基准、「AI scientist」标签含义差异极大)来自二手转述,标注 。
  5. Robin 系统的两处效率数字口径不一:「一天内完成数月工作」与「872~937 人时压缩到 2 小时以内」来自不同转述,本文仅并列说明、未采用任何一处作为结论。
  6. AlphaFold 使用者规模(190 国 / 200 万研究者)托卡马克 RL 控制细节(90 路 / 19 路 / 10 kHz)Aletheia 解决 4 个 Erdős 开放问题等数字来自多来源交叉的转述(B 级),引用时建议核对一手论文。
  7. AI 气象模型低估极端事件的结论来自 Science Advances 相关研究的转述,方向性结论与学界共识一致,具体数字未引用;建议核对原文。
  8. 风乌、伏羲等国产气象大模型的部分性能数字(如有效预报 11.25 天、0.09° 分辨率)来自厂商与媒体转述(B 级),引用时建议核对论文或官方发布。
  9. 科学智能物质创制中心「投资超百亿元」及承载 AI4S 重大专项的表述来自二手转述,本文件未采用;中国 AI4S 专项的具体资助金额与立项数未检索到官方公开明细。
  10. 4.2 节示例中的材料体系、判据阈值与候选数量均为示意性构造,不代表任何真实研究。

6. 参考资料

  1. AI for Science 成果编年与关键数据 — AIWiki。https://aiwiki.ai/wiki/ai_for_science
  2. AI for Science 能力与里程碑汇编 — Achievements AI。https://achievements.ai/type/capability
  3. AI for Science 知识树(AlphaFold / GNoME / Aletheia 等条目)— 智源社区。https://hub.baai.org.cn/knowledge-tree/8fd127c3-5f22-493d-8fa2-d23da3c3ace7
  4. AI in Physics 综述页(AlphaGeometry / 托卡马克 RL 控制 / GNoME 争议等)— Scale Physics。https://scalephysics.com/horizons/aiinphysics/
  5. AI 气象模型低估极端事件相关研究(Science Advances / Earth & Environment 系)— Nature。https://www.nature.com/articles/s43247-025-02502-y
  6. Machine Learning Weather Forecasting 分析页(ERA5 依赖与混合建模方向)— Research。https://research.mental-momentum.ai/r/machine-learning-weather-forecasting-ijl0v7
  7. AI 天气预报词条(风清 / 风雷 / 风顺 / 第一梯队 / 示范计划)— 百度百科。https://baike.baidu.com/item/AI天气预报/68593824
  8. Agents for R&D Science(AI Scientist 4/7 评估、Robin ripasudil 质疑、斯坦福虚拟 CSO B7-H3 案例)— Orchestra Bio。https://orchestra.bio/blog/agents-for-r-d-science
  9. 自主科学发现系统 2026 盘点(Co-Scientist / Kosmos / 融资与工具生态)— Mixflow。https://mixflow.ai/blog/the-ai-pulse-whats-new-in-autonomous-scientific-discovery-for-2026
  10. Nature「Which AI scientist suits your lab?」转述页(领域缺共享基准、自主层级选型)— Sinapti。https://sinapti.ca/post/es/nature-compara-los-sistemas-de-ia-que-automatizan-la-investi-xd43nve2
  11. 上海人工智能实验室司南科学智能评测体系(SciEvalKit 七大能力 × 六学科)— 上海人工智能实验室。https://www.shlab.org.cn/news/5444233
  12. SciEvalKit 官方仓库 — InternScience(GitHub)。https://github.com/InternScience/SciEvalKit
  13. Awesome AI for Science(科学智能体基准全景:SciCode、MLAgentBench、NewtonBench 等)— GitHub。https://github.com/ai-boost/awesome-ai-for-science
  14. 「人工智能驱动的科学研究」专项部署相关报道 — 中国记协网。https://union.china.com.cn/cmdt/txt/2024-03/26/content_42737343.html
  15. Agentic MOF Screening on Aurora(AI4S 与 HPC 交叉案例)— supercomputing.news。https://www.supercomputing.news/hpc/agentic-mof-screening-aurora

AI4S · Scientific Discovery

1. Introduction

1.1 Background

AI for Science (AI4S, AI-driven scientific research) is the direction with the highest scientific return on AI agent technology, and also the one most easily overhyped by narratives. On one hand, protein structure prediction since 2018, and weather foundation models and materials discovery since 2023, mark the first time AI has reached or surpassed the level of traditional methods on several core scientific tasks; on the other hand, autonomous "AI Scientist" systems have been independently evaluated and found to exhibit novelty misjudgment and numerical hallucination, and AI weather models have been shown to systematically underestimate extreme events—an AI4S without a strict validation layer is an amplifier of hallucination.

Policy is likewise accelerating. In 2023-03, the Ministry of Science and Technology and the National Natural Science Foundation of China launched the "AI-driven scientific research" special deployment, focusing on key areas such as drug R&D, gene research, biological breeding, and new materials R&D, advancing algorithm-model innovation, open sharing of scientific research data, and computing-platform construction, while supporting the heterogeneous integrated development of high-performance computing centers and intelligent computing centers. In 2026-05, the China Meteorological Administration launched the second AI weather forecast model demonstration program and released meteorological industry standards such as the Verification and Evaluation of Forecast Performance of AI Weather Forecast Models"verification and evaluation" has become a precondition for operational admission, which is highly isomorphic with the role of L5 in the six-layer Harness model.

This direction has a fundamental difference from the other five directions that is worth declaring up front: here the ground truth is not a human-annotated dataset, but physical laws, experimental verification, and conservation-law constraints. A SQL query can be compared, a review's citations can be checked, but a crystal structure proposed by AI can only be adjudicated by physical laws and experimental synthesis.

1.2 Definition

The AI Harness for the scientific discovery direction refers to the engineering support layer that, around hypothesis generation, literature-evidence assembly, simulation and experiment planning, result verification, and conclusion grading, provides agents with domain-context assembly, simulation and experiment tool contracts, lab-in-the-loop orchestration, hypothesis and failure-record memory, and an evaluation basis centered on physical laws and experimental verification.

It does not replace the judgment of domain scientists, does not replace experiments themselves, and does not replace peer review. It carries the layer that "constrains each of the agent's scientific claims with verifiable evidence."

Three boundary points need to be clarified:

  • The AI4S Harness is not a "machine that automatically produces papers." Independent evaluations of autonomous scientific discovery systems have shown that unvalidated automatic output carries errors in a respectable form (see section 3.3).
  • The AI4S Harness does not adjudicate experiments. The execution authority for physical experiments and clinical validation rests with humans; agents are responsible for hypotheses, design, and interpretation.
  • The AI4S Harness has a distinctive evaluation-cost structure: cheap verification (conservation laws, convergence, comparison against existing models) can be performed at high frequency, while expensive verification (experimental synthesis, clinical trials) must be human-in-the-loop and triggered on demand.

1.3 Positioning in the AI Harness System — L5 Is the First-Order Problem

图 1-1|AI4S 六层能力模型:L5 第一性问题与 L3 lab-in-the-loop 编排

AI4S 六层能力模型:L5 第一性问题,L3 lab-in-the-loop 侧重评级基于正文六层能力模型分析 · 示意:基于本文分析绘制 L6 治理与安全层 侧重 ★★★★★ 科研诚信 · 伦理审查 · 生物安全 · 数据汇交 L5 评估与观测层 · 第一性问题(本图重点) 侧重 ★★★★★ ground truth = 物理定律 · 实验 · 守恒律 | 廉价高频 · 昂贵按需 L4 记忆与状态层 侧重 ★★★★ 假说库 · 引用库 · 实验与失败记录 L3 编排与控制层 · lab-in-the-loop(本图重点) 侧重 ★★★★★ 智能体做假说 · 设计,人做定义 · 执行 L2 工具与执行层 侧重 ★★★★★ 模拟实验 MCP · 实验跟踪 · 仪器接口 L1 上下文工程层 侧重 ★★★★ 文献证据 · 领域数据库 · 失败实验记录 结构解读:AI4S 的核心在 L5(ground truth 是物理定律/实验/守恒律)与 L3(lab-in-the-loop),其余四层为这两层服务。

数据来源:基于本文分析绘制的示意图。

The emphasis of the scientific discovery direction in the six-layer capability model is as follows (this emphasis analysis is based on public case facts, not a standard):

LayerEmphasisConcrete Content in the AI4S Direction
L1 Context Engineering Layer★★★★Literature evidence and domain knowledge, domain databases, protocols and data dictionaries, failed experiment records
L2 Tools and Execution Layer★★★★★Simulation and experiment MCP (e.g., a domain server exposing simulation-launch and data-sorting tools), experiment tracking, robot and instrument interfaces
L3 Orchestration and Control Layer★★★★★lab-in-the-loop architecture: agents do hypothesis generation, experiment design, and data interpretation; humans do problem definition, adoption judgment, and physical experiment execution
L4 Memory and State Layer★★★★Hypothesis library, literature citation library, experiment records, and failure records (a searchable debugging knowledge base)
L5 Evaluation and Observation Layer★★★★★ground truth = physical laws / experimental verification / conservation laws; tiered evaluation: cheap physical constraints executed at high frequency, expensive experimental validation triggered on demand
L6 Governance and Safety Layer★★★★★Research integrity (no data or citation fabrication), experimental ethics review, biosafety, model interpretability, and the Scientific Data Management Measures' "deposit before acceptance"

The core lies in L5 and L3.

L5 is a first-order problem because all of the value in this direction is ultimately adjudicated by "whether it withstands physical laws and experiments." The verifiable forms of ground truth include:

  • Physical constraints: conservation of energy, mass, and momentum; convergence; comparison against analytical solutions or mature physical models (e.g., ECMWF HRES / ENS in meteorology).
  • Experimental verification: synthesis confirmation by independent laboratories (e.g., the 736 external independent syntheses of materials predicted by GNoME), or clinical or statistical significance (e.g., survival analysis in target evaluation).
  • Reproducibility: the same result can be recomputed or reproduced by independent teams on different setups (corresponding to the ACM four-tier reproducibility definition).

The counterexample warning likewise comes from the absence of L5: Sakana's AI Scientist was found by an independent evaluation to misjudge existing concepts as novel and to produce hallucinated numerical results in 4 of 7 manuscripts (per public reporting). Without strict L5, the larger the scale of AI4S output, the greater the accumulated hallucination.

L3 is core because the field is converging on the lab-in-the-loop architecture: agents are responsible for hypothesis generation, experiment design, and data interpretation, while human researchers are responsible for posing questions, judging which AI outputs to adopt, and executing physical experiments. The instance in Stanford's virtual CSO system where "the science-review agent pointed out that a ligand–receptor inference lacked spatial context, thereby triggering spatial transcriptomics analysis" (see section 3.3) demonstrates the value of self-criticism among agents as an orchestration mechanism.

Bottleneck layer: L5's cost structure. Cheap verification (conservation laws, convergence, comparison against existing models) can be automated and executed at high frequency; expensive verification (experimental synthesis, wet experiments, clinical) cannot be run at high frequency and must be human-in-the-loop. Designing a tiered evaluation system of "surrogate metrics first, ground-truth calibration periodically" is the greatest engineering challenge for Harness implementation in this direction.

1.4 Value and Bottlenecks

Value is reflected in three places:

  • Expanding the hypothesis space: models can systematically propose hypotheses among hundreds of millions of candidates (protein structures, crystal structures, drug–indication pairings), freeing humans from enumeration labor.
  • Compressing cognitive labor: FutureHouse's Robin system is described by the institution as able to complete in one day part of the literature review and hypothesis generation work that human researchers would originally need months to do; another report claims it compressed the cognitive labor of a discovery cycle from an estimated 872–937 person-hours to under 2 hours (the two figures have different calibers; see the information gap statement).
  • Turning verification standards from verbal into automatic: conservation-law checks, convergence determination, and citation verifiability—steps that originally depended on the quality of reviewers—can be solidified into gates that every output must pass.

Bottlenecks are likewise threefold:

  • Novelty misjudgment and numerical hallucination: AI Scientist's 4/7 problem rate shows that the stronger the generative capability, the more dangerous the unvalidated output.
  • Out-of-distribution failure: AI weather models systematically underestimate record-breaking extreme events, and the gap grows with extremity (see section 3.2)—the tail of the training distribution is precisely the most scientifically valuable region.
  • Verification cost and the trust gap: expensive verification cannot be run at high frequency, making the "AI proposes, humans select" step the bottleneck; the controversy over GNoME's "tenfold" claim is an embodiment of this gap (see section 3.3).

2. Glossary

TermEnglish / AbbreviationDefinition
AI for ScienceAI4SA research paradigm that accelerates scientific discovery with AI methods, spanning proteins, materials, meteorology, mathematics, physics, and other fields
Protein structure predictionProtein Structure PredictionThe task of predicting a protein's three-dimensional structure from its amino-acid sequence; the CASP competition is its traditional evaluation venue
CASPCritical Assessment of protein Structure PredictionA biennial global evaluation competition for protein structure prediction; AlphaFold2 became famous at CASP14
GDTGlobal Distance TestA measure of the correctness of protein structure prediction, out of 100; AlphaFold2's median GDT at CASP14 was 92.4
Deep-learning structure prediction databaseAlphaFold DBThe protein structure database released by DeepMind, covering over 200 million protein structures
Structural motifs diffusion generationRFdiffusionA generative model that denoising-fine-tunes on top of RoseTTAFold, moving from "reading structures" to "inventing structures"
Weather foundation modelAI Weather ModelModels that replace or supplement numerical weather prediction with neural networks; representatives include Pangu-Weather, GraphCast, GenCast, Fengwu, FuXi, etc.
Reanalysis dataReanalysis, ERA5Historical atmospheric datasets produced by combining physical models with data assimilation (4D-Var) and observations; the training-data source for most AI weather models
Ensemble forecastEnsemble ForecastA method that produces probabilistic forecasts by perturbing multiple ensemble members; GenCast produces ensemble forecasts with a diffusion model
Physics-informed neural networkPINN, Physics-Informed Neural NetworkA neural network method embedding physical equations into the training loss as constraint terms; a frontier direction of hybrid AI-physics modeling
Conservation lawsConservation LawsPhysical laws such as conservation of energy, mass, and momentum; the core automated verification basis for AI4S output
Materials discoveryMaterials DiscoverySearching for and designing new materials with target properties; representative systems include GNoME and MatterGen
Metal-organic frameworkMOF, Metal-Organic FrameworkPorous crystalline materials commonly used in scenarios such as gas adsorption; a typical target of AI material screening
Grand-canonical Monte CarloGCMC, Grand-Canonical Monte CarloA computational method for simulating statistical-physics problems such as adsorption
TokamakTokamakA magnetically confined nuclear fusion device; since 2022 reinforcement learning controllers have successfully run real plasmas
lab-in-the-loopLab-in-the-loopA collaboration architecture where agents handle hypotheses, design, and interpretation, and humans handle problem definition, adoption judgment, and physical experiment execution
AI ScientistAI ScientistAn end-to-end system that autonomously generates hypotheses, runs experiments, writes papers, and simulates review; the representative is Sakana AI's system of the same name
Novelty misjudgmentNovelty MisjudgmentMisjudging existing concepts as new discoveries; independent evaluation found this issue in 4 of AI Scientist's 7 manuscripts
Numerical hallucinationNumerical HallucinationValues generated by the model that were never executed or conflict with execution results
Peer reviewPeer ReviewThe quality-review mechanism of scientific publishing; note that review at the workshop level of a conference is not equivalent to acceptance at the main conference
Scientific data depositionScientific Data DepositionThe mechanism, under the Scientific Data Management Measures, whereby scientific data from government-budget-funded projects must be deposited before acceptance

3. Case Studies

3.1 Case Study 1: Protein Structure Prediction — From CASP to Programmable Design

3.1.1 Background

Protein structure prediction is the first "theorem-level" success scenario of AI4S: a clearly defined task (from sequence to structure), a mature evaluation system (the CASP competition and the GDT metric), and a ground truth determined by experiment (X-ray / cryo-EM structures). In 2018, AlphaFold first demonstrated deep learning's potential on this task at CASP13.

3.1.2 Approach

The key milestones from 2020–2024 are as follows (multiple sources cross-agree; source grades are mainly A/B, see references):

TimeAchievementKey Data
2020-12AlphaFold2 at CASP14Median GDT 92.4 / 100; best prediction on 88 of 97 targets; average error about 1.6 Å (about one atom's width)
2021RoseTTAFold (Baker Lab) published in Science"Three-track" neural network (sequence / distance / coordinates processed simultaneously); a single gaming PC computes a structure in as little as 10 minutes; open source
2022AlphaFold protein structure database releasedCovers over 200 million proteins, almost all known protein sequences
2023-03ESMFold (Meta) published in ScienceNo sequence alignment needed; folded 617 million metagenomic proteins
2023-07RFdiffusion (Baker Lab) published in NatureMoved from "reading structures" to "inventing structures"; hundreds of designs were manufactured in the lab
2023-09AlphaMissense published in ScienceScored all 71 million possible missense variants of the human proteome, classifying 89% of them (previously only about 0.1% were confirmed by lab or clinic)
2024-05AlphaFold 3 (DeepMind and Isomorphic Labs) published in NatureExtended from monomeric proteins to complexes (including DNA, RNA, ions, small molecules); interaction accuracy for protein and other molecular types improved by at least 50% over existing methods

As for adoption scale, per public reporting, 2 million researchers across 190 countries have used AlphaFold 2 or 3 (grade B source).

3.1.3 Results

This case establishes a path of the greatest reference value for Harness design—evaluation first, generation follows, open source amplifies:

  1. The CASP competition provided physical-level ground truth: predictions were directly compared against experimentally determined structures, with no dispute over winners. This is exactly the ideal form of the L5 evaluation layer—first a trustworthy criterion, then model competition, and finally narrative convergence.
  2. Open source and database-ification amplified marginal value: RoseTTAFold's open source and AlphaFold DB's coverage of 200 million structures let downstream researchers consume the results without training their own. In Harness terms: model results only become scientific productivity once solidified into searchable, reusable artifacts (L4).
  3. The leap from prediction to design depends on a feedback loop: RFdiffusion's designs must be manufactured and validated in the lab—the closed loop of generation (L2/L3) and experimental verification (L5) is what makes a complete scientific discovery cycle.

Boundaries that must be stated honestly: AlphaFold 3's full model was not entirely open-sourced, and the community has discussed its reproducibility and fair use; "at least 50% improvement in accuracy" is the paper's conclusion against existing methods, limited in scope to the interaction types it evaluated. These boundaries do not affect the conclusions above, but should be preserved when citing.

3.2 Case Study 2: AI Weather Foundation Models — Surpassing Numerical Forecasting and Systematically Underestimating Extreme Events

3.2.1 Background

Medium-range weather forecasting is the field in AI4S where "AI first challenged, on a routine task, equations that physicists wrote for a century." Since 2023, models such as Pangu-Weather, GraphCast, FourCastNet, FuXi, and Fengwu have been published one after another, and reported to surpass operational numerical weather prediction in both accuracy and speed.

3.2.2 Approach

A chronology of key facts (multiple sources cross-agree):

TimeAchievementKey Data
2023-07Pangu-Weather (Huawei Cloud) published in NatureTrained on 39 years of atmospheric records; the first ML model reported to surpass operational numerical weather prediction in both accuracy and speed; global forecast on a single GPU in about 1.4 seconds; over 10,000× faster than traditional numerical methods
2023-11GraphCast (DeepMind) published in Science0.25° resolution, 10-day forecast; surpassed ECMWF HRES on 90.3% of 1,380 verification targets; under 1 minute on a single TPU
2023FuXi (Fudan); FourCastNet (NVIDIA)FuXi achieves 15-day forecasts with a cascade model, its accuracy hitting the ensemble mean of ECMWF's 15-day ensemble forecast for the first time; FourCastNet is the first 0.25°-resolution model
2023-04 onwardFengwu (Shanghai AI Laboratory)First to achieve effective forecasts beyond 10 days; FengWu-GHR resolution raised to 0.09° (about 9 km), effective forecast length raised from 10.75 to 11.25 days (traditional models improve by only about 1 day every 10 years)
2024-12GenCast (DeepMind) published in NatureDiffusion model producing probabilistic ensemble forecasts; outperformed ECMWF ENS on 97.2% of 1,320 variable × lead-time combinations, reaching 99.8% for lead times beyond 36 hours; a single TPU v5 produces a 15-day forecast in about 8 minutes
2025-02ECMWF's own operational model AIFS Single put into operationsEnsemble version put into operations on 2025-07-01
2025-10China Meteorological Administration evaluationA first tier formed around "FuXi", "Puyun", "Fengwu", "Fengqing", and "Pangu"; usable forecast days for key circulation elements reached 10 days; while also noting that AI models are better positioned as an effective supplement to traditional numerical forecasting
2026-05Second AI weather forecast model demonstration program launchedEstablished a standardized evaluation system covering nowcasting, short-to-medium-range, and sub-seasonal time scales; released meteorological industry standards such as forecast-performance verification and evaluation, providing the basis for operational admission of models

Critical content (must be presented alongside the success narrative): according to research related to Science Advances, the major AI weather models systematically underestimate record-breaking heat, cold, and wind events, and the more extreme the event, the larger the gap; among them, AIFS especially underestimates storm peak wind speeds because its MSE training objective tends to smooth the atmospheric gradients that produce sharp peaks. The conclusion: physical models still prevail in the tail of the distribution—where extreme events live.

A deeper structural dependency must also be declared: AI weather models rely on ERA5 reanalysis data for both training and inference, and ERA5 is itself the output of physical models plus data assimilation (4D-Var). If the traditional physical models and assimilation pipeline were abandoned, AI models would lose the strictly structured initial conditions they need for training and routine inference. A pure AI model is essentially a highly complex pattern-matching interpolator that lacks a structural understanding of conservation laws and could, in principle, produce physically impossible atmospheric states when pushed beyond the boundary of its training distribution. The frontier direction is a deep hybrid of PINN and differentiable solvers with machine-learning parameterization schemes.

3.2.3 Results

This case yields three structural conclusions for AI4S (and indeed for all data-driven scientific directions):

  1. Benchmark superiority does not equal task superiority: GraphCast beating HRES on 90.3% of verification targets is a fact; but extreme events fall precisely in the remaining tail, and the tail is where forecast value is highest. "90% superior" and "systematically underestimating extremes" are two sides of the same set of facts and must be stated together.
  2. Training-data dependence sets the capability ceiling: the ERA5 dependency means AI weather models are methodologically not substitutes for physical models but consumers of them. The China Meteorological Administration's "effective supplement" positioning, and its use of verification-evaluation standards as an operational admission threshold, is a direct expression of Harness L5 thinking at the industry-governance layer.
  3. The emergence of industry standardization marks the direction's maturity: the 2026 meteorological industry standard turned "verification and evaluation" into an admission condition, which equals turning L5 from a recommendation into a mandatory gate—a governance path all AI4S subfields can follow.

3.3 Case Study 3: Testing the Mettle of Autonomous Scientific Discovery Systems — Success, Controversy, and lab-in-the-loop

3.3.1 Background

In 2024–2026, "AI Scientist"-type systems entered a period of dense releases: Sakana AI's The AI Scientist, FutureHouse's Robin and Kosmos, Google's co-scientist, Stanford's virtual CSO, ChemAgents, Lila Sciences, and others. There is a significant gap between the field's promotional narratives and its actual mettle; independent evaluations and scholarly criticism provide valuable calibration data.

3.3.2 Approach

First the facts, then the challenges.

Positive facts:

  • Stanford's virtual CSO B7-H3 case demonstrated a complete loop of multi-agent collaboration and self-criticism: when evaluating B7-H3 as a lung cancer treatment target, the statistical-genetics agent found no significant germline association but flagged regulatory activity; the single-cell atlas agent identified B7-H3 enrichment in cancer-associated fibroblasts; the science-review agent pointed out that the ligand–receptor inference lacked spatial context—this self-criticism triggered subsequent spatial transcriptomics analysis that confirmed an immune-exclusion microenvironment around B7-H3's high-expression regions; the clinical-trial agent ran a stratified survival analysis on the lung adenocarcinoma genome atlas (disease-specific survival HR = 1.82, p = 0.031); it ultimately recommended the antibody–drug conjugate (ADC) modality, and downgraded the priority of small molecules and PROTACs after the pharmacologist agent found no druggable pocket.
  • Mathematics: AlphaGeometry2 solved 84% (42/50) of IMO geometry problems at gold-medal level; in 2026-02, DeepMind's Aletheia autonomously solved 4 open problems among the 700 Erdős conjectures and generated complete research papers.
  • Physics control: in 2022-02, a reinforcement learning controller from DeepMind and EPFL ran on a real tokamak with 90 measurement inputs, 19 voltage commands, and a 10 kHz frequency, maintaining a plasma configuration that previously required weeks of expert hand-tuning—the first time ever.

Independent challenges (must be presented with as much space as the success cases):

  1. AI Scientist's novelty misjudgment and numerical hallucination: Sakana's AI Scientist was found by an independent evaluation to misjudge existing concepts as novel in 4 of 7 manuscripts and to produce hallucinated numerical results. One paper generated by its improved version once reached the peer-review acceptance threshold at a workshop of ICLR 2025—this is at the conference-workshop level and is not equivalent to acceptance at the main top-conference track; this must be noted when citing.
  2. The controversy over GNoME's "tenfold" claim: GNoME discovered 2.2 million new crystal structures, of which 381,000 were judged stable, and external laboratories independently synthesized and confirmed 736. But the implicit claim that "AI expanded the catalog of useful materials tenfold" has been retracted in spirit (if not in form); the Nobel-style excitement around GNoME has settled into a more restrained narrative: AI is good at proposing candidates, while human chemists still have to pick the few worth doing.
  3. The challenge to Robin's ripasudil: FutureHouse's Robin proposed using the glaucoma drug ripasudil for dry age-related macular degeneration, mechanistically via ROCK inhibition—but this was challenged: Robin's contribution is the "drug–indication pairing," not the underlying mechanism itself; this is a meaningful contribution but more limited than the original framing suggested.
  4. Humans still lead on complex scientific tasks: a recent Nature study further notes that humans still outperform the best AI agents on complex scientific tasks, though the gap is narrowing (this secondhand source has limited credibility; marked [To be verified]).
  5. Lack of shared benchmarks: a Nature systematic selection guide notes that the field still lacks shared benchmarks to distinguish real progress from marketing, and that the "AI scientist" label varies greatly in meaning from vendor to vendor (secondhand source; recommended to check the original).

3.3.3 Results

With the positive and negative facts juxtaposed, three conclusions most central to Harness design in this direction can be distilled:

  1. lab-in-the-loop is an architecture validated by both sides. The Stanford case shows that self-criticism among agents (the review agent triggering spatial analysis) can improve conclusion quality; the challenges to AI Scientist and Robin show that end-to-end autonomous output lacking external validation has limited contribution density. Convergent conclusion: agents handle hypotheses, design, and interpretation, while humans handle problem definition, adoption judgment, and physical experiments—this is exactly the standard form of the L3 orchestration layer in this direction.
  2. Numbers must come from execution, and novelty must pass retrieval comparison. AI Scientist's two classes of failure (numerical hallucination and novelty misjudgment) each map to two engineering-able gates: numbers can only come from simulation / code-execution results; "first discovery"-type claims must first pass literature retrieval comparison. Both are written into the red lines in section 4.1.
  3. Contribution narratives must be tied to verification grades. GNoME's 381,000 stable candidates are a computation-level conclusion, while the 736 independent syntheses are an experiment-level conclusion—nearly three orders of magnitude apart. Harness output norms should mandatorily distinguish the three evidence grades of "computational prediction / experimental confirmation / independent reproduction" and prohibit conflating them.

4. Practice Standards

4.1 AGENTS.md Specification

Standard-source statement: the following is a proposed standard draft of the AGENTS.md for the scientific discovery direction (AI4S) put forward in this document. As of now, there is no original AGENTS.md specification for the AI4S direction issued by any official body, industry association, or standards organization; AGENTS.md is a community convention, not a standard. This draft inherits all clauses of the data-science team-level AGENTS.md and tightens and extends them for the AI4S direction.

# AGENTS.md —— 科学发现(AI for Science)

> 继承数据科学组级 AGENTS.md 全部条款。本节为科学发现方向的收紧与扩展。
> 本文件为标准建议稿,业界尚无官方标准版本。

## 角色与边界

- 本 Agent 是**科学假说的起草者与验证的组织者**:可以综述文献、提出假说、设计模拟与实验、执行计算验证、解读结果。
- 不可以直接执行的操作:执行物理 / 化学 / 生物实验、进行临床或伦理裁决、单方面宣布「首次发现」、替人类决定采纳哪个假说。
- **ground truth 不是人类标注,而是物理定律、实验验证与守恒律**。任何与物理约束冲突的计算结果,默认为错误,除非有明确证据说明约束条件差异。
- 问题的定义权、假说的采纳权与实验的执行权归领域科学家,不归本 Agent。

## 环境假设

- 领域工具可用:模拟引擎(分子模拟、CFD、气象模型等)、领域数据库、文献检索接口。
- 模拟与实验工具经工具契约暴露(如 MCP 服务器);工具不直接承载计算,而是产出可由工作流引擎 / 调度系统承载的作业描述。
- 验证判据可用:守恒律容差、收敛阈值、基准算例、(如适用)实验对照协议。
- 文献库与引用数据库可检索,且检索结果带可点击来源。
- 失败实验记录库可用(可为空集,但需有沉淀路径)。
- **任一验证判据缺失时,先与领域科学家共同定义判据,再开始计算。无判据的模拟产出一律标记为未验证。**

## 上下文加载顺序(Context Budget)

1. 科学问题与验收判据(必须):假说、目标性质、判据类型与容差
2. 领域约束:守恒律、边界条件、适用范围(必须)
3. 文献证据:与假说直接相关的既有工作与引用(**新颖性判断前必须**)
4. 领域数据库条目:结构、性质、已验证记录
5. 同类既往模拟:参数、判据、结果与失败记录
6. 计算资源约束:精度、规模、时长与成本上限
7. 实验可行性与安全约束:涉及湿实验 / 生物 / 临床时的伦理与安全前提

- 文献证据必须携带可点击来源;无来源的领域知识不得作为新颖性或正确性的依据。

## 工具契约

| 工具 | 用途 | 模式 | 约束 |
|---|---|---|---|
| 文献检索 | 查询既有工作、引用核查 | 只读 | 结果必须带来源;新颖性比对结果留痕 |
| 领域数据库 | 查询结构、性质、既有记录 | 只读 | 不得修改公共数据库 |
| 模拟工具接口 | 发起模拟任务描述 | 受限写 | emit 作业描述给工作流引擎;不直接操控调度器 |
| 数值与符号计算 | 计算、推导、统计分析 | 受限执行 | 数值必须由执行产生并记录输入输出 |
| 验证判据执行器 | 守恒律检查、收敛判定、基准对照 | 自动 | 每次模拟产出必经;结果不可被跳过 |
| 实验记录本 | 读写实验设计与结果记录 | 受限写 | 实验执行状态由人更新,Agent 不得代填 |

- 工具参数必须做模式校验;格式错误在调用前拦截。
- 每次模拟记录:判据、参数、版本、运行时长、守恒律偏差与收敛残差。

## 任务执行流程(SOP)

1. **问题定义**:与科学家对齐科学问题、目标性质与验收判据;判据缺失则先补判据。
2. **文献综述**:检索并汇总既有工作,标注来源;识别可借鉴方法与已知失败。
3. **假说生成**:给出假说及其可检验推论;每个假说附验证判据草案。
4. **新颖性比对**:对「首次 / 新发现」类主张执行文献检索比对;比对结果留痕,未过比对不得声称新颖。
5. **计算设计**:选择方法与参数,声明适用范围与已知的近似。
6. **执行与物理验证**:经工作流引擎执行模拟;自动执行守恒律检查、收敛判定与基准对照。
7. **结果解读**:区分「计算预测 / 实验确认 / 独立复现」三档证据等级;列出与预期不符的反例。
8. **实验设计移交**:需要湿实验 / 物理 / 临床验证时,输出实验设计草案,交由人类执行;Agent 不得代执行。
9. **结论分级与移交评审**:结论按证据等级分级输出;所有引用可点击核查;进入人工评审与(如适用)同行评审。
10. **归档**:假说、判据、参数、结果、失败记录与文献比对记录归档;失败原因写入失败记录库。

## 验证与证据要求

- **每条数值必须由执行产生**:来自模拟、代码或统计计算的实际输出,禁止由模型凭记忆或估算生成。
- **物理约束优先于拟合优度**:结果违反守恒律或产生物理上不可能的状态时,判定为无效,不得以「整体误差小」辩护。
- **证据等级三档制**:计算预测(未经实验)→ 实验确认 → 独立复现;输出中必须显式标注当前档位。
- **新颖性主张必须先过文献比对**:比对关键词、命中结果与差异说明留痕。
- **引用必须可点击且逐条可核查**:无来源的引用禁止出现在任何产出物中。
- 主动报告反例:与假说不符的结果必须完整呈现,禁止只报告支持性结果。
- 规模结论与已验证结论分开陈述:如「候选 38 万」与「已合成确认 736」是两句话,不是一句话。

## 失败与升级策略

- 同类失败重试不超过 2 次;第 3 次改变策略或升级。
- **物理约束检查失败**:默认结果无效;先排查边界条件与参数,再考虑方法适用性;不得放宽判据。
- **文献比对发现已有工作**:更新假说定位,如实引用既有工作,禁止重复包装。
- **模拟与既有基准冲突**:报告冲突并升级给科学家,禁止选择性地忽略不利的基准结果。
- **工具参数格式错误**:停止重试,修正调用构造;多次出现上报工具契约缺陷。
- **涉及生物安全 / 伦理敏感的实验设计**:停止生成细节,升级给伦理与安全负责人。
- **成本或时长超预算**:暂停并报告,获批后继续。
- 升级时携带:科学问题、假说、判据、已执行验证、反例清单、建议下一步。

## 安全与合规红线

- 不得伪造、篡改或选择性报告数据。
- 不得伪造引用、编造文献或夸大既有工作的支持程度。
- 不得在未过文献比对的情况下使用「首次」「首创」类表述。
- 不得代替人类执行或声称已执行物理、化学、生物实验。
- 涉及病原体、危险化学品、临床与动物实验的设计,须先通过伦理与安全审查。
- 依《科学数据管理办法》,政府预算资助项目形成的科学数据须先汇交、再验收;不得私匿应汇交数据。
- 不得绕过同行评审流程直接对外发布科学结论。

## 禁止事项

- 禁止生成未经验证的数值结论。
- 禁止把计算预测表述为实验确认;禁止把单点确认表述为普遍规律。
- 禁止把会议研讨会层面的接收表述为主会录用或「顶会发表」。
- 禁止跳过新颖性比对直接宣称突破。
- 禁止在判据未定义时开始大规模模拟。
- 禁止删除或淡化失败记录。
- 禁止跨方向复制通用模板;AI4S 的验证判据与治理要求与数据分析、HPC 有实质差异。

## 输出格式

- 结论先行 → 证据等级(计算预测 / 实验确认 / 独立复现)→ 判据执行记录 → 反例与局限 → 建议。
- 数值带单位与判据上下文;范围用「~」连接;百分比数值与 % 之间无空格。
- 文献引用带可点击来源;「首次发现」类表述必须附比对记录编号。
- 候选规模与已验证数量分开陈述,用表格列出:证据等级、数量、验证方式、验证方。
- 失败记录使用固定格式:假说、判据、失败现象、根因分析、可复用教训。

## 评估与自检

- [ ] 验收判据在计算开始前已定义并双方确认
- [ ] 全部数值来自实际执行,输入输出留痕
- [ ] 守恒律检查与收敛判定已自动执行且通过,或失败已如实报告
- [ ] 新颖性主张已过文献比对,比对记录留痕
- [ ] 全部引用可点击且逐条核查
- [ ] 证据等级三档制已执行,候选与已验证分开陈述
- [ ] 反例与不利结果已完整呈现
- [ ] 需实验验证的部分已移交人类,未代执行
- [ ] 伦理与安全约束已确认(如适用)
- [ ] 假说、判据、结果与失败记录已归档

4.2 SKILL.md Specification

Standard-source statement: the following is a proposed standard draft of the SKILL.md for the scientific discovery direction put forward in this document; likewise, there is no original official standard.

---
name: ai4s-hypothesis-validation
description: 科学假说从生成到计算验证与实验移交的标准执行流程。适用于文献综述与新颖性比对、假说生成、模拟验证设计、守恒律与收敛性检查、证据分级与实验设计移交等任务。触发场景:任何以「提出并验证科学假设」为目标的智能体任务。
version: 1.0
created: 2026-09-12
---

# 科学假说验证标准流程

## 适用场景

- 文献综述与新颖性比对(某方向已知什么、还有什么空隙)。
- 科学假说的生成、形式化与可检验推论设计。
- 计算验证:模拟参数设计、守恒律与收敛性检查、基准对照。
- 实验设计草案编制与移交(湿实验 / 物理 / 临床)。
- 不适用场景:实验的实际执行与裁决、临床决策、伦理与安全审批。

## 前置条件

- 科学问题与目标性质已与领域科学家对齐。
- 验证判据(守恒律容差、收敛阈值、基准算例)已定义;未定义则先补。
- 文献检索接口与领域数据库可用且结果带来源。
- 模拟工具经工具契约可用;工作流引擎与资源预算已确认。
- 失败记录库可用(可为空集)。

## 输入

| 输入项 | 必需 | 说明 |
|---|---|---|
| 科学问题与目标性质 | 是 | 待发现或待验证的对象 |
| 验收判据 | 是 | 守恒律容差、收敛阈值、基准与对照 |
| 领域约束与适用范围 | 是 | 物理边界条件、方法近似声明 |
| 资源预算 | 是 | 计算规模、时长与成本上限 |
| 既有工作与失败记录 | 否 | 文献综述初稿、既往模拟记录 |

## 输出

| 输出项 | 必需 | 说明 |
|---|---|---|
| 假说书 | 是 | 假说、可检验推论、验证判据草案 |
| 新颖性比对记录 | 是 | 检索关键词、命中结果、差异说明 |
| 计算验证报告 | 是 | 参数、版本、守恒律偏差、收敛残差、基准对照 |
| 证据分级结论 | 是 | 计算预测 / 实验确认 / 独立复现三档 |
| 反例与局限 | 是 | 与预期不符的结果及解释 |
| 实验设计草案 | 否 | 交由人类执行的实验方案与安全前提 |

## 执行步骤

1. **对齐问题与判据**
   与科学家确认科学问题、目标性质与验收判据;判据缺失时先共同定义,禁止无判据开算。

2. **文献综述**
   检索既有工作并标注来源;汇总已知方法、已知结论与已知失败;形成新颖性空隙假设。

3. **假说生成**
   给出假说及其可检验推论;每个假说附验证判据草案与证伪条件。

4. **新颖性比对**
   对「首次 / 新发现」类主张执行文献检索比对:记录关键词、命中结果、差异说明;未过比对的主张降级为「改进 / 应用」。

5. **计算设计**
   选择方法与参数;声明近似与适用范围;设计对照算例与负对照;估算资源与时长。

6. **执行与自动验证**
   经工作流引擎提交模拟;自动执行守恒律检查、收敛判定与基准对照;检查不通过则结果无效并进入失败分析。

7. **结果解读与分级**
   区分计算预测 / 实验确认 / 独立复现三档;规模结论与已验证结论分开陈述;列出全部反例。

8. **实验设计移交**
   需要实验验证时输出实验设计草案,含安全与伦理前提;执行权归人类,Agent 不得代执行。

9. **人工评审**
   结论连同判据记录、比对记录与反例一并提交评审;评审意见回流入假说书。

10. **归档与失败沉淀**
    假说书、验证报告、比对记录、反例与失败原因归档;失败教训写入失败记录库。

## 质量标准(DoD)

判据与验证:

- [ ] 判据在计算开始前定义并经科学家确认
- [ ] 守恒律检查与收敛判定已自动执行并留痕
- [ ] 全部数值来自实际执行,禁止模型生成数值
- [ ] 与基准 / 解析解 / 既有模型的对照已完成

证据与引用:

- [ ] 新颖性主张已过文献比对,记录留痕
- [ ] 全部引用可点击且逐条核查
- [ ] 证据等级三档制已执行:计算预测 / 实验确认 / 独立复现
- [ ] 规模结论与已验证数量分开陈述

诚信与治理:

- [ ] 反例与不利结果完整呈现
- [ ] 实验设计已移交,未代执行任何实验
- [ ] 伦理与安全前提已确认(如适用)
- [ ] 假说、判据、结果与失败记录已归档

## 常见失败与处理

| 失败现象 | 根因 | 处理方式 |
|---|---|---|
| 结果违反守恒律 | 边界条件错误 / 方法不适用 / 时间步长过大 | 判定无效;排查条件与方法;禁止以整体误差小辩护 |
| 收敛残差不达标 | 判据过严或数值方法限制 | 报告残差与判据差距;升级给科学家;禁止放宽判据求完成 |
| 文献比对发现已有相同工作 | 综述不足 | 更新假说定位为改进 / 应用;如实引用既有工作 |
| 数值「过于完美」 | 数值由模型生成而非执行 | 回溯输入输出记录;补执行;报告诚信事件 |
| 候选规模被误读为已验证数量 | 证据等级未分开陈述 | 重写结论:计算预测 N 项、实验确认 M 项,分开成句 |
| 假说无法证伪 | 判据缺失或表述含糊 | 回到判据定义步骤;补证伪条件后再开算 |
| 实验团队拒绝移交方案 | 安全 / 伦理前提缺失 | 补齐安全与伦理评估;升级给负责人 |
| 模拟与既有基准冲突 | 方法近似或基准适用范围不同 | 如实报告冲突与两种解释;升级裁决 |

## 示例

**任务**:为某类多孔材料筛选吸附性能候选,并评估新颖性。

1. 对齐判据:目标性质为吸附容量;判据为能量守恒偏差小于 1e-6、与已发表基准材料偏差方向一致、统计显著性 p 小于 0.05。
2. 文献综述:检索既有筛选工作 24 篇,汇总已知方法与最佳已报道吸附容量。
3. 假说生成:提出「孔径分布与吸附容量存在非线性关系」的假说,附证伪条件。
4. 新颖性比对:以三个关键词组检索,命中 6 篇相近工作;差异说明留痕,主张降级为「在特定子族中的系统验证」。
5. 计算设计:蒙特卡洛模拟,声明力场近似;2,000 候选,预算内分批。
6. 执行与自动验证:经工作流引擎分批执行;守恒律检查全部通过;2 批收敛失败,进入失败分析。
7. 结果解读:计算预测前 20 候选(证据等级:计算预测);与既有实验对照的 40 个已知材料偏差在判据内。
8. 实验设计移交:输出前 5 候选的合成与测定草案,含安全前提;执行归人类。
9. 人工评审:结论、判据记录、比对记录与反例一并提交。
10. 归档:全部工件归档;2 批收敛失败的原因写入失败记录库。

4.3 Implementation Checklist

4.3.1 Context Layer (L1)

  • [ ] The literature evidence base is searchable, and results come with clickable sources
  • [ ] Domain database entries (structures, properties, verified records) are included in the context
  • [ ] Acceptance-criterion templates (conservation-law tolerance, convergence threshold, benchmark cases) are available
  • [ ] Failed experiment records can be retrieved and injected into the context
  • [ ] A method's approximations and scope of applicability are explicitly declared along with the context

4.3.2 Tools and Execution Layer (L2)

  • [ ] Simulation and experiment planning tools are exposed through tool contracts, emitting job descriptions rather than directly manipulating the scheduler
  • [ ] Tool parameters have schema validation, and malformed parameters are intercepted before invocation
  • [ ] Numerical and symbolic computation leaves input/output traces
  • [ ] Each simulation records criteria, parameters, version, duration, and constraint deviations

4.3.3 Orchestration and Control Layer (L3)

  • [ ] The lab-in-the-loop architecture is in place: hypothesis adoption and experiment execution are decided by humans
  • [ ] An inter-agent self-criticism mechanism is available (a review agent questions conclusions and triggers supplementary analysis)
  • [ ] Novelty comparison is a mandatory checkpoint before generating "first discovery"-type claims
  • [ ] Experiment design handoff has a clear state flow (draft → human execution → result feedback)

4.3.4 Memory and State Layer (L4)

  • [ ] The hypothesis library and literature citation library can accumulate sustainably
  • [ ] A failure-record library is established, and failure causes and lessons are searchable
  • [ ] Simulation parameters and versions are traceable and recomputable
  • [ ] The correspondence between review comments and new-version hypotheses is recorded

4.3.5 Evaluation and Observation Layer (L5)

  • [ ] Physical-level automated verification (conservation laws, convergence, benchmark comparison) is a mandatory gate for every simulation output
  • [ ] The three-tier evidence grading (computational prediction / experimental confirmation / independent reproduction) is applied across all outputs
  • [ ] Tiered evaluation: cheap physical constraints executed at high frequency, expensive experimental validation human-in-the-loop and triggered on demand
  • [ ] Scale conclusions and verified counts are stated separately
  • [ ] There is a complete channel for presenting counterexamples and adverse results

4.3.6 Governance and Safety Layer (L6)

  • [ ] Numbers must come from execution and citations must be verifiable—the two integrity red lines are technically enforced
  • [ ] "First"-type claims are constrained by comparison-record identifiers
  • [ ] Designs involving biosafety, hazardous chemicals, and clinical or animal experiments have an ethics and safety review checkpoint
  • [ ] The scientific-data deposit obligation (deposit before acceptance) is incorporated into the process
  • [ ] Unverified conclusions must not be released externally, and the peer-review process cannot be bypassed

5. Summary

The core demand of the scientific discovery direction on AI Harness can be summed up in one sentence: ground truth is not human annotation, but physical laws, experimental verification, and conservation-law constraints—so L5 becomes the first-order problem.

The evidence chain in this direction shows a rare symmetry:

  1. Every successful case stands on strong verification. AlphaFold2's persuasiveness comes from comparison against experimentally determined structures at CASP; GraphCast / GenCast's persuasiveness comes from comparison against ECMWF HRES / ENS; Stanford's virtual CSO's persuasiveness comes from spatial analysis and statistical significance triggered by the review agent. Every AI4S result that holds up is backed by an executable, reproducible verification system.
  2. Every case that failed lacked the same piece of verification. AI Scientist had 4 of 7 manuscripts with novelty misjudgment and numerical hallucination; the controversy over GNoME's "tenfold" claim (381,000 stable candidates versus 736 independent syntheses, nearly three orders of magnitude apart); Robin's ripasudil was pointed out to contribute a "drug–indication pairing" rather than the underlying mechanism; AI weather models systematically underestimate extreme events, and the gap grows with extremity—the failure forms differ, but the root cause is the same: generative capability exceeded verification capability.
  3. Humans still lead on complex tasks, but the industry has found the right collaboration form: lab-in-the-loop. Agents do hypotheses, design, and interpretation; humans do problem definition, adoption judgment, and physical experiments. This is both a calibration of the "human vs AI" narrative and the standard answer of the L3 orchestration layer in this direction.

Therefore the focus of this direction falls on L5 and L3: L5 solidifies physical laws and experimental verification into automatic gates and an evidence-grading regime (the three tiers of computational prediction / experimental confirmation / independent reproduction, with scale conclusions stated separately from verified conclusions); L3 solidifies the human–machine division of labor into a lab-in-the-loop orchestration mechanism. As for the remaining four layers, they all serve these two.

To sum up this direction's claim in one sentence: attach an evidence grade to every scientific claim, turn every physical constraint into an automatic gate, run every "first discovery" through literature comparison first, and keep every experiment in human hands.

Information Gap Statement

  1. There is no official or industry-recognized standard text for the scientific discovery direction's AGENTS.md / SKILL.md. Sections 4.1 and 4.2 are both standard drafts proposed in this document.
  2. "AI Harness × data science" and the authoritative definitions of AI4S, market size, adoption rate, and dedicated funding amounts and project counts have no reliable public source; this document provides no such figures.
  3. "Humans still outperform the best AI agents on complex scientific tasks" (a Nature study) comes from secondhand reporting, marked [To be verified]; it is recommended to check the original before citing.
  4. The content of the Nature "Which AI scientist suits your lab?" selection guide (the field lacks shared benchmarks, and the "AI scientist" label varies greatly in meaning) comes from secondhand reporting, marked [To be verified].
  5. The two efficiency figures for the Robin system have inconsistent calibers: "completing months of work in a day" and "compressing 872–937 person-hours to under 2 hours" come from different reports; this document only presents them side by side and does not adopt either as a conclusion.
  6. Figures such as AlphaFold's user base (190 countries / 2 million researchers), the tokamak RL-control details (90 inputs / 19 commands / 10 kHz), and Aletheia solving 4 open Erdős problems come from multi-source cross-referenced reporting (grade B); it is recommended to check the primary papers when citing.
  7. The conclusion that AI weather models underestimate extreme events comes from reporting of research related to Science Advances; the directional conclusion is consistent with scholarly consensus, and specific figures are not cited; the original is recommended for verification.
  8. Some performance figures of domestic Chinese weather foundation models such as Fengwu and FuXi (e.g., effective forecasts of 11.25 days, 0.09° resolution) come from vendor and media reporting (grade B); it is recommended to check the papers or official releases when citing.
  9. The statement about the Science Intelligence and Materials Creation Center's "investment exceeding ten billion yuan" and its hosting of the major AI4S special program comes from secondhand reporting and is not adopted by this document; no official public breakdown of the specific funding amounts and project counts of China's AI4S program was found.
  10. The material systems, criterion thresholds, and candidate counts in the section 4.2 example are illustrative constructions and do not represent any real research.

6. References

  1. AI for Science achievements chronology and key data — AIWiki. https://aiwiki.ai/wiki/ai_for_science
  2. AI for Science capabilities and milestones compilation — Achievements AI. https://achievements.ai/type/capability
  3. AI for Science knowledge tree (AlphaFold / GNoME / Aletheia entries, etc.) — BAAI Community. https://hub.baai.org.cn/knowledge-tree/8fd127c3-5f22-493d-8fa2-d23da3c3ace7
  4. AI in Physics review page (AlphaGeometry / tokamak RL control / GNoME controversy, etc.) — Scale Physics. https://scalephysics.com/horizons/aiinphysics/
  5. Research related to AI weather models underestimating extreme events (Science Advances / Earth & Environment) — Nature. https://www.nature.com/articles/s43247-025-02502-y
  6. Machine Learning Weather Forecasting analysis page (ERA5 dependency and hybrid-modeling directions) — Research. https://research.mental-momentum.ai/r/machine-learning-weather-forecasting-ijl0v7
  7. AI weather forecast entry (Fengqing / Fenglei / Fengshun / first tier / demonstration program) — Baidu Baike. https://baike.baidu.com/item/AI天气预报/68593824
  8. Agents for R&D Science (AI Scientist 4/7 evaluation, Robin ripasudil challenge, Stanford virtual CSO B7-H3 case) — Orchestra Bio. https://orchestra.bio/blog/agents-for-r-d-science
  9. Autonomous scientific discovery systems 2026 roundup (Co-Scientist / Kosmos / funding and tool ecosystem) — Mixflow. https://mixflow.ai/blog/the-ai-pulse-whats-new-in-autonomous-scientific-discovery-for-2026
  10. Nature "Which AI scientist suits your lab?" retelling page (the field lacks shared benchmarks, autonomy-tier selection) — Sinapti. https://sinapti.ca/post/es/nature-compara-los-sistemas-de-ia-que-automatizan-la-investi-xd43nve2
  11. Shanghai AI Laboratory's Sinan scientific-intelligence evaluation system (SciEvalKit's seven capabilities × six disciplines) — Shanghai AI Laboratory. https://www.shlab.org.cn/news/5444233
  12. SciEvalKit official repository — InternScience (GitHub). https://github.com/InternScience/SciEvalKit
  13. Awesome AI for Science (a full panorama of scientific-agent benchmarks: SciCode, MLAgentBench, NewtonBench, etc.) — GitHub. https://github.com/ai-boost/awesome-ai-for-science
  14. Reporting on the "AI-driven scientific research" special deployment — China Journalists Association website. https://union.china.com.cn/cmdt/txt/2024-03/26/content_42737343.html
  15. Agentic MOF Screening on Aurora (an AI4S and HPC crossover case) — supercomputing.news. https://www.supercomputing.news/hpc/agentic-mof-screening-aurora