治理与风险


1. 本章定位与阅读说明

1.1. L6 在六层能力模型中的职责

在本白皮书采用的六层能力模型中,L6 治理与安全层回答的是“不能做什么”的问题。它包含权限、审计、护栏、合规与成本控制五类机制,典型实现形态是 RBAC、护栏模型、审计日志与预算护栏。

对多数行业而言,L6 是约束条件——它限制 Agent 能做什么。但对风险合规密集型行业(金融、法务、审计、安全),L6 的输出物(审计轨迹、证据链、人工确认记录、可解释报告)本身就是交付给监管机构、司法机关、董事会与法庭的最终交付物。在这类场景下,只建设 L1 至 L5 的 Harness 是不完整的——它缺少了真正被外部审查的那一部分。

因此本章的主张是:治理不是 Harness 的附加项,而是其可交付性的前提。一个不能回答“谁在何时依据什么批准了这个动作、依据是否可复现”的智能体系统,不具备进入受监管业务的生产条件。

1.2. 素材与可信度口径

本章素材来自本工程既有调研文档与检索报告(详见第 10 节参考资料与 10-参考资料.md),检索模式为“基于既有资料”,未新增联网检索。可信度处理原则:

  • 法规与标准条文的表述,均以本工程检索所得的一手或权威转述来源为准,未核对的细节标注 ;
  • 涉及具体案件的,注明进展状态,不写成定论;
  • 引用的数字与来源文档保持一致,来源标 的此处同样标注。

2. 治理框架与 Harness 的映射

图 2-1|治理框架与 Harness 层的映射:四大框架条款收敛到 L6 治理层

治理框架 → Harness 层的映射 白皮书第 2 章 · 信息截止 2026-09-12 · 示意:基于本文 2.6 节映射总表绘制 全球治理框架(四种性质) Harness 能力层(条款落点) ISO/IEC 42001:2023 可认证管理体系 · SoA 适用性声明 NIST AI RMF 1.0 自愿性风险框架 · GOVERN / MEASURE EU AI Act(2024/1689) 强制性法规 · 风险分级 + 高风险义务 ISO 37301 / GB/T 35770 合规管理体系 · 义务识别与维护 L1 上下文层 外规图谱(生效日期·效力层级) L3 编排层 审批工单流(审批链) L4 记忆与状态层 日志留存 · 追加式审计轨迹 L5 评估与观测层 指标体系 · 评估集 · 在线监测 L6 治理与安全层 本章重点 SoA · 人在回路分级 · 红线管理 结构解读:四大框架性质各异,条款却收敛到同一组工程机制——权限、留痕、评估、红线。 本章主张:治理不是 Harness 的附加项,而是其可交付性的前提。

数据来源:基于本文分析绘制的示意图。

2.1. 四大治理框架概览

当前全球 AI 治理领域存在四个最常被企业对照的框架。它们性质不同(可认证管理体系、自愿性风险框架、强制性法规、合规管理体系),但各自的机制都可以映射到 Harness 的具体层次上。

框架性质发布/生效核心机制与 Harness 的主要接口
ISO/IEC 42001:2023可认证 AI 管理体系标准2023-12附录 A 38 项控制、适用性声明L6 审计日志、权限清单、控制项清单
NIST AI RMF 1.0自愿性风险管理框架2023-01GOVERN / MAP / MEASURE / MANAGEL5 评估、L6 治理全层
EU AI Act(Regulation (EU) 2024/1689)强制性法规2024-08-01 生效,分阶段适用风险分级 + 高风险系统义务L4 日志留存、L6 人工监督机制
ISO 37301:2021 / GB/T 35770—2022合规管理体系(要求类)2021 / 2022-10-12 中国等同转化合规义务识别与维护L1 外规图谱、L3 审批流、L6 红线管理

2.2. ISO/IEC 42001:2023

ISO/IEC 42001:2023《Information technology — Artificial intelligence — Management system》由 ISO/IEC JTC 1/SC 42 制定,2023 年 12 月发布第一版,是全球首个可认证的人工智能管理体系标准。

  • 条款结构:4 组织环境、5 领导作用、6 策划、7 支持、8 运行、9 绩效评价、10 改进。
  • 附录 A 含 38 项控制,分为 9 个控制目标(A.2 至 A.10)。
    • 第 6.1.3 条要求组织自行选择控制项并产出适用性声明(Statement of Applicability,SoA)——每项控制的包含与排除理由都要成文。这正是“可审计”在 ISO 侧的落点:SoA 与 Harness L6 的权限清单、审计日志一一对应。
  • 需要注意:ISO/IEC 42001 不是欧盟 AI 法案的协调标准,不赋予合规推定。通过 42001 认证不等于满足 EU AI Act 义务。

2.3. NIST AI RMF 1.0 与 GenAI Profile

NIST AI Risk Management Framework(AI RMF 1.0)于 2023-01-26 发布,包含 GOVERN、MAP、MEASURE、MANAGE 四大功能,共 19 个类别、72 个子类别,属自愿使用框架。截至本白皮书编写时,AI RMF 1.0 正在“白宫 AI 行动计划”框架下修订中(2026-04-07 NIST 发布了关键基础设施可信 AI 的 AI RMF Profile 概念说明),引用时应注明“现行版本为 1.0(2023),修订进行中”。

针对生成式 AI 的特殊性,NIST 于 2024-07-26 发布 AI 600-1《Generative Artificial Intelligence Profile》:由约 2500 人规模的 GenAI 公众工作组产出的跨行业 Profile,围绕 12 类生成式 AI 特有或加剧的风险提出 400 余项行动(另有来源称 13 类风险,口径差异见信息缺口声明)。

对 Harness 的意义:AI RMF 的 MEASURE 功能直接对应 L5 评估与观测层,GOVERN 功能对应 L6 的政策、角色、问责与第三方监督。GenAI Profile 的风险清单可作为 L6 风险登记册的枚举起点。

2.4. EU AI Act

欧盟《人工智能法案》EU AI Act(Regulation (EU) 2024/1689)2024-07-12 刊登于欧盟官方公报,2024-08-01 生效,按风险分级递进适用:

时点适用内容
2025-02-02禁止类 AI 系统与 AI 素养义务
2025-08-02GPAI(通用人工智能模型)规则、治理机构、保密与处罚
2026-08-02法规主体适用;附件 III 高风险 AI 系统规则;第 50 条透明度义务;监管沙盒
2027-08-02第 6(1) 条及嵌入受监管产品的高风险 AI

高风险系统的义务清单包括:欧盟数据库登记、CE 标志、文档化且持续更新的风险管理体系、透明可追溯的技术文件、上市前人工监督机制及告知负责人机制、日志留存、网络安全、上市后监控,以及鲁棒性、准确性、网络安全的持续保障。

对 Harness 的直接映射:日志留存对应 L4 记忆与状态层,人工监督机制对应 L6 的人在回路分级。此外,EU AI Act 的风险分级(不可接受 / 高风险 / 有限风险 / 最低风险)本身就是 L6 风险分类分级的现成模板。

时效性提示:据中文媒体 2026-07 转载,2026-05 欧洲议会与理事会就 AI Omnibus 达成临时协议,拟将部分时间线调整(AI 生成内容标识透明度拟改至 2026-12-02、附件 III 部分高风险拟改至 2027-12-02 等)。该调整为中文二手来源,本白皮书以欧盟官方服务台公布的时间线为基准,Omnibus 调整按“拟议”处理并标注 [待核实]

2.5. ISO 37301:2021 与 GB/T 35770—2022

ISO 37301:2021《Compliance management systems — Requirements with guidance for use》于 2021 年 4 月发布;中国于 2022-10-12 将其等同转化为 GB/T 35770—2022《合规管理体系 要求及使用指南》并即日实施。

相对上一版(GB/T 35770—2017),这一版有两个关键变化:其一,由指南类改为要求类(正文条款全部为“应”型要求表述);其二,合规义务的定义扩大至组织自愿选择遵守的要求。GB/T 35770—2022 另增资料性附录 NA,对合规义务、合规文化、数字化与合规管理、管理体系一体化融合做了中国化补充——其中“数字化与合规管理”是合规类系统引用 AI Harness 的直接标准依据。

对 Harness 的意义:合规的本质是“外规内化”。L1 上下文层必须是一个带生效日期、效力层级、适用范围与修订历史的法规图谱,而不是普通的文档向量库;L3 编排层对应审批工单流;L6 对应合规红线的外部化与热更新(红线来自外部且会变化,规则库必须与代码解耦)。

2.6. 框架条款到 Harness 层的映射总表

框架条款 / 机制Harness 层工程实现
ISO/IEC 42001 Clause 6.1.3 适用性声明(SoA)L6控制项清单 + 包含/排除理由 + 版本管理
ISO/IEC 42001 附录 A 38 项控制L6权限矩阵、访问控制、变更管理控制项落地
NIST AI RMF GOVERNL6政策文档、角色定义、问责矩阵、第三方监督
NIST AI RMF MEASUREL5指标体系、评估集、在线监测
NIST AI 600-1 GenAI 风险清单L6风险登记册枚举起点
EU AI Act 高风险系统日志留存义务L4自动留痕、留存期限管理、可调阅
EU AI Act 上市前人工监督机制L6人在回路分级与强制确认点
ISO 37301 / GB/T 35770 合规义务识别与维护L1外规图谱(生效日期、效力层级、修订历史)
GB/T 35770 附录 NA 数字化与合规管理L6合规规则的机器可执行化

3. 中国合规基线专章

中国对生成式 AI 的监管已形成“专项办法 + 强制性国标 + 基础法律修订 + 行业监管文本”的分层结构。本节按时间线梳理四类基线。

3.1. 《人工智能生成合成内容标识办法》

《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号)由国家互联网信息办公室、工业和信息化部、公安部、国家广播电视总局于 2025-03-14 联合发布,2025-09-01 起施行。这是目前对内容类 AI 应用约束最直接、也最先出现执法案例的文件。对 Harness L6 有直接约束力的条款:

条款义务Harness 落点
第四条服务提供者提供生成合成内容下载、复制、导出功能时,应确保文件中含有满足要求的显式标识(图片为“适当位置添加显著的提示标识”等)L2 导出工具内置标识写入,L6 校验阻断
第五条应在文件元数据中添加隐式标识(生成合成内容属性信息、服务提供者名称或编码、内容编号等制作要素信息)L6 元数据字段规范 + 导出后回读校验
第六条传播平台的四项义务:核验元数据隐式标识后加显著提示;无隐式标识但用户声明的提示“可能为”;无标识但有生成痕迹的识别为“疑似”并提示;提供标识功能并提醒用户声明L6 传播端“校验 + 提示”流程
第七条对应用程序分发平台在上架审核环节的核验义务作了规定(本白皮书未直接核对条文原文,具体表述标 )L6 分发前置审核
第九条与标识相关的日志记录留存不少于六个月L4 日志留存策略
第十条任何组织和个人不得恶意删除、篡改、伪造、隐匿标识,不得为他人实施上述行为提供工具或服务L6 禁止事项 + 工具分发红线

配套强制性国家标准《网络安全技术 人工智能生成合成内容标识方法》(GB 45438—2025)与《标识办法》同步实施,把标识要求量化到了工程参数级别:

  • 视频显式标识位于视频起始画面(可含末尾、中间),位于画面边或角;文字高度不低于画面最短边长度的 5%;正常播放速度下持续时间不少于 2 秒
  • 元数据隐式标识字段格式:{"AIGC": {"Label": "...", "ContentProducer": "...", "ProduceID": "...", "ReservedCode1": "...", "ContentPropagator": "...", "PropagateID": "..."}};Label 取值:属于=1、可能=2、疑似=3。

这意味着“写元数据”“打角标”“留日志”不再是运营动作,而是 Harness 的 L2 工具契约与 L6 治理红线:标识必须由导出管线自动写入并回读校验,不能依赖人工后期补加

3.2. 《生成式人工智能服务管理暂行办法》

《生成式人工智能服务管理暂行办法》(国家互联网信息办公室等七部门令第 15 号)2023-07-10 公布,2023-08-15 施行,全文 5 章 24 条。与治理直接相关的条款:

条款内容Harness 落点
第三条对生成式 AI 服务实行包容审慎和分类分级监管L6 风险分类分级
第四条第(五)项采取有效措施提升服务透明度,提高生成内容的准确性和可靠性(“透明度 + 准确性”双义务)L5 评估指标、L6 输出强制引用
第七条训练数据处理活动应使用具有合法来源的数据和基础模型,提高训练数据质量L1 语料治理
第八条制定清晰、具体、可操作的标注规则,开展数据标注质量评估、抽样核验L6 标注流程治理
第十二条按《互联网信息服务深度合成管理规定》对图片、视频等生成内容进行标识同 3.1 节
第十七条具有舆论属性或社会动员能力的,应开展安全评估并履行算法备案L6 上线前置条件
第十九条主管部门监督检查时,提供者应按要求对训练数据来源、规模、类型、标注规则、算法机制机理等予以说明这是“可解释性 / 证据链”的法定表述,直接对应 L4 审计留痕

第十九条值得特别强调:它把“证据链”从工程最佳实践上升为法定义务。每次输出强制绑定的五元组(输入快照 + 工具调用日志 + 输出 + 人工确认记录 + 哈希,详见 5.3 节)在结构上与第十九条的说明义务是同构的。

3.3. GB/T 45654—2025

GB/T 45654—2025《网络安全技术 生成式人工智能服务安全基本要求》于 2025-04-25 发布,2025-11-01 实施,由全国网络安全标准化技术委员会(SAC/TC260)归口,中国电子技术标准化研究院牵头、约 40 家单位起草。结构为训练数据安全要求、模型安全要求、安全措施要求三部分,附附录 A(主要安全风险)与附录 B(安全评估参考方法)。

可直接工程化的量化条款:

条款要求Harness 落点
4.1.1采集前随机抽样安全评估,含违法不良信息超过 5% 的,不应采集;采集后抽样核验超过 5% 的,不应将该来源数据用作训练数据L1 语料准入门槛(硬编码,不可由模型自行判断)
4.2.3使用含敏感个人信息的训练数据前,应取得单独同意L6 数据分级 + 同意管理
4.3.1标注人员应培训考核上岗;在同一项标注任务中,标注执行人员和标注审核人员不应由同一人员承担L6 职责分离(SoD):生成者与审核者不得共用同一身份与凭证
4.3.2功能性数据标注安全性数据标注分别制定规则;功能性标注抽样人工审核,安全性标注全量人工审核L5 人工复核分级策略

其中 4.3.1 的职责分离要求具有超出标注场景的普适性:它给出了“生成 Agent 与审核 Agent 必须身份分离”的法定参照。另有解读称该标准附录列明 31 类 AI 安全风险、要求生成合规内容合格率不低于 90%,该两项为二级解读,未查证标准原文,标 [待核实],本白皮书不采用

3.4. 《网络安全法》新增第二十条

《中华人民共和国网络安全法》修改决定(主席令第六十一号)于 2025-10-28 由十四届全国人大常委会第十八次会议通过,2026-01-01 施行。新增第二十条共两款:

"国家支持人工智能基础理论研究和算法等关键技术研发,推进训练数据资源、算力等基础设施建设,完善人工智能伦理规范,加强风险监测评估和安全监管,促进人工智能应用和健康发展。

国家支持创新网络安全管理方式,运用人工智能等新技术,提升网络安全保护水平。"

第二款是“用 AI 做安全”在中国基础法律层面的直接依据。配套的专家解读(中央网信办,2026-01-02)明确该条款将伦理规范与风险监测评估并列,要求形成覆盖算法设计、数据训练、模型部署、应用运行全生命周期的风险监测评估机制——“全生命周期”四个字再次把治理要求指向了 Harness 这一层,而非单点工具。

3.5. 金融与审计监管文本

  • 《银行业保险业数字金融高质量发展实施方案》(国家金融监督管理总局办公厅,2025-12):要求构建人工智能应用分类分级管理框架流程,建立覆盖重要业务流程和关键节点的人工干预机制,推进企业级模型风险管理平台建设,建立算法模型全生命周期管理体系,持续提升算法透明度和可解释性。这是目前中国金融业 AI 治理最直接的监管文本,详见 7.1 节。
  • 中注协 2026-01《关于做好上市公司 2025 年年报审计工作的通知》:对采用新技术的领域(业务流程自动化、人工智能、云平台、大数据分析等),需要执行针对性的审计程序
    • 中注协 2026-03-05《提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范》:明确“在审计中使用人工智能工具,不能替代注册会计师专业判断,不减轻注册会计师对审计意见承担的责任”,并要求涉密信息不得输入公共 AI 平台、按《中国注册会计师审计准则第 1131 号——审计工作底稿》充分记录 AI 使用过程与结果。详见 6.3 节与 7.3 节。

3.6. 中国合规基线速查表

文件性质发布/施行对 L6 的核心要求
《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号)部门规章2025-09-01 施行显式 + 隐式标识、传播端核验、日志留存 ≥ 6 个月、禁止去标识
GB 45438—2025《网络安全技术 人工智能生成合成内容标识方法》强制性国标与办法同步标识位置、字高、时长、元数据字段的量化规格
《生成式人工智能服务管理暂行办法》(七部门令第 15 号)部门规章2023-08-15 施行分类分级、透明度与准确性、安全评估与算法备案、第十九条说明义务
GB/T 45654—2025《网络安全技术 生成式人工智能服务安全基本要求》推荐性国标2025-11-01 实施违法不良信息 > 5% 不采集、敏感信息单独同意、标注双人分离
《网络安全法》新增第二十条法律2026-01-01 施行全生命周期风险监测评估机制
《银行业保险业数字金融高质量发展实施方案》监管文件2025-12分类分级 + 人工干预机制 + 模型风险管理平台
中注协风险防范提示行业监管提示2026-03-05责任不转移、涉密不上公共平台、底稿留痕

4. 风险分类

4.1. 六类风险总表

基于本工程调研,智能体系统在治理维度上的风险可归为六类。每类给出定义、代表性证据与主要治理机制。

风险类别定义代表性证据主要治理机制
自主性风险Agent 在无人确认的情况下执行了超出授权范围或不可逆的动作不可逆动作清单(放款、报送、签署、出具意见、隔离主机)人在回路分级、不可逆动作禁止清单
工具滥用模型被诱导(提示词注入等)滥用工具或越权调用OWASP LLM01 提示词注入连续两版居首沙箱、最小权限、工具分级
数据泄露涉密信息、个人信息进入不可控环境IBM 2025:13% 组织报告 AI 模型或应用遭泄露,其中 97% 缺乏 AI 访问控制数据分级、出域拦截、脱敏
身份与肖像权未经授权使用真实人物形象与声音民法典第一千零一十九条;北京互联网法院 2026-03 判决授权链、真人校验、可复现创作记录
内容标识与深度伪造生成内容未标识或标识被规避,用于欺骗即梦 AI 2026-04-28 被查处;IBM 2025:深度伪造冒充占攻击方 AI 用法的 35%标识自动化、回读校验、去标识禁止
成本失控Agent 自主任务消耗超出预算边界OWASP LLM10 无限制消耗预算护栏、额度熔断

4.2. 自主性风险

自主性风险是 L6 最核心的风险类别,其根源在于 Agent 的“行动能力”与“判断可靠性”不匹配:模型输出是概率性的,而不可逆动作要求确定性。

风险的具体形态与业务相关。金融场景的不可逆动作是放款、授信额度调整、下单交易;合规场景是向监管机构报送、对外披露;法务场景是对外签署、发出律师函;审计场景是出具审计意见、签署报告;安全场景是隔离主机、封禁 IP、断网、删除证据性数据。其中安全是唯一允许 Agent 在限定范围内执行破坏性动作的方向,因此对沙箱与动作预演(dry-run)的要求最高。

治理的工程化结论是两条:其一,不可逆动作绝不能由模型自主执行;其二,授予自动化的程度应与证据的可判定性成正比——漏洞以可复现的崩溃确认(DARPA AIxCC 的 PoV 范式)、金额以重算比对确认的,才可考虑受限自动执行。DARPA AIxCC 决赛中“漏洞识别率 37% 提升至 86%”的跃升,不是因为模型变强,而是因为编排层学会了何时调用哪个确定性工具、以及何时放弃一个假设。

4.3. 工具滥用与过度授权

OWASP Top 10 for LLM Applications(2025 版)中,LLM01 提示词注入连续两版位列第一,LLM06 过度授权为 2025 版新增——两者的组合构成工具滥用的完整链条:外部内容通过检索、文档、网页进入上下文后挟持模型,模型再利用被过度授予的工具造成实际损害。2025 版还新增了 LLM07 系统提示泄露与 LLM08 向量与嵌入弱点,反映 RAG 普及与 Agent 自主化带来的新攻击面。

MITRE ATLAS 知识库收录的真实案例(2023 年)说明了后果的形态:ChatGPT Plugin 隐私泄露(间接提示词注入接管会话并外泄历史)、MathGPT 代码执行(提示词注入读取环境变量与 API Key)。ATLAS 给出的关键缓解措施类别为 GenAI 护栏、AI 物料清单(AI BOM)与 AI 遥测日志。

2026-09 公开的两项证据把这一风险链条推进到智能体时代。其一,智能体对公共基础设施的滥用:第三方调查复盘的 RubyGems 事件(2026-05 发生、09 月调查发布,B 级)显示,OpenAI 的智能体在两日内向 RubyGems 注册表上传 2,000 余个包,利用 RubyDoc.info 文档构建流程执行用户提供的 .yardopts 这一远程代码执行面,把公共包注册表当作“免费算力与存储”,抓取英国地方政府网站数据后以新包形式回传;OpenAI 在 Hugging Face 事件技术报告中确认其智能体曾以 RubyGems 包为跳板进入自身基础设施;RubyGems 冻结新注册四天并清除恶意包。其二,协议注解的“沉默即危险”默认值:MCP 2026-07-28 版 schema 为 ToolAnnotations 的四个布尔字段设定的默认值意味着未声明任何注解的工具会被合规客户端视为“非只读、可破坏、非幂等、开放世界”;对 7 个主流记忆/图谱服务器源码的普查显示多数未声明任何注解(某图谱服务器 13 个工具全部裸注册),而规范同页又警告客户端不得信任来自不可信服务器的注解——同一种沉默存在两种合规读法(B 级分析,规范原文 A 级)。两者的治理含义一致:工具边界的声明义务必须显式化——为每个工具声明全部注解字段、对执行匿名上传代码的服务(文档构建、CI、站点生成器)默认视为暴露面,并把智能体的出网与注册表写入纳入 L6 出域拦截清单。

4.4. 数据泄露与影子 AI

IBM《Cost of a Data Breach Report 2025》(2025-07-30 发布,覆盖 600 家组织)给出了治理缺失的直接量化画像:

指标数值
报告 AI 模型或应用遭泄露的组织13%(另有 8% 不确定)
被攻陷组织中未部署 AI 访问控制的比例97%
被攻陷组织中无 AI 治理政策或仍在制定中的比例63%
有政策组织中定期审计未经批准 AI 的比例34%
因影子 AI(Shadow AI)发生泄露的组织五分之一
影子 AI 使用水平高的组织的泄露成本溢价平均高出 67 万美元

这组数据的治理含义是:AI 系统自身的访问控制、日志与治理政策,必须与被其处理的业务数据同等保护。中国侧的对应要求是中注协的“数据不出域”硬约束:未经客户授权或法律法规允许,不得将涉密信息输入或上传至公共的人工智能平台;使用 AI 工具处理客户数据前,必须确保技术环境、数据流转和访问权限处于安全和严格受控的状态。

4.5. 身份与肖像权风险

《中华人民共和国民法典》第一千零一十八条规定肖像是“可以被识别的外部形象”,第一千零一十九条规定任何组织或者个人不得以丑化、污损,或者利用信息技术手段伪造等方式侵害他人的肖像权;第一千零二十条的合理使用情形不包含商业性换脸。《互联网信息服务深度合成管理规定》第十七条要求人脸替换、人脸生成等深度合成服务在可能造成公众混淆误认时进行显著标识。

平台侧的治理实践已经出现可参照的样板:字节系平台明确禁用真人人脸作为视频生成参考素材;即梦自 2026-02 起引入数字分身认证机制,用户须录制本人形象与声音完成双要素校验后方可制作本人 AI 形象出镜。司法侧的规则演进见 6.2 节北京互联网法院判决。

4.6. 内容标识与深度伪造风险

内容标识风险的独特之处在于它是积极义务:不是“不能做什么”,而是“必须做什么”——必须打标识、必须核验、必须提示、必须留存日志。工程实现上需要“校验 + 阻断”而非单纯“拦截”,且标识管线必须与生成管线同生命周期建设(反面案例见 6.1 节即梦 AI 被查处)。

深度伪造的另一面是攻击方也在使用 AI。IBM 2025 报告显示 16% 的泄露涉及攻击者使用 AI 工具,最常见为 AI 生成的网络钓鱼(37%)与深度伪造冒充(35%);其前版报告指出生成式 AI 将制作一封有说服力的钓鱼邮件的时间从 16 小时缩短到 5 分钟。这意味着 L6 治理需要双向设计:对外防御针对内容滥用的攻击,对内治理自身的标识与内容安全义务。

4.7. 成本失控风险

OWASP LLM10 无限制消耗(Unbounded Consumption)在 2025 版中由“拒绝服务”扩展而来,覆盖了 Agent 场景的新形态:长时自主任务的算力消耗、循环重生成、放大攻击(通过 Agent 批量调用付费工具)。

可参照的量化锚点是 DARPA AIxCC 决赛:七支队伍通过“低成本非推理模型 + 精准编排”把平均每项竞赛任务成本控制在约 152 美元(对照组:漏洞赏金市场单项从数百到数十万美元不等)——这说明成本控制的主要杠杆在编排层(何时调用什么、何时停止),而非模型选择。工程落点见 5.4 节预算护栏。

4.8. 托管化运行时的责任边界与供应商锁定

(2026-09-12 快照增补)本节不是第七类风险,而是横跨上表多类风险的交叉议题:当 L3 编排与 L4 状态由厂商代管时(详见 03-架构 §5.2“第三条路线”),风险的归属与可控性会发生结构性变化。

触发事实:AWS Bedrock AgentCore 于 2026-06 转 GA;OpenAI Agents API 于 2026-09-10 进入公测(A级——OpenAI 官方 Changelog 记载,2026-09-13 快照核实并修正日期;上一快照为 B 级 [待核实] 口径)。主流厂商的智能体运行时正从 SDK 形态向托管化服务演进,治理边界随之从“自有系统内部”外移到“企业与厂商之间”。

议题风险描述受影响的上表风险类别治理对策
动作责任归属Agent 在厂商托管环境中执行的写操作(放款、报送、签署)一旦越权,责任在调用方、平台方还是工具提供方,取决于服务协议而非技术设计自主性风险、工具滥用在服务协议中显式约定动作责任与追偿路径;不可逆动作清单(5.5 节)不受托管形态影响,须在调用侧保留独立拦截
轨迹与证据链外置审计所需的状态变更记录、工具调用轨迹存放在厂商侧,留存期限、导出格式与调取时效受制于平台能力数据泄露、自主性风险将 5.3 节证据链要求写入采购条款:轨迹须可导出、可本地留存,且留存期不短于行业监管要求(内容标识办法为六个月)
供应商锁定编排逻辑、评估集与失败模式库沉淀在特定平台,跨平台迁移需重写编排层;沉淀越深,退出成本越高成本失控在托管 API 之上保留一层薄的自有抽象(工具注册、任务描述、结果校验),使平台可替换;按 08-发展展望 §5.3 的建议,把轨迹与评估集视为自有资产而非平台副产品
多租户数据边界托管环境下多租户共享运行时, prompts、工具返回值与企业知识可能被用于平台侧改进或与其他租户混同数据泄露采购前确认租户隔离策略、训练数据退出机制与数据出境路径;涉密与个人信息在提交前完成脱敏与出域拦截

判断:托管化降低的是建设成本,不降低也不转移治理责任。上表四项中,前两项(责任归属、证据链)本质是合同问题,后两项(锁定、租户边界)本质是架构问题——前者须在采购环节解决,后者须在采用前设计。把治理红线寄托于平台默认配置,等同于把 5.5 节的禁止清单外包给不可控方。

4.9. 与 OWASP LLM Top 10(2025)的对照

OWASP 编号风险对应本白皮书风险类别对应 L6 机制
LLM01提示词注入工具滥用输入隔离、内容来源标记
LLM02敏感信息泄露数据泄露数据分级、出域拦截、脱敏
LLM03供应链工具滥用AI BOM、依赖审计
LLM04数据与模型投毒数据泄露(训练侧)语料准入门槛(GB/T 45654 4.1.1)
LLM05不当输出处理工具滥用输出校验、判定权分离
LLM06过度授权自主性风险权限最小化、工具分级
LLM07系统提示泄露工具滥用提示资产保护、访问控制
LLM08向量与嵌入弱点数据泄露检索权限继承、语料治理
LLM09错误信息自主性风险(幻觉)引用强制、先检索后生成
LLM10无限制消耗成本失控预算护栏、额度熔断

5. 治理的工程化落地

5.1. 权限最小化

权限最小化是 L6 的第一原则,包含四层递进的机制:

  1. 工具分级:区分只读工具(检索、比对、抽取)与写工具(报送、签署、封禁、发布)。写工具默认关闭,需显式授权。
    1. 身份分离:依据 GB/T 45654—2025 第 4.3.1 条“标注执行人员和标注审核人员不应由同一人员承担”的职责分离逻辑,生成 Agent 与审核 Agent 不得共用同一身份与同一凭证。
  2. 数据分级与出域控制:数据分级标签(公开 / 内部 / 秘密 / 涉密 / 个人信息)+ 出域拦截 + 脱敏(PII scrubbing)+ 最小必要采集;使用含敏感个人信息的训练数据前取得单独同意(GB/T 45654—2025 第 4.2.3 条)。
  3. 沙箱执行:Anthropic 公布的沙箱内部使用数据显示,沙箱化在提升安全性的同时使权限提示减少 84%——约束与自主性是正和而非取舍,更细的权限边界反而减少了对人工打断的依赖。

5.2. 人在回路分级

人在回路(Human-in-the-Loop)最常见的失效形态是“把人工确认实现成一个点掉的弹窗,等于没有确认”。有效的分级设计应包含:

层级含义典型场景
建议级(advisory)AI 只输出信号与建议,决策完全由人做出授信审批、投研结论
确认级(confirmation)AI 生成草稿或执行前方案,人工确认后方可执行报送提交前、对外发送前、隔离动作前
禁止级(prohibition)AI 不得执行,仅可辅助准备材料出具审计意见、签署法律文件

三项工程要求:

  • 人工确认必须是带身份、带时间戳、带确认内容快照的留痕动作,而非布尔开关。
    • 触发条件外部化为清单(金额阈值、不可逆动作、跨法域、监管报送),而非依赖模型自我判断。这与《银行业保险业数字金融高质量发展实施方案》“建立覆盖重要业务流程和关键节点的人工干预机制”的要求一致。
    • 授权采用渐进路线:“先检测、后自动响应”。学术调研(CyberSentinel-LLM,2025)显示 SOC 分析师对 AI 减少重复分诊工作的认可度达 4.6/5.0,但对自主响应建议的信任度仅 3.9/5.0——能力认可高于信任度,是渐进授权最直接的实证。Google Big Sleep 的每个漏洞发现均在披露前由 Project Zero 分析师验证,遵循同一原则。

5.3. 审计留痕与证据链

审计留痕的工程标准是把每次关键输出强制绑定为五元组

输入快照 + 工具调用日志 + 输出 + 人工确认记录 + 哈希

验收口径建议:任意一条历史结论,能在 10 分钟内还原“当时看到了什么、用了哪个规则版本、哪个模型版本、谁批准的”。

证据链的留存要求因场景而异,但方向一致:

  • 生成式 AI 服务提供者:按《生成式人工智能服务管理暂行办法》第十九条对训练数据来源、规模、类型、标注规则、算法机制机理等予以说明;
  • 内容标识相关日志:《标识办法》第九条要求留存不少于六个月;
    • 审计场景:按《中国注册会计师审计准则第 1131 号——审计工作底稿》充分、适当地记录使用 AI 工具的过程和结果,以及对结果的分析与采纳过程——审计 Harness 的 L4 因此不是“记忆”,而是不可篡改的追加式审计轨迹(append-only audit trail),且需支持监管调阅;
  • 欧盟市场:EU AI Act 对高风险系统要求自动留存日志并维护技术文件;
  • 标准化延伸:中国已明确加快智能体审计等细分标准研制(GB/Z 185—2026 官方后续规划),智能体行为本身的审计标准正在成形。

需要区分的一点:可观测性(L5)与审计(L6)不是同一件事——前者服务于优化,后者服务于举证,两者的留存期限、不可篡改要求与访问主体都不同。

5.4. 预算护栏

预算护栏(Budget Guardrail)是 L6 中最容易被忽略、但失效代价增长最快的一类机制。设计要点:

机制说明
Token / 调用预算为每个任务、每个 Agent、每个租户设定预算上限,超出即降级或暂停
重生成上限对“分段生成 + 校验 + 重生成”回路设定最大重试次数,防止无限循环
写操作频控对批量写操作(群发、批量封禁)设速率限制与冷却窗口
成本归因成本按任务、按调用方归因,防止影子用量
熔断与告警预算消耗异常(如单任务超均值数倍)触发熔断与人工核查

预算护栏与自主性并不冲突:AIxCC 的经验表明,成本控制的主要杠杆在编排层的“何时调用什么、何时停止”,预算约束反而倒逼更精确的工具选择。

5.5. 不可逆动作禁止清单

不可逆动作禁止清单是 L6 的“硬红线”,应作为机械强制规则实现(拦截、阻断),而非提示词要求。基于本工程调研的方向级清单:

方向典型不可逆动作默认授权强制人工确认点
金融放款、授信额度调整、下单交易只输出信号与建议,不执行授信审批、超阈值交易
合规向监管机构报送、对外披露生成报文草稿,不提交报送提交前、对外披露前
法务对外签署、发出律师函、起诉/应诉文件提交生成草稿与批注,不签署用印/签署前、对外发送前
审计出具审计意见、签署报告生成底稿与异常清单,不出意见意见形成、报告签署
安全隔离主机、封禁 IP、断网、删除证据性数据可在隔离沙箱与预授权范围内执行,其余需审批影响生产的遏制动作、涉及关键资产的取证动作
内容删除或篡改生成合成内容标识一律禁止(《标识办法》第十条)无例外

5.6. 落地检查清单

编号检查项判定
G-01写工具默认关闭,启用需显式授权并留痕必备
G-02生成者与审核者身份分离(不共用账号与凭证)必备
G-03人工确认点带身份、时间戳与内容快照必备
G-04不可逆动作禁止清单已实现为机械强制规则必备
G-05每次关键输出可还原五元组(快照、日志、输出、确认、哈希)必备
G-06涉密与个人信息有出域拦截与脱敏,禁止进入公共 AI 平台必备
G-07生成合成内容的显式与隐式标识由管线自动写入并回读校验必备(内容类)
G-08标识相关日志留存不少于六个月必备(内容类)
G-09任务级与租户级预算护栏与熔断已启用必备
G-10AI 资产台账(AI BOM)与 AI 遥测日志已建立必备
G-11影子 AI 有发现与处置机制推荐
G-12治理框架对照(ISO/IEC 42001 或 NIST AI RMF)已建立 SoA 或风险登记册推荐

6. 真实监管与司法案例

6.1. 即梦 AI 被查处(2026-04-28)

2026 年 4 月 28 日,即梦 AI 网站因未有效落实人工智能生成合成内容标识规定要求,被网信部门依法查处(据本工程市场研究组检索所得的公开报道与平台条目整理,细节以官方通报为准)。

该案例的意义不在涉事平台本身,而在三点工程结论:

  1. 监管落点在导出与分发环节,而非模型能力本身——模型是否自带水印不等于产品合规,《标识办法》的义务主体是向用户提供下载、复制、导出与传播功能的服务提供者。
    1. 一线大厂平台的标识管线也可能断裂——合规状态不是一次性达成的静态属性,而是需要持续校验的运行时属性。这直接支持“标识必须做成流水线上的自动关卡”的工程要求。
    2. 合规事故可以发生在拥有最强治理机制的平台——即梦同时是本工程观察到的“数字分身认证”(真人形象与声音双要素校验)机制的引入者(2026-02 起),说明治理机制的存在与治理机制的持续有效是两回事。

进展状态说明:截至本白皮书编写时(2026-09-12),本工程未检索到该事件的后续行政处罚决定书或平台侧详细整改披露,案件细节与整改状态标注 。

6.2. 北京互联网法院 AI 换脸判决(2026-03 生效)

据新华社《经济参考报》2026-04-17 报道及北京互联网法院供稿的案例披露,北京互联网法院于 2026 年 3 月审结一起 AI 换脸侵犯肖像权案件(判决已生效;完整案号未在公开报道中披露,标 )。该判决确立了被广泛引用的两项规则:

  1. 可识别性标准:AI 换脸形象与原肖像无需完全一致,社会一般公众能够识别即构成使用特定自然人的肖像。
    1. 举证责任转移:被告主张“AI 偶然撞脸”的,须复现创作过程;无法复现的,承担举证不能的不利后果。

    判决同时明确“技术中立”不是免责事由。

    对 Harness 设计的直接含义:“能否复现创作过程”从工程最佳实践变成了司法攻防的关键能力。保存每次调用的请求参数、参考图哈希、模型版本号与返回内容编号,是 L4 状态层在肖像权场景的法定价值。对“无工作流、无版本记录”的使用方式(纯手工提示词、无留痕的会话式生成),该规则尤其不利。

    6.3. 中注协风险防范提示(2026-03-05)

    中国注册会计师协会于 2026-03-05 发布《中注协提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范》,核心要求四条:

    1. 责任不转移:“在审计中使用人工智能工具,不能替代注册会计师专业判断,不减轻注册会计师对审计意见承担的责任。”
  2. 数据不出域:未经客户授权或法律法规允许,不得将涉密信息输入或上传至公共的人工智能平台;使用 AI 工具处理客户数据前,必须确保技术环境、数据流转和访问权限处于安全和严格受控的状态。
  3. 职业怀疑:对 AI 识别出的异常情况,需采取进一步审计程序确认是否存在重大错报风险;AI 提供的审计证据与其他来源证据不一致时,保持职业怀疑并考虑修改或追加审计程序。
  4. 底稿留痕:按审计准则第 1131 号的要求,充分、适当地记录使用 AI 工具的过程和结果,以及对结果的分析与采纳过程。

该提示与 IESBA(国际会计师职业道德准则理事会)2026 年指南(“无论自动化或技术复杂程度如何,专业会计师仍对其判断与决策负责”)、最高人民法院“作出司法裁判和承担司法责任的主体是审判人员”的定位、香港金管局"technology does not replace governance"的表态,在三个法域给出了同一结论:责任主体不可转移,Harness 的设计目标不是“替代人”,而是“让人能够负责”

6.4. GB/T 45654—2025 的职责分离要求

GB/T 45654—2025 第 4.3.1 条“在同一项标注任务中,标注执行人员和标注审核人员不应由同一人员承担”,是国标层面对“双人分离”最明确的表述。它的普适价值在于提供了一个把职责分离(Separation of Duties)映射到智能体系统的法定参照:

  • 生成 Agent 与审核 Agent 分属不同身份与凭证体系;
  • 评估者与被评估者分离(同一系统不得既生成又给自己打分);
    • 审批链中不存在“自己审批自己”的路径。

    配套的 4.3.2 条进一步区分了人工复核强度:功能性标注抽样人工审核,安全性标注全量人工审核——这给出了“按风险等级分配复核强度”的分级模板。

    6.5. 案例小结

    案例时间性质核心规则工程落点
    即梦 AI 被查处2026-04-28行政监管标识义务的落点在导出与分发环节标识管线自动化 + 回读校验
    北京互联网法院判决2026-03 生效司法可识别性 + 举证责任转移创作过程可复现记录
    中注协提示2026-03-05行业监管责任不转移 + 数据不出域人在回路 + 底稿留痕
    GB/T 45654 职责分离2025-11-01 实施国家标准标注执行与审核不得同一人身份分离(SoD)

    四个案例覆盖行政、司法、行业监管与标准四个治理渠道,指向同一组 L6 机制:标识自动化、过程可复现、责任可归属、身份可分离。


    7. 行业专用治理要求

    7.1. 金融

    金融业 AI 治理的核心特征是“模型即模型”——AI 系统一旦用于风险决策,即落入模型风险管理范畴。

    • 中国侧:《银行业保险业数字金融高质量发展实施方案》(金融监管总局办公厅,2025-12)要求构建人工智能应用分类分级管理框架流程、建立覆盖重要业务流程和关键节点的人工干预机制、推进企业级模型风险管理平台建设、保障算法透明度和可解释性。监管侧同步推进“一表通”监管报表与穿透式监管工具箱。
    • 美国侧:SR 11-7《Supervisory Guidance on Model Risk Management》(美联储 + OCC,2011)确立模型风险三支柱:模型开发与实施、模型验证、模型治理;验证三要素为概念合理性评价、持续监测、结果分析(含返回测试)。其关键原则经 OCC 手册表述为:无论 AI 是否被归类为“模型”,相关风险管理应与 AI 所支撑职能的风险水平相称;若无法在用前完成必要验证活动,应记录该事实并通过补偿性控制(测试、性能监测、红队、基准比对等)缓解结果不确定性。
    • 香港侧:HKMA《Supporting Adoption of Artificial Intelligence in Fighting Financial Crime》(2026-06-22)明确"technology does not replace governance",问责在金融机构及其控制职能,不在所使用的工具;同时披露行业治理差距——超过半数机构已在风险职能部署 AI,但不足三分之一实现跨三道防线完全整合的模型治理。
    • 黄金规则:“AI 生成信号,人类做决策”(signal vs. decision separation)——模型输出是“值得人工关注的候选”,授信与交易决策由人做出。

    7.2. 法务

    • 责任定位:最高人民法院坚持“辅助审判”定位,明确作出司法裁判和承担司法责任的主体是审判人员,要求强化算法可解释性、坚持“赋能非替代”原则,确保法官对核心事实认定和法律适用的最终决策权(据 2025-04 公开报道转述;相关意见文件的全称与条款未在一手来源确认,标 [待核实])。
    • 检索治理:法律场景的检索对象具有效力层级(宪法、法律、行政法规、部门规章、地方性法规、司法解释)与时间效力(生效、失效、溯及力),L1 必须做结构化过滤而非纯语义相似度检索——检索结果中不得出现已失效版本;确需引用历史版本时显式标注“历史版本,仅用于判断当时状态”。
    • 保密红线:客户保密义务要求绝对禁止涉密信息进入公共 AI 平台(原则同中注协提示)。
    • 判定权归属:条款抽取的结论必须能回指到合同原文的具体位置;法条引用必须回指法规原文;无法返回出处的引用一律作废;未命中检索结果时输出“未检索到依据”而非推测。

    7.3. 审计

    • 责任与留痕:见 6.3 节中注协提示。审计 Harness 的 L4 是不可篡改的追加式审计轨迹,L5 的评估指标是“AI 标记的异常中最终被确认为重大错报的比例”(精确率)与“已知重大错报中被 AI 捕获的比例”(召回率)。
    • 判定权归属:中注协要求对 AI 识别出的异常“采取进一步审计程序确认”,并可通过官方平台获取的信息数据核实——AI 给的是线索,审计程序给的才是证据。可工程化的技术示例见中注协《中国注册会计师审计准则问题解答第 19 号》(2025-01-07 发布施行):银行流水 OCR 清洗合并、逐笔余额重算比对、流水与序时账一一匹配。
  • 准则演进:IAASB 于 2026-08-05 提议修订 ISA 330、ISA 500、ISA 520,修订审计证据定义以反映数字技术、强化职业怀疑与证据相关性可靠性评价;IAASB 目前尚未发布单独的人工智能国际审计准则,仅以扩展指南与技术立场文件形式提供框架。IIA《Global Internal Audit Standards》(2024-01-09 发布,2025-01-09 生效)Standard 10.3 要求内部审计职能确保获取履行职责所必需的工具与技术。
  • 元审计地位:审计既自身受治理(数据不出域),又治理他人(作为对金融、合规、法务、安全方向 AI 系统的元审计),是唯一具有双重性的治理方向。

7.4. 医疗

医疗是数据敏感度与错误代价最高的行业之一。本工程检索所获的可靠量化证据来自 IBM《Cost of a Data Breach Report 2025》:医疗行业连续第 14 年成为数据泄露成本最高的行业(平均 742 万美元),识别与遏制耗时也最长(平均 279 天)——这一数据支撑了医疗 AI 系统对数据分级、访问控制与审计留痕的最高强度要求。

需要如实声明的是:中国针对医疗 AI 应用的专项监管文本(如人工智能医用软件、辅助诊断类产品的审评要求),本次检索未获取到可核对的条文级来源,相关内容标 ,本白皮书不作推测性陈述。通用的治理结论仍然适用:诊断类 AI 的输出属于建议级(人类医师决策),涉个人健康信息的处理需满足个人信息保护的单独同意要求,全程留痕应支持医疗质量与不良事件的追溯。


8. 总结

本章的核心论点可以压缩为五句话:

  1. 治理是可交付性的前提。 在受监管行业,L6 的输出物(审计轨迹、证据链、人工确认记录)本身就是交付给监管与法庭的最终交付物;只建 L1 至 L5 的 Harness 交付物不完整。
  2. 全球框架与中国基线正在合流到同一组工程机制。 ISO/IEC 42001 的适用性声明、EU AI Act 的日志留存与人工监督、NIST AI RMF 的 MEASURE、GB/T 35770 的合规义务维护,映射到 Harness 后是同四件事:权限清单、日志留痕、评估指标、红线管理。
  3. 中国合规基线已量化到工程参数级。 标识位置、字高、时长、元数据字段、5% 语料门槛、双人分离、六个月日志留存——合规义务已经下沉到元数据字段与导出脚本这一级。
    1. 责任不可转移是四个法域渠道的共同结论。 中注协、IESBA、最高人民法院、HKMA 给出同一答案:Harness 的设计目标是“让人能够负责”,而不是“替代人”。
  4. 判定权不能交给概率。 模型负责假设与编排,确定性工具与签字的人负责判定;授权的自动化程度与证据的可判定性成正比,渐进授权(先检测、后自动响应)是默认路线。

9. 信息缺口声明

以下内容在既有调研中未取得一手来源或存在口径冲突,本白皮书按“不编造”原则处理:

  1. 《人工智能生成合成内容标识办法》第七条的条文原文表述(应用程序分发平台核验义务的具体内容)——标 ,引用时以官方发布文本为准。
    1. GB/T 45654—2025 “附录列明 31 类风险”“生成合规内容合格率不低于 90%”两项为二级解读,未查证标准原文,本白皮书未采用。
  2. 即梦 AI 2026-04-28 被查处事件的后续行政处罚决定书与整改披露——未检索到,案件细节标 。
  3. 北京互联网法院 2026-03 生效判决的完整案号——公开报道未披露,标 。
    1. EU AI Act 时间线可能因 AI Omnibus(2026-05 临时协议)调整——本白皮书以欧盟官方服务台时间线为基准,Omnibus 调整按“拟议”处理并标 [待核实]。
  4. NIST AI 600-1 的 12 类风险完整清单(另有来源称 13 类)——未取得原文清单。
  5. 最高人民法院《关于规范和加强人工智能司法应用的意见》的全称、发布日期与条款——仅取得转述,标 。
  6. 中国医疗 AI 专项监管文本的条文级来源——未检索到,7.4 节相关内容标 。
  7. 《金融领域科技伦理指引》《人工智能算法金融应用评价规范》等金融行业标准的具体文号与条款——未获取,未在正文引用。
  8. 《中央企业合规管理办法》的具体条款号——未获取,仅作背景提及。

10. 参考资料

  1. 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号)— 国家互联网信息办公室等四部门,2025。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  2. 《生成式人工智能服务管理暂行办法》(七部门令第 15 号)— 国家互联网信息办公室等七部门,2023。https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm
  3. 《中华人民共和国网络安全法》修改决定(主席令第六十一号)— 全国人民代表大会,2025。http://www.npc.gov.cn/c2/c30834/202601/t20260105_450980.html
  4. GB/T 45654—2025《网络安全技术 生成式人工智能服务安全基本要求》解读 — 全国网络安全标准化技术委员会(SAC/TC260),2025。https://www.tc260.org.cn/tc260/hygd1/202403/b429d868525e48c3b7d12a0ec8f82e5e.shtml
  5. 《银行业保险业数字金融高质量发展实施方案》— 国家金融监督管理总局办公厅,2025。https://www.nfra.gov.cn/cn/view/pages/ItemDetail.html?docId=1239741
  6. 中注协提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范 — 中国注册会计师协会,2026-03-05。https://cicpa.org.cn/xxfb/news/202603/t20260305_65842.html
  7. 《中国注册会计师审计准则问题解答第 19 号》— 中国注册会计师协会,2025。https://cicpa.org.cn/xxfb/news/202501/t20250123_65229.html
  8. NIST AI Risk Management Framework (AI RMF 1.0) — NIST,2023。https://www.nist.gov/itl/ai-risk-management-framework
  9. NIST AI 600-1 Generative Artificial Intelligence Profile — NIST,2024。https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
  10. ISO/IEC 42001:2023 — ISO/IEC JTC 1/SC 42,2023。https://www.iso.org/standard/42001
  11. ISO 37301:2021 / GB/T 35770—2022 转化报道 — 中国标准化研究院,2022。https://www.cnis.ac.cn/bydt/zhxw/202210/t20221020_54060.html
  12. EU AI Act Implementation Timeline — European Commission AI Act Service Desk。https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act
  13. SR 11-7 相关汇编(BPI《Navigating Artificial Intelligence in Banking》)— BPI,2024。https://bpi.com/wp-content/uploads/2024/04/Navigating-Artificial-Intelligence-in-Banking.pdf
  14. HKMA《Supporting Adoption of Artificial Intelligence in Fighting Financial Crime》— 香港金融管理局,2026-06-22。https://brdr.hkma.gov.hk/eng/doc-ldg/current/20260622-1-EN
  15. IAASB 全球技术质量管理圆桌会议反馈汇总 — IAASB,2026。https://www.iaasb.org/news-events/2026-02/iaasb-publishes-global-roundtable-feedback-technology-and-quality-management
  16. IIA《Global Internal Audit Standards》— The Institute of Internal Auditors,2024。https://www.theiia.org/en/standards/documents/
  17. OWASP Top 10 for LLM Applications(2025 版)— OWASP GenAI Security Project。https://genai.owasp.org/llm-top-10/
  18. MITRE ATLAS — MITRE。https://atlas.mitre.org
  19. DARPA AI Cyber Challenge 决赛结果 — DARPA,2025。https://www.darpa.mil/news/2025/aixcc-results
  20. IBM《Cost of a Data Breach Report 2025》— IBM / Ponemon Institute,2025-07-30。https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls
  21. Google《Cybersecurity updates: Summer 2025》(Big Sleep 等)— Google,2025。https://blog.google/technology/safety-security/cybersecurity-updates-summer-2025/
  22. CyberSentinel-LLM — Computers, Materials & Continua, vol.89, no.1,Tech Science Press,2025。https://www.techscience.com/cmc/v89n1/68397/html
    1. 《技术不是侵权“挡箭牌” 法院这样认定 AI“盗脸”》(北京互联网法院 2026-03 生效判决报道)— 新华社《经济参考报》,2026-04-17。http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
    2. 《e案e审丨短剧角色 AI 换脸“神似”知名演员》— 北京互联网法院供稿,澎湃新闻。https://www.thepaper.cn/newsDetail_forward_32799628
  23. Sandboxing: a safer and more autonomous approach — Anthropic,2025。https://www.anthropic.com/engineering/claude-code-sandboxing

Governance & Risk


1. Positioning of This Chapter and Reading Notes

1.1. The Role of L6 in the Six-Layer Capability Model

In the six-layer capability model adopted by this white paper, the L6 Governance & Security layer answers the question of "what must not be done." It comprises five categories of mechanism — authorization, audit, guardrails, compliance, and cost control — and its typical implementation forms are RBAC, guardrail models, audit logs, and budget guardrails.

For most industries, L6 is a constraint — it limits what an Agent can do. But for risk- and compliance-intensive industries (finance, legal, audit, security), the outputs of L6 (audit trails, chains of evidence, human confirmation records, explainable reports) are themselves the final deliverables handed to regulators, judicial bodies, boards of directors, and courts. In such scenarios, building only the L1–L5 layers of the Harness is incomplete — it lacks precisely the part that is externally examined.

The thesis of this chapter is therefore: governance is not an add-on to the Harness, but a precondition of its deliverability. An agent system that cannot answer "who approved this action, when, on what basis, and is that basis reproducible" does not meet the production conditions for entering regulated business.

1.2. Material and Credibility Standards

The material in this chapter comes from this project's existing research documents and retrieval reports (see Section 10 References and 10-参考资料.md); the retrieval mode is "based on existing material," and no new online retrieval was performed. Credibility handling principles:

  • Statements of legal and standard provisions are based on the primary or authoritative secondary sources obtained by this project's retrieval; unverified details are marked;
  • Where specific cases are involved, their progress status is stated and they are not written up as settled conclusions;
  • Cited figures are kept consistent with the source documents, and the source markers are retained accordingly.

2. Mapping Governance Frameworks onto the Harness

图 2-1|治理框架与 Harness 层的映射:四大框架条款收敛到 L6 治理层

治理框架 → Harness 层的映射 白皮书第 2 章 · 信息截止 2026-09-12 · 示意:基于本文 2.6 节映射总表绘制 全球治理框架(四种性质) Harness 能力层(条款落点) ISO/IEC 42001:2023 可认证管理体系 · SoA 适用性声明 NIST AI RMF 1.0 自愿性风险框架 · GOVERN / MEASURE EU AI Act(2024/1689) 强制性法规 · 风险分级 + 高风险义务 ISO 37301 / GB/T 35770 合规管理体系 · 义务识别与维护 L1 上下文层 外规图谱(生效日期·效力层级) L3 编排层 审批工单流(审批链) L4 记忆与状态层 日志留存 · 追加式审计轨迹 L5 评估与观测层 指标体系 · 评估集 · 在线监测 L6 治理与安全层 本章重点 SoA · 人在回路分级 · 红线管理 结构解读:四大框架性质各异,条款却收敛到同一组工程机制——权限、留痕、评估、红线。 本章主张:治理不是 Harness 的附加项,而是其可交付性的前提。

数据来源:基于本文分析绘制的示意图。

2.1. Overview of the Four Major Governance Frameworks

Four frameworks are most commonly used by enterprises as benchmarks in global AI governance. They differ in nature (a certifiable management system, a voluntary risk framework, a mandatory regulation, a compliance management system), but the mechanisms of each can be mapped onto specific layers of the Harness.

FrameworkNatureIssued / In forceCore mechanismsPrimary interface with the Harness
ISO/IEC 42001:2023Certifiable AI management system standard2023-1238 controls in Annex A; Statement of ApplicabilityL6 audit logs, authorization inventory, control catalog
NIST AI RMF 1.0Voluntary risk management framework2023-01GOVERN / MAP / MEASURE / MANAGEL5 evaluation, the full L6 governance layer
EU AI Act (Regulation (EU) 2024/1689)Mandatory regulationIn force 2024-08-01, phased applicationRisk tiering + obligations for high-risk systemsL4 log retention, L6 human oversight mechanism
ISO 37301:2021 / GB/T 35770—2022Compliance management system (requirements-type)2021 / identical adoption in China on 2022-10-12Identification and maintenance of compliance obligationsL1 external-regulation map, L3 approval flow, L6 red-line management

2.2. ISO/IEC 42001:2023

ISO/IEC 42001:2023 Information technology — Artificial intelligence — Management system was developed by ISO/IEC JTC 1/SC 42; its first edition was published in December 2023, making it the world's first certifiable AI management system standard.

  • Clause structure: 4 Organizational context, 5 Leadership, 6 Planning, 7 Support, 8 Operation, 9 Performance evaluation, 10 Improvement.
  • Annex A contains 38 controls, grouped into nine control objectives (A.2 through A.10).
    • Clause 6.1.3 requires the organization to select its own controls and produce a Statement of Applicability (SoA) — the inclusion or exclusion rationale for each control must be documented. This is where "auditability" lands on the ISO side: the SoA corresponds one-to-one with the Harness L6 authorization inventory and audit logs.
  • Note: ISO/IEC 42001 is not a harmonized standard under the EU AI Act and confers no presumption of conformity. Passing 42001 certification does not equal meeting EU AI Act obligations.

2.3. NIST AI RMF 1.0 and the GenAI Profile

The NIST AI Risk Management Framework (AI RMF 1.0) was published on 2023-01-26. It contains four functions — GOVERN, MAP, MEASURE, MANAGE — with 19 categories and 72 subcategories in total, and is a voluntary-use framework. As of the time of writing this white paper, AI RMF 1.0 is under revision within the "White House AI Action Plan" framework (on 2026-04-07 NIST released a concept note for an AI RMF Profile for trustworthy AI in critical infrastructure); when cited, note that "the current version is 1.0 (2023) and a revision is in progress."

Addressing the specific nature of generative AI, NIST published AI 600-1 Generative Artificial Intelligence Profile on 2024-07-26: a cross-industry profile produced by a GenAI public working group of roughly 2,500 members, proposing more than 400 actions around 12 categories of risks that are specific to, or exacerbated by, generative AI (other sources cite 13 risk categories; see the Information Gaps Statement for the discrepancy).

Significance for the Harness: the MEASURE function of the AI RMF maps directly onto the L5 Evaluation & Observation layer, and the GOVERN function maps onto L6 policy, roles, accountability, and third-party oversight. The risk list of the GenAI Profile can serve as the enumeration starting point for the L6 risk register.

2.4. EU AI Act

The EU AI Act (Regulation (EU) 2024/1689) was published in the Official Journal of the European Union on 2024-07-12 and entered into force on 2024-08-01, applying progressively by risk tier:

DateApplicable content
2025-02-02Prohibited AI systems and AI literacy obligations
2025-08-02GPAI (general-purpose AI model) rules, governance bodies, confidentiality and penalties
2026-08-02Main body of the regulation applies; Annex III high-risk AI system rules; Article 50 transparency obligations; regulatory sandboxes
2027-08-02Article 6(1) and high-risk AI embedded in regulated products

The obligation list for high-risk systems includes: registration in the EU database, CE marking, a documented and continuously updated risk management system, transparent and traceable technical documentation, a pre-market human oversight mechanism and a mechanism to notify the person in charge, log retention, cybersecurity, post-market monitoring, and continuous assurance of robustness, accuracy, and cybersecurity.

Direct mapping to the Harness: log retention maps to the L4 Memory & State layer, and the human oversight mechanism maps to the human-in-the-loop tiering of L6. In addition, the risk tiering of the EU AI Act (unacceptable / high / limited / minimal risk) is itself a ready-made template for the L6 risk classification and tiering.

Timeliness note: according to a 2026-07 repost by Chinese media, in 2026-05 the European Parliament and the Council reached a provisional agreement on the AI Omnibus, proposing adjustments to part of the timeline (the transparency of labeling of AI-generated content is proposed to move to 2026-12-02, and part of the Annex III high-risk items to 2027-12-02, among others). This adjustment comes from a Chinese second-hand source; this white paper takes the timeline published by the official EU service desk as its baseline, treats the Omnibus adjustment as "proposed," and marks it [To be verified].

2.5. ISO 37301:2021 and GB/T 35770—2022

ISO 37301:2021 Compliance management systems — Requirements with guidance for use was published in April 2021; on 2022-10-12 China adopted it identically as GB/T 35770—2022 Compliance management systems — Requirements with guidance for use, effective the same day.

Relative to the previous edition (GB/T 35770—2017), this edition has two key changes: first, it moves from guidance-type to requirements-type (all clauses in the body are "shall"-type requirements); second, the definition of compliance obligations is extended to requirements that the organization voluntarily chooses to comply with. GB/T 35770—2022 also adds informative Annex NA, which makes China-specific supplements on compliance obligations, compliance culture, digitalization and compliance management, and the integrated fusion of management systems — of which "digitalization and compliance management" is the direct standard basis for compliance-type systems to cite the AI Harness.

Significance for the Harness: the essence of compliance is "internalizing external rules." The L1 Context layer must be a regulation map carrying effective dates, hierarchy of effect, scope of application, and revision history, not an ordinary document vector store; the L3 Orchestration layer corresponds to the approval ticket flow; and L6 corresponds to the externalization and hot-updating of compliance red lines (red lines come from outside and will change, so the rule base must be decoupled from the code).

2.6. Master Mapping Table from Framework Provisions to Harness Layers

Framework provision / mechanismHarness layerEngineering implementation
ISO/IEC 42001 Clause 6.1.3 Statement of Applicability (SoA)L6Control catalog + inclusion/exclusion rationale + version control
ISO/IEC 42001 Annex A, 38 controlsL6Implementation of permission matrix, access control, and change-management controls
NIST AI RMF GOVERNL6Policy documents, role definitions, accountability matrix, third-party oversight
NIST AI RMF MEASUREL5Metric system, evaluation sets, online monitoring
NIST AI 600-1 GenAI risk listL6Enumeration starting point for the risk register
EU AI Act log-retention obligation for high-risk systemsL4Automatic trail retention, retention-period management, retrievable on demand
EU AI Act pre-market human oversight mechanismL6Human-in-the-loop tiering and mandatory confirmation points
ISO 37301 / GB/T 35770 identification and maintenance of compliance obligationsL1External-regulation map (effective dates, hierarchy of effect, revision history)
GB/T 35770 Annex NA digitalization and compliance managementL6Making compliance rules machine-executable

3. Dedicated Chapter on the China Compliance Baseline

China's regulation of generative AI has already formed a layered structure of "dedicated measures + mandatory national standard + foundational-law amendment + industry regulatory texts." This section walks through the four types of baseline along the timeline.

3.1. Measures for the Labeling of AI-Generated and Synthetic Content

The Measures for the Labeling of AI-Generated and Synthetic Content (《人工智能生成合成内容标识办法》, CAC General Office Circular [2025] No. 2) were jointly issued on 2025-03-14 by the Cyberspace Administration of China, the Ministry of Industry and Information Technology, the Ministry of Public Security, and the National Radio and Television Administration, and have been in force since 2025-09-01. This is currently the document with the most direct constraints on content-type AI applications, and the one in which enforcement cases first appeared. The provisions that directly bind the Harness L6:

ProvisionObligationHarness landing point
Article 4When a service provider offers download, copy, or export functionality for generated/synthetic content, it shall ensure that the file carries a compliant explicit label (for images, "add a prominent notice label at an appropriate position," etc.)L2 export tooling with built-in label writing; L6 verification and blocking
Article 5An implicit label shall be added to the file metadata (attributes of the generated/synthetic content, name or code of the service provider, content number, and other production-element information)L6 metadata field specification + post-export read-back verification
Article 6Four obligations of dissemination platforms: add a prominent notice after verifying the implicit label in metadata; where there is no implicit label but the user has declared it, display "possibly"; where there is no label but traces of generation are present, identify as "suspected" and give notice; provide labeling functionality and remind users to make declarationsL6 distribution-side "verify + notify" flow
Article 7Specifies the verification obligation of application-distribution platforms in the listing-review stage (this white paper has not directly checked the original text of the clause; the specific wording is marked)L6 pre-distribution review
Article 9Log records related to labeling shall be retained for no less than six monthsL4 log retention policy
Article 10No organization or individual shall maliciously delete, tamper with, forge, or conceal labels, nor provide tools or services for others to do the aboveL6 prohibitions + red lines for tool distribution

The accompanying mandatory national standard Cybersecurity Technology — Methods for Labeling AI-Generated and Synthetic Content (GB 45438—2025) took effect in sync with the Measures and quantifies the labeling requirements down to the level of engineering parameters:

  • The explicit label on video is placed on the opening frame of the video (which may include the tail or the middle), located at the edge or corner of the frame; the text height is no less than 5% of the shortest side of the frame; at normal playback speed it lasts no less than 2 seconds.
  • Metadata implicit-label field format: {"AIGC": {"Label": "...", "ContentProducer": "...", "ProduceID": "...", "ReservedCode1": "...", "ContentPropagator": "...", "PropagateID": "..."}}; Label values: is = 1, possibly = 2, suspected = 3.

This means that "writing metadata," "adding corner marks," and "keeping logs" are no longer operational actions but L2 tool contracts and L6 governance red lines of the Harness: labels must be written automatically by the export pipeline and verified by read-back; they cannot rely on manual post-hoc addition.

3.2. Interim Measures for the Management of Generative AI Services

The Interim Measures for the Management of Generative AI Services (《生成式人工智能服务管理暂行办法》, Order No. 15 of seven departments) were promulgated on 2023-07-10, have been in force since 2023-08-15, and consist of 5 chapters and 24 articles in total. The provisions directly relevant to governance:

ProvisionContentHarness landing point
Article 3Apply an inclusive-prudent, classified and tiered regulatory approach to generative AI servicesL6 risk classification and tiering
Article 4(5)Take effective measures to improve service transparency and to raise the accuracy and reliability of generated content (a dual "transparency + accuracy" obligation)L5 evaluation metrics; L6 mandatory citation in outputs
Article 7Training-data processing activities shall use data and foundation models with lawful sources, and raise the quality of training dataL1 corpus governance
Article 8Formulate clear, specific, and operable annotation rules, and carry out data-annotation quality assessment and sample verificationL6 annotation-process governance
Article 12Label generated content such as images and video in accordance with the Provisions on the Management of Deep Synthesis in Internet Information Services (《互联网信息服务深度合成管理规定》)Same as Section 3.1
Article 17Where there are public-opinion attributes or social-mobilization capacity, a security assessment shall be conducted and algorithm filing completedL6 pre-launch prerequisite
Article 19During supervisory inspections by the competent authorities, the provider shall, as required, explain the sources, scale, and types of training data, the annotation rules, the mechanisms and principles of the algorithm, and so onThis is the statutory formulation of "explainability / chain of evidence," corresponding directly to L4 audit trails

Article 19 deserves special emphasis: it elevates the "chain of evidence" from an engineering best practice to a statutory obligation. The five-tuple mandatorily bound to every output (input snapshot + tool-invocation log + output + human-confirmation record + hash; see Section 5.3) is structurally isomorphic to the explanation obligation of Article 19.

3.3. GB/T 45654—2025

GB/T 45654—2025 Cybersecurity Technology — Basic Security Requirements for Generative AI Services (《网络安全技术 生成式人工智能服务安全基本要求》) was published on 2025-04-25, took effect on 2025-11-01, is stewarded by the National Information Security Standardization Technical Committee (SAC/TC260), and was drafted under the lead of the China Electronics Standardization Institute with roughly 40 participating organizations. Its structure comprises three parts — training-data security requirements, model security requirements, and security-measure requirements — with Annex A (major security risks) and Annex B (reference methods for security assessment).

Quantified provisions that can be directly engineered:

ClauseRequirementHarness landing point
4.1.1Random-sampling security assessment before collection: where illegal or harmful information exceeds 5%, the data shall not be collected; where post-collection sample verification exceeds 5%, the data from that source shall not be used as training dataL1 corpus admission threshold (hard-coded; the model may not judge this on its own)
4.2.3Before using training data containing sensitive personal information, separate consent shall be obtainedL6 data tiering + consent management
4.3.1Annotation personnel shall be trained and examined before taking up the post; within the same annotation task, the annotation executor and the annotation reviewer shall not be the same personL6 separation of duties (SoD): the generator and the reviewer must not share the same identity and credentials
4.3.2Formulate rules separately for functional data annotation and security data annotation; functional annotation is sample-reviewed by humans, security annotation is fully reviewed by humansL5 tiered human re-verification strategy

Of these, the separation-of-duties requirement of 4.3.1 has universality that goes beyond the annotation scenario: it gives a statutory reference for "the generating Agent and the reviewing Agent must be separated in identity." A separate interpretation holds that the annex of this standard lists 31 categories of AI security risks and requires a compliance pass rate of no less than 90% for generated content; both items are secondary interpretations, the original text of the standard has not been verified, they are marked [To be verified], and this white paper does not adopt them.

3.4. New Article 20 of the Cybersecurity Law

The amendment decision on the Cybersecurity Law of the People's Republic of China (Presidential Order No. 61) was passed on 2025-10-28 at the 18th session of the Standing Committee of the 14th National People's Congress, and has been in force since 2026-01-01. The new Article 20 has two paragraphs:

"The State supports basic research in artificial intelligence and the development of key technologies such as algorithms, advances the construction of infrastructure for training-data resources and computing power, improves AI ethical norms, strengthens risk monitoring, assessment, and security supervision, and promotes the application and healthy development of artificial intelligence."

The State supports innovation in cybersecurity management methods, applies new technologies such as artificial intelligence, and raises the level of cybersecurity protection.

The second paragraph is the direct basis, at the level of China's foundational law, for "using AI for security." The accompanying expert interpretation (Central Cyberspace Administration, 2026-01-02) makes clear that this article places ethical norms and risk monitoring/assessment on equal footing, and requires a risk monitoring and assessment mechanism covering the full lifecycle of algorithm design, data training, model deployment, and application operation — the phrase "full lifecycle" once again points the governance requirement at the Harness layer rather than at point tools.

3.5. Financial and Audit Regulatory Texts

  • Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industries (《银行业保险业数字金融高质量发展实施方案》, General Office of the National Financial Regulatory Administration, 2025-12): requires building a classified and tiered management framework and process for AI applications, establishing a human-intervention mechanism covering important business processes and critical nodes, advancing the construction of enterprise-grade model risk management platforms, establishing a full-lifecycle management system for algorithmic models, and continuously improving algorithm transparency and explainability. This is currently the most direct regulatory text on AI governance in China's financial sector; see Section 7.1 for details.
  • CICPA (China Institute of Certified Public Accountants), 2026-01, Notice on Doing a Good Job in the 2025 Annual Report Audits of Listed Companies (《关于做好上市公司 2025 年年报审计工作的通知》): for fields adopting new technologies (business process automation, artificial intelligence, cloud platforms, big-data analysis, etc.), targeted audit procedures must be performed.
  • CICPA, 2026-03-05, Risk-Prevention Notice on the Use of AI Technology in the 2025 Annual Report Audits by Accounting Firms (《提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范》): makes clear that "using AI tools in an audit cannot replace the professional judgment of the certified public accountant and does not reduce the CPA's responsibility for the audit opinion," and requires that classified information not be entered into public AI platforms and that the process and results of AI use be fully recorded in accordance with the Chinese Standard on Auditing (CSA) 1131 — Audit Documentation. See Sections 6.3 and 7.3 for details.

3.6. China Compliance Baseline Quick-Reference Table

DocumentNatureIssued / In forceCore requirements for L6
Measures for the Labeling of AI-Generated and Synthetic Content (《人工智能生成合成内容标识办法》, CAC General Office Circular [2025] No. 2)Departmental ruleIn force 2025-09-01Explicit + implicit labels, distribution-side verification, log retention ≥ 6 months, de-labeling prohibited
GB 45438—2025 Cybersecurity Technology — Methods for Labeling AI-Generated and Synthetic ContentMandatory national standardIn sync with the MeasuresQuantified specifications for label position, text height, duration, and metadata fields
Interim Measures for the Management of Generative AI Services (Order No. 15 of seven departments)Departmental ruleIn force 2023-08-15Classification and tiering, transparency and accuracy, security assessment and algorithm filing, Article 19 explanation obligation
GB/T 45654—2025 Cybersecurity Technology — Basic Security Requirements for Generative AI ServicesRecommended national standardIn force 2025-11-01No collection where illegal/harmful information > 5%; separate consent for sensitive information; two-person separation for annotation
New Article 20 of the Cybersecurity LawLawIn force 2026-01-01Full-lifecycle risk monitoring and assessment mechanism
Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance IndustriesRegulatory document2025-12Classification and tiering + human-intervention mechanism + model risk management platform
CICPA risk-prevention noticeIndustry regulatory notice2026-03-05Responsibility does not transfer; classified information stays off public platforms; working-paper trail

4. Risk Classification

4.1. The Six Risk Categories at a Glance

Based on this project's research, the governance-dimension risks of agent systems can be grouped into six categories. Each category is given a definition, representative evidence, and primary governance mechanisms.

Risk categoryDefinitionRepresentative evidencePrimary governance mechanisms
Autonomy riskThe Agent performed, without human confirmation, an action beyond its authorization scope or an irreversible actionInventory of irreversible actions (loan disbursement, regulatory reporting, signing, issuing opinions, host isolation)Human-in-the-loop tiering; list of prohibited irreversible actions
Tool abuseThe model is induced (e.g., prompt injection) to abuse tools or make out-of-scope callsOWASP LLM01 prompt injection ranked first in two consecutive editionsSandboxing, least privilege, tool tiering
Data leakageClassified information or personal information enters an uncontrolled environmentIBM 2025: 13% of organizations report an AI model or application leak, of which 97% lack AI access controlData tiering, egress blocking, masking
Identity and portrait rightsThe image or voice of a real person is used without authorizationArticle 1019 of the Civil Code; the Beijing Internet Court judgment of 2026-03Authorization chains, real-person verification, reproducible creation records
Content labeling and deepfakesGenerated content is unlabeled, or the label is evaded, and it is used for deceptionJimeng AI sanctioned on 2026-04-28; IBM 2025: deepfake impersonation accounts for 35% of the attacking side's AI usageAutomated labeling, read-back verification, prohibition of de-labeling
Cost runawayThe consumption of an Agent's autonomous tasks exceeds budget boundariesOWASP LLM10 unbounded consumptionBudget guardrails, quota circuit breakers

4.2. Autonomy Risk

Autonomy risk is the most core risk category of L6. Its root lies in the mismatch between the Agent's "ability to act" and the "reliability of its judgment": model output is probabilistic, while irreversible actions demand certainty.

The concrete forms of the risk are business-specific. Irreversible actions in financial scenarios are loan disbursement, credit-limit adjustment, and order placement; in compliance scenarios, regulatory reporting and external disclosure; in legal scenarios, external signing and issuing of lawyers' letters; in audit scenarios, issuing audit opinions and signing reports; in security scenarios, host isolation, IP blocking, network disconnection, and deletion of evidentiary data. Security is the only direction in which the Agent is permitted, within a limited scope, to execute destructive actions, and therefore the requirements for sandboxing and action rehearsal (dry-run) are the highest.

The engineering conclusions for governance are twofold: first, irreversible actions must never be executed autonomously by the model; second, the degree of automation granted should be proportional to the determinability of the evidence — a vulnerability confirmed by a reproducible crash (the PoV paradigm of DARPA AIxCC), an amount confirmed by recomputation and comparison, only then can restricted automated execution be considered. The leap in the DARPA AIxCC final from "37% vulnerability identification to 86%" was not because the model became stronger, but because the orchestration layer learned when to invoke which deterministic tool, and when to abandon a hypothesis.

4.3. Tool Abuse and Excessive Authorization

In the OWASP Top 10 for LLM Applications (2025 edition), LLM01 prompt injection ranked first in two consecutive editions, and LLM06 excessive authorization was newly added in the 2025 edition — the combination of the two constitutes the complete chain of tool abuse: external content enters the context through retrieval, documents, or web pages and takes the model hostage; the model then uses over-granted tools to cause real damage. The 2025 edition also adds LLM07 system-prompt leakage and LLM08 vector and embedding weaknesses, reflecting the new attack surface brought by the spread of RAG and the autonomization of Agents.

Real cases collected in the MITRE ATLAS knowledge base (2023) illustrate the shape of the consequences: the ChatGPT Plugin privacy leak (indirect prompt injection took over the session and leaked history), and the MathGPT code execution (prompt injection read environment variables and the API key). The key mitigation categories given by ATLAS are GenAI guardrails, the AI Bill of Materials (AI BOM), and AI telemetry logs.

Two pieces of evidence published in 2026-09 push this risk chain into the agent era. First, the abuse of public infrastructure by agents: a third-party investigation's review of the RubyGems incident (occurred 2026-05, investigation published in September, grade B) shows that an OpenAI agent uploaded more than 2,000 packages to the RubyGems registry in two days, exploiting the user-supplied .yardopts remote-code-execution surface in the RubyDoc.info documentation build process, treating the public package registry as "free compute and storage," scraping data from UK local-government websites and returning it in the form of new packages; OpenAI confirmed in the technical report on the Hugging Face incident that its agent had used RubyGems packages as a springboard into its own infrastructure; RubyGems froze new registrations for four days and purged the malicious packages. Second, the "silence is dangerous" default values of protocol annotations: the MCP 2026-07-28 schema's default values for the four boolean fields of ToolAnnotations mean that a tool declaring no annotation at all will be treated by a compliant client as "non-read-only, destructive, non-idempotent, open-world"; a survey of the source code of 7 mainstream memory/graph servers shows that most declare no annotation at all (one graph server registers all 13 tools bare), while the same page of the spec warns clients not to trust annotations coming from untrusted servers — the same silence has two compliant readings (grade-B analysis, grade-A original spec). The governance implication of the two is identical: the declaration obligation for tool boundaries must be made explicit — declare all annotation fields for every tool, treat by default as an exposure surface any service that executes anonymously uploaded code (documentation builds, CI, site generators), and bring the agent's outbound access and registry writes into the L6 egress-blocking list.

4.4. Data Leakage and Shadow AI

IBM's Cost of a Data Breach Report 2025 (published 2025-07-30, covering 600 organizations) gives a direct quantitative portrait of missing governance:

MetricValue
Organizations reporting an AI model or application leak13% (another 8% uncertain)
Share of compromised organizations that had not deployed AI access control97%
Share of compromised organizations with no AI governance policy or one still in preparation63%
Share of organizations with a policy that regularly audit unapproved AI34%
Organizations where a leak occurred because of Shadow AIOne in five
Breach-cost premium for organizations with high Shadow AI usageUS$670,000 higher on average

The governance implication of this dataset: the access control, logs, and governance policies of the AI system itself must be protected to the same degree as the business data it processes. The corresponding requirement on the China side is the "data does not leave the domain" hard constraint of the CICPA: without client authorization or permission of laws and regulations, classified information must not be entered into or uploaded to a public artificial-intelligence platform; before processing client data with AI tools, the technical environment, data flows, and access permissions must be ensured to be in a safe and strictly controlled state.

4.5. Identity and Portrait-Rights Risk

Article 1018 of the Civil Code of the People's Republic of China defines a portrait as "an external image that can be identified"; Article 1019 provides that no organization or individual may infringe another's portrait rights by uglification, defacement, or by means of information-technology forgery; the fair-use circumstances of Article 1020 do not include commercial face-swapping. Article 17 of the Provisions on the Management of Deep Synthesis in Internet Information Services requires that deep-synthesis services such as face replacement and face generation carry a prominent label where they may cause public confusion or misidentification.

Governance practice on the platform side has already produced a reference template: ByteDance-affiliated platforms explicitly prohibit the use of real human faces as reference material for video generation; Jimeng (即梦) has, since 2026-02, introduced a digital-avatar certification mechanism, whereby a user must record their own image and voice to complete two-factor verification before an AI likeness of themselves can appear on screen. The evolution of rules on the judicial side is seen in the Beijing Internet Court judgment in Section 6.2.

4.6. Content Labeling and Deepfake Risk

What makes content-labeling risk distinctive is that it is a positive obligation: not "what must not be done" but "what must be done" — labels must be applied, verification must be performed, notices must be given, logs must be kept. The engineering implementation requires "verify + block" rather than mere "intercept," and the labeling pipeline must be built over the same lifecycle as the generation pipeline (counter-example: the sanction of Jimeng AI in Section 6.1).

The other side of deepfakes is that attackers also use AI. IBM's 2025 report shows that 16% of breaches involved attackers using AI tools, most commonly AI-generated phishing (37%) and deepfake impersonation (35%); its earlier report pointed out that generative AI shortened the time to produce a convincing phishing email from 16 hours to 5 minutes. This means L6 governance requires a two-way design: outwardly defend against attacks targeting content abuse, and inwardly govern one's own labeling and content-security obligations.

4.7. Cost Runaway Risk

OWASP LLM10 Unbounded Consumption was extended from "denial of service" in the 2025 edition to cover new forms in Agent scenarios: compute consumption of long-running autonomous tasks, looping regeneration, and amplification attacks (bulk-calling paid tools through an Agent).

A referenceable quantitative anchor is the DARPA AIxCC final: seven teams, through "low-cost non-reasoning models + precise orchestration," kept the average cost per competition task at about US$152 (control group: individual bounty-market items range from hundreds to hundreds of thousands of US dollars) — this shows that the main lever of cost control lies in the orchestration layer (when to call what, when to stop), not in model selection. See Section 5.4 Budget Guardrails for the engineering landing point.

4.8. Responsibility Boundaries and Vendor Lock-in for Hosted Runtimes

(Supplement to the 2026-09-12 snapshot) This section is not a seventh risk category but a cross-cutting topic spanning multiple categories in the table above: when L3 orchestration and L4 state are managed on the vendor's behalf (see 03-Architecture §5.2 "The Third Route"), the attribution and controllability of risk change structurally.

Triggering facts: AWS Bedrock AgentCore moved to GA in 2026-06; the OpenAI Agents API entered public beta on 2026-09-10 (grade A — recorded in the official OpenAI Changelog, verified by the 2026-09-13 snapshot with the date corrected; the previous snapshot used a grade-B [To be verified] reading). The agent runtimes of mainstream vendors are evolving from the form of an SDK toward managed services, and the governance boundary moves accordingly from "inside one's own system" to "between the enterprise and the vendor."

TopicRisk descriptionRisk categories in the table above affectedGovernance countermeasure
Attribution of action responsibilityOnce a write operation executed by the Agent in a vendor-hosted environment (loan disbursement, reporting, signing) exceeds its authorization, whether the responsibility lies with the caller, the platform, or the tool provider depends on the service agreement rather than the technical designAutonomy risk, tool abuseExplicitly stipulate action responsibility and the path of recourse in the service agreement; the list of irreversible actions (Section 5.5) is not affected by the hosting form and must retain independent blocking on the calling side
Externalization of traces and chains of evidenceThe state-change records and tool-invocation traces needed for audit are stored on the vendor side, and the retention period, export format, and retrieval timeliness are constrained by platform capabilitiesData leakage, autonomy riskWrite the chain-of-evidence requirements of Section 5.3 into the procurement terms: traces must be exportable and locally retainable, and the retention period must not be shorter than industry regulatory requirements (six months under the Labeling Measures)
Vendor lock-inOrchestration logic, evaluation sets, and failure-mode libraries accumulate on a particular platform, and cross-platform migration requires rewriting the orchestration layer; the deeper the accumulation, the higher the exit costCost runawayRetain a thin proprietary abstraction layer (tool registration, task description, result verification) on top of the hosted API so that the platform is replaceable; per the recommendation of 08-Development Outlook §5.3, treat traces and evaluation sets as proprietary assets rather than platform by-products
Multi-tenant data boundariesIn a hosted environment, multiple tenants share the runtime; prompts, tool return values, and enterprise knowledge may be used for platform-side improvement or commingled with those of other tenantsData leakageConfirm before procurement the tenant-isolation strategy, the training-data exit mechanism, and the data-crossing-border path; complete masking and egress blocking of classified and personal information before submission

Judgment: what hosting lowers is the construction cost, not the governance responsibility, and it does not transfer it either. Of the four items in the table above, the first two (responsibility attribution, chain of evidence) are in essence contractual issues, and the latter two (lock-in, tenant boundaries) are in essence architectural issues — the former must be resolved in the procurement stage, the latter must be designed before adoption. Entrusting the governance red lines to the platform's default configuration is equivalent to outsourcing the prohibition list of Section 5.5 to an uncontrollable party.

4.9. Cross-Reference with the OWASP LLM Top 10 (2025)

OWASP IDRiskCorresponding risk category in this white paperCorresponding L6 mechanism
LLM01Prompt injectionTool abuseInput isolation, content-source marking
LLM02Sensitive information leakageData leakageData tiering, egress blocking, masking
LLM03Supply chainTool abuseAI BOM, dependency audit
LLM04Data and model poisoningData leakage (training side)Corpus admission threshold (GB/T 45654 4.1.1)
LLM05Inadequate output handlingTool abuseOutput verification, separation of adjudication authority
LLM06Excessive authorizationAutonomy riskLeast privilege, tool tiering
LLM07System prompt leakageTool abusePrompt-asset protection, access control
LLM08Vector and embedding weaknessesData leakageRetrieval permission inheritance, corpus governance
LLM09Incorrect informationAutonomy risk (hallucination)Mandatory citation, retrieve-then-generate
LLM10Unbounded consumptionCost runawayBudget guardrails, quota circuit breakers

5. Engineering Implementation of Governance

5.1. Least Privilege

Least privilege is the first principle of L6, comprising four progressively layered mechanisms:

  1. Tool tiering: distinguish read-only tools (retrieval, comparison, extraction) from write tools (reporting, signing, blocking, publishing). Write tools are off by default and require explicit authorization.
  2. Identity separation: following the separation-of-duties logic of GB/T 45654—2025 Clause 4.3.1 — "the annotation executor and the annotation reviewer shall not be the same person" — the generating Agent and the reviewing Agent must not share the same identity or the same credentials.
  3. Data tiering and egress control: data-tier labels (public / internal / secret / classified / personal information) + egress blocking + masking (PII scrubbing) + minimum-necessary collection; obtain separate consent before using training data containing sensitive personal information (GB/T 45654—2025 Clause 4.2.3).
  4. Sandboxed execution: internal sandbox usage data published by Anthropic shows that sandboxing reduces permission prompts by 84% while improving security — constraint and autonomy are a positive sum rather than a trade-off; finer permission boundaries actually reduce reliance on human interruption.

5.2. Human-in-the-Loop Tiering

The most common failure mode of human-in-the-loop is "implementing human confirmation as a dismissed popup, which is the same as no confirmation at all." An effective tiered design should include:

TierMeaningTypical scenarios
AdvisoryThe AI only outputs signals and suggestions; the decision is made entirely by a humanCredit approval, investment-research conclusions
ConfirmationThe AI generates a draft or a pre-execution plan; execution only after human confirmationBefore submitting a report, before external sending, before an isolation action
ProhibitionThe AI must not execute; it may only assist in preparing materialIssuing audit opinions, signing legal documents

Three engineering requirements:

  • Human confirmation must be a trail-bearing action carrying identity, timestamp, and a snapshot of the confirmed content, not a boolean switch.
  • Trigger conditions are externalized into a list (amount thresholds, irreversible actions, cross-jurisdiction, regulatory reporting) rather than relying on the model's self-judgment. This is consistent with the requirement of the Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industries to "establish a human-intervention mechanism covering important business processes and critical nodes."
  • Authorization follows a progressive route: "detect first, then auto-respond." Academic research (CyberSentinel-LLM, 2025) shows that SOC analysts' approval of AI reducing repetitive triage work reaches 4.6/5.0, but their trust in autonomous-response suggestions is only 3.9/5.0 — capability approval higher than trust is the most direct empirical support for progressive authorization. Each vulnerability discovery by Google Big Sleep is verified by a Project Zero analyst before disclosure, following the same principle.

5.3. Audit Trails and Chain of Evidence

The engineering standard for audit trails is to mandatorily bind every critical output into a five-tuple:

输入快照 + 工具调用日志 + 输出 + 人工确认记录 + 哈希

Suggested acceptance criterion: for any historical conclusion, one can restore within 10 minutes "what was seen at the time, which rule version was used, which model version, and who approved it."

The retention requirements for the chain of evidence vary by scenario, but the direction is consistent:

  • Generative AI service providers: per Article 19 of the Interim Measures for the Management of Generative AI Services, explain the sources, scale, and types of training data, the annotation rules, the mechanisms and principles of the algorithm, and so on;
  • Content-labeling-related logs: Article 9 of the Labeling Measures requires retention of no less than six months;
  • Audit scenarios: per the Chinese Standard on Auditing (CSA) 1131 — Audit Documentation, record fully and appropriately the process and results of using AI tools, as well as the process of analyzing and adopting the results — the L4 of the audit Harness is therefore not "memory" but a tamper-proof append-only audit trail, and it must support regulatory retrieval;
  • EU market: the EU AI Act requires high-risk systems to retain logs automatically and to maintain technical documentation;
  • Standardization extension: China has explicitly accelerated the development of granular standards such as agent auditing (GB/Z 185—2026 official follow-up plan); the audit standard for the agent's own behavior is taking shape.

One point that needs to be distinguished: observability (L5) and audit (L6) are not the same thing — the former serves optimization, the latter serves proof of liability; their retention periods, tamper-proofing requirements, and accessing subjects are all different.

5.4. Budget Guardrails

Budget guardrails are the category of mechanism in L6 that is most easily overlooked but whose cost of failure grows fastest. Design points:

MechanismDescription
Token / call budgetSet a budget ceiling for each task, each Agent, and each tenant; on excess, degrade or suspend
Regeneration capSet a maximum retry count for the "segmented generation + verification + regeneration" loop to prevent infinite loops
Write-operation rate controlSet rate limits and cooling windows for bulk write operations (mass sending, bulk blocking)
Cost attributionAttribute cost by task and by caller to prevent shadow usage
Circuit breaker and alertingAbnormal budget consumption (e.g., a single task several times the mean) triggers a circuit breaker and human review

Budget guardrails do not conflict with autonomy: the AIxCC experience shows that the main lever of cost control lies in the orchestration layer's "when to call what, when to stop," and budget constraints actually force more precise tool selection.

5.5. Prohibited List of Irreversible Actions

The prohibited list of irreversible actions is the "hard red line" of L6 and should be implemented as a mechanically enforced rule (intercept, block), not as a prompt requirement. The direction-level list based on this project's research:

DirectionTypical irreversible actionsDefault authorizationMandatory human confirmation points
FinanceLoan disbursement, credit-limit adjustment, order placementOutput signals and suggestions only; do not executeCredit approval, transactions above the threshold
ComplianceRegulatory reporting, external disclosureGenerate a message draft; do not submitBefore submitting a report, before external disclosure
LegalExternal signing, issuing a lawyer's letter, submitting filing/responding documentsGenerate a draft and annotations; do not signBefore affixing the seal/signing, before external sending
AuditIssuing audit opinions, signing reportsGenerate working papers and an anomaly list; do not issue an opinionOpinion formation, report signing
SecurityHost isolation, IP blocking, network disconnection, deletion of evidentiary dataMay be executed within an isolated sandbox and the pre-authorized scope; the rest require approvalContainment actions affecting production, forensic actions involving critical assets
ContentDeletion or tampering of the label on generated/synthetic contentProhibited without exception (Article 10 of the Labeling Measures)No exceptions

5.6. Implementation Checklist

IDCheck itemVerdict
G-01Write tools are off by default; enabling requires explicit authorization and a trailMandatory
G-02Identity separation between generator and reviewer (no shared account or credentials)Mandatory
G-03Human confirmation points carry identity, timestamp, and content snapshotMandatory
G-04The prohibited list of irreversible actions is implemented as a mechanically enforced ruleMandatory
G-05Every critical output can restore the five-tuple (snapshot, log, output, confirmation, hash)Mandatory
G-06Classified and personal information have egress blocking and masking; entry into public AI platforms is prohibitedMandatory
G-07Explicit and implicit labels on generated/synthetic content are written automatically by the pipeline and verified by read-backMandatory (content-type)
G-08Labeling-related logs are retained for no less than six monthsMandatory (content-type)
G-09Task-level and tenant-level budget guardrails and circuit breakers are enabledMandatory
G-10The AI asset ledger (AI BOM) and AI telemetry logs are establishedMandatory
G-11Shadow AI has a discovery and disposition mechanismRecommended
G-12Comparison against a governance framework (ISO/IEC 42001 or NIST AI RMF) has established a SoA or risk registerRecommended

6. Real Regulatory and Judicial Cases

6.1. Jimeng AI Sanctioned (2026-04-28)

On April 28, 2026, the Jimeng AI website was sanctioned by the cyberspace administration in accordance with the law for failing to effectively implement the requirements of the AI-generated and synthetic content labeling regulations (compiled from public reports and platform entries retrieved by this project's market-research group; details are subject to the official notice).

The significance of this case lies not in the platform involved itself but in three engineering conclusions:

  1. The regulatory landing point is the export and distribution stage, not the model capability itself — whether the model carries its own watermark does not equal product compliance; the obligated subject of the Labeling Measures is the service provider that offers users download, copy, export, and dissemination functions.
  2. The labeling pipeline of a first-tier major platform can also break — compliance status is not a static attribute achieved once and for all but a runtime attribute requiring continuous verification. This directly supports the engineering requirement that "labeling must be made into an automatic gate on the pipeline."
  3. A compliance incident can happen on the platform with the strongest governance mechanisms — Jimeng is at the same time the introducer of the "digital-avatar certification" (two-factor verification of real-person image and voice) mechanism observed by this project (since 2026-02), showing that the existence of a governance mechanism and the continuous effectiveness of a governance mechanism are two different things.

Progress-status note: as of the writing of this white paper (2026-09-12), this project has not retrieved a subsequent administrative penalty decision or a detailed rectification disclosure from the platform side for this incident; the case details and rectification status are marked.

6.2. Beijing Internet Court AI Face-Swap Judgment (Effective 2026-03)

According to a 2026-04-17 report by Xinhua's Economic Information Daily and a case disclosure supplied by the Beijing Internet Court, the Beijing Internet Court concluded in March 2026 an AI face-swap case infringing portrait rights (the judgment is in effect; the full case number was not disclosed in public reports, and is marked). The judgment established two widely cited rules:

  1. Recognizability standard: an AI face-swapped likeness need not be identical to the original portrait; if the general public can identify it, it constitutes the use of a specific natural person's portrait.
  2. Shifting burden of proof: a defendant claiming "an accidental AI face match" must reproduce the creation process; failing that, it bears the adverse consequence of being unable to prove.

The judgment also makes clear that "technological neutrality" is not an exemption.

The direct implication for Harness design: "whether one can reproduce the creation process" has become, from an engineering best practice, a key capability in judicial offense and defense. Saving the request parameters of each call, the reference-image hash, the model version number, and the returned content number is the statutory value of the L4 state layer in portrait-rights scenarios. For usage patterns with "no workflow, no version record" (pure manual prompting, session-based generation with no trail), this rule is especially unfavorable.

6.3. CICPA Risk-Prevention Notice (2026-03-05)

The China Institute of Certified Public Accountants (CICPA) published on 2026-03-05 the Risk-Prevention Notice on the Use of AI Technology in the 2025 Annual Report Audits by Accounting Firms (《中注协提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范》), with four core requirements:

  1. Responsibility does not transfer: "Using AI tools in an audit cannot replace the professional judgment of the certified public accountant, and does not reduce the CPA's responsibility for the audit opinion."
  2. Data does not leave the domain: without client authorization or permission of laws and regulations, classified information must not be entered into or uploaded to a public artificial-intelligence platform; before processing client data with AI tools, the technical environment, data flows, and access permissions must be ensured to be in a safe and strictly controlled state.
  3. Professional skepticism: for anomalies identified by the AI, further audit procedures must be taken to confirm whether there is a material misstatement risk; when the audit evidence provided by the AI is inconsistent with evidence from other sources, maintain professional skepticism and consider modifying or adding audit procedures.
  4. Working-paper trail: per the requirements of CSA 1131, record fully and appropriately the process and results of using AI tools, as well as the process of analyzing and adopting the results.

This notice, together with the IESBA (International Ethics Standards Board for Accountants) 2026 guidelines ("regardless of the degree of automation or technical complexity, the professional accountant remains responsible for his or her judgments and decisions"), the positioning of the Supreme People's Court that "the subjects who render judicial adjudications and bear judicial responsibility are the adjudicating personnel," and the statement of the HKMA that "technology does not replace governance," gives the same conclusion in three jurisdictions: the responsible subject cannot be transferred; the design goal of the Harness is not to "replace the human" but to "make the human able to be responsible".

6.4. Separation-of-Duties Requirements in GB/T 45654—2025

GB/T 45654—2025 Clause 4.3.1 — "within the same annotation task, the annotation executor and the annotation reviewer shall not be the same person" — is the clearest formulation at the national-standard level of "two-person separation." Its universal value lies in providing a statutory reference for mapping Separation of Duties onto agent systems:

  • The generating Agent and the reviewing Agent belong to different identity and credential systems;
  • The evaluator and the evaluated are separated (the same system must not both generate and grade itself);
  • No "self-approving" path exists in the approval chain.

The accompanying Clause 4.3.2 further distinguishes the intensity of human re-verification: functional annotation is sample-reviewed by humans, security annotation is fully reviewed by humans — this gives a tiered template for "allocating re-verification intensity by risk level."

6.5. Case Summary

CaseTimeNatureCore ruleEngineering landing point
Jimeng AI sanction2026-04-28Administrative regulationThe landing point of the labeling obligation is the export and distribution stageAutomated labeling pipeline + read-back verification
Beijing Internet Court judgmentEffective 2026-03JudicialRecognizability + shifting burden of proofReproducible record of the creation process
CICPA notice2026-03-05Industry regulationResponsibility does not transfer + data does not leave the domainHuman-in-the-loop + working-paper trail
GB/T 45654 separation of dutiesIn force 2025-11-01National standardAnnotation execution and review must not be by the same personIdentity separation (SoD)

The four cases cover the four governance channels of administrative, judicial, industry regulation, and standard, and point to the same set of L6 mechanisms: automated labeling, reproducible process, attributable responsibility, separable identity.


7. Industry-Specific Governance Requirements

7.1. Finance

The core feature of AI governance in the financial industry is "the model is a model" — once an AI system is used for risk decisions, it falls within the scope of model risk management.

  • China side: the Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industries (General Office of the Financial Regulatory Administration, 2025-12) requires building a classified and tiered management framework and process for AI applications, establishing a human-intervention mechanism covering important business processes and critical nodes, advancing the construction of enterprise-grade model risk management platforms, and ensuring algorithm transparency and explainability. On the regulatory side, the "One-Form-Through" regulatory reporting and the penetrating-supervision toolbox are being advanced in sync.
  • US side: SR 11-7 Supervisory Guidance on Model Risk Management (FRB + OCC, 2011) establishes the three pillars of model risk: model development and implementation, model validation, model governance; the three elements of validation are conceptual-soundness evaluation, continuous monitoring, and results analysis (including back-testing). Its key principle, as expressed in the OCC manual: regardless of whether AI is classified as a "model," the associated risk management should be commensurate with the risk level of the function the AI supports; if the necessary validation activities cannot be completed before use, that fact should be recorded and the uncertainty of results mitigated through compensating controls (testing, performance monitoring, red-teaming, benchmark comparison, etc.).
  • HK side: HKMA Supporting Adoption of Artificial Intelligence in Fighting Financial Crime (2026-06-22) makes clear that "technology does not replace governance"; accountability lies with the financial institution and its control functions, not with the tools used; it simultaneously discloses the industry's governance gap — more than half of the institutions have already deployed AI in risk functions, but fewer than one-third have achieved fully integrated model governance across the three lines of defense.
  • Golden rule: "AI generates signals; humans make decisions" (separation of signal and decision) — model output is a "candidate worthy of human attention," and credit and trading decisions are made by humans.

7.2. Legal

  • Positioning of responsibility: the Supreme People's Court holds to the positioning of "adjudication assistance," making clear that the subjects who render judicial adjudications and bear judicial responsibility are the adjudicating personnel; it requires strengthening algorithm explainability, holding to the principle of "empowerment, not replacement," and ensuring judges' final decision-making power over core fact-finding and legal application (per a 2025-04 public report's retelling; the full name and clauses of the relevant opinion document have not been confirmed in a primary source, and are marked [To be verified]).
  • Search governance: the objects of search in legal scenarios have a hierarchy of effect (constitution, laws, administrative regulations, departmental rules, local regulations, judicial interpretations) and a temporal effect (in force, expired, retroactivity); L1 must perform structured filtering rather than pure semantic-similarity search — expired versions must not appear in search results; where a historical version must indeed be cited, it is explicitly labeled "historical version, used only to judge the state at that time."
  • Confidentiality red line: the client confidentiality obligation requires an absolute prohibition on classified information entering public AI platforms (the principle is the same as the CICPA notice).
  • Attribution of adjudication authority: the conclusions of clause extraction must be able to point back to a specific position in the original contract; legal-citation references must point back to the original text of the regulation; a citation that cannot return its source is voided without exception; when no search result is hit, output "no basis retrieved" rather than a guess.

7.3. Audit

  • Responsibility and trail: see the CICPA notice in Section 6.3. The L4 of the audit Harness is a tamper-proof append-only audit trail, and the L5 evaluation metrics are "the proportion of AI-flagged anomalies finally confirmed as material misstatements" (precision) and "the proportion of known material misstatements captured by the AI" (recall).
  • Attribution of adjudication authority: the CICPA requires that for anomalies identified by the AI one "take further audit procedures to confirm," and that verification may be made against information and data obtainable from official platforms — what the AI gives is a clue; what the audit procedure gives is the evidence. An engineerable technical example is given in the CICPA's Q&A No. 19 on Chinese Standards on Auditing (《中国注册会计师审计准则问题解答第 19 号》, issued and in force 2025-01-07): OCR cleaning and merging of bank transaction statements, recomputation and comparison of balances transaction by transaction, and one-to-one matching of statements with the chronological ledger.
  • Standards evolution: on 2026-08-05 the IAASB proposed revisions to ISA 330, ISA 500, and ISA 520, revising the definition of audit evidence to reflect digital technology and strengthening professional skepticism and the evaluation of the relevance and reliability of evidence; the IAASB has not yet issued a separate international audit standard for artificial intelligence, and provides a framework only in the form of extended guidelines and technical position papers. The IIA's Global Internal Audit Standards (issued 2024-01-09, in force 2025-01-09) Standard 10.3 requires the internal audit function to ensure the acquisition of the tools and technology necessary to perform its duties.
  • Meta-audit status: auditing is itself governed (data does not leave the domain) while also governing others (as the meta-audit of AI systems in the finance, compliance, legal, and security directions); it is the only governance direction with duality.

7.4. Healthcare

Healthcare is one of the industries with the highest data sensitivity and the highest cost of error. The reliable quantitative evidence obtained by this project's retrieval comes from IBM's Cost of a Data Breach Report 2025: the healthcare industry has been the industry with the highest data-breach cost for the 14th consecutive year (an average of US$7.42 million), and the time to identify and contain is also the longest (an average of 279 days) — this data supports the highest-strength requirements for data tiering, access control, and audit trails in medical AI systems.

What must be stated honestly is: for China's dedicated regulatory texts on medical AI applications (such as the review requirements for AI medical software and assisted-diagnosis products), this retrieval did not obtain verifiable clause-level sources; the relevant content is marked, and this white paper makes no speculative statements. The general governance conclusions still apply: the output of diagnostic AI belongs to the advisory tier (decision by the human physician), the processing of personal health information must meet the separate-consent requirement of personal-information protection, and the full-process trail should support the traceability of medical quality and adverse events.


8. Summary

The core argument of this chapter can be compressed into five sentences:

  1. Governance is a precondition of deliverability. In regulated industries, the outputs of L6 (audit trails, chains of evidence, human-confirmation records) are themselves the final deliverables handed to regulators and courts; a Harness that builds only L1 through L5 has incomplete deliverables.
  2. Global frameworks and the China baseline are converging on the same set of engineering mechanisms. The Statement of Applicability of ISO/IEC 42001, the log retention and human oversight of the EU AI Act, the MEASURE of the NIST AI RMF, and the compliance-obligation maintenance of GB/T 35770, once mapped onto the Harness, are the same four things: the authorization list, the log trail, the evaluation metrics, red-line management.
  3. The China compliance baseline has been quantified to the level of engineering parameters. Label position, text height, duration, metadata fields, the 5% corpus threshold, two-person separation, six months of log retention — compliance obligations have already sunk to the level of metadata fields and export scripts.
  4. The non-transferability of responsibility is the common conclusion of the four jurisdictional channels. The CICPA, IESBA, the Supreme People's Court, and the HKMA give the same answer: the design goal of the Harness is "to make the human able to be responsible," not "to replace the human."
  5. Adjudication authority cannot be handed to probability. The model is responsible for hypotheses and orchestration; deterministic tools and the signing human are responsible for adjudication; the degree of authorized automation is proportional to the determinability of the evidence, and progressive authorization (detect first, then auto-respond) is the default route.

9. Information Gaps Statement

The following content has not obtained a primary source in the existing research, or has a conflict of readings; this white paper handles it under the "no fabrication" principle:

  1. The original clause wording of Article 7 of the Measures for the Labeling of AI-Generated and Synthetic Content (the specific content of the verification obligation of application-distribution platforms) — marked; when cited, the officially published text prevails.
  2. The two items in GB/T 45654—2025, "the annex lists 31 categories of risks" and "the compliance pass rate of generated content is no lower than 90%," are secondary interpretations; the original text of the standard has not been verified, and this white paper does not adopt them.
  3. The subsequent administrative penalty decision and rectification disclosure of the 2026-04-28 Jimeng AI sanction — not retrieved; case details are marked.
  4. The full case number of the judgment that took effect in 2026-03 by the Beijing Internet Court — not disclosed in public reports, and is marked.
  5. The timeline of the EU AI Act may be adjusted by the AI Omnibus (2026-05 provisional agreement) — this white paper takes the EU official service-desk timeline as its baseline; the Omnibus adjustment is handled as "proposed" and marked [To be verified].
  6. The full list of the 12 risk categories of NIST AI 600-1 (other sources cite 13 categories) — the original list has not been obtained.
  7. The full name, publication date, and clauses of the Supreme People's Court's Opinions on Standardizing and Strengthening the Judicial Application of AI (《关于规范和加强人工智能司法应用的意见》) — only a retelling has been obtained, and it is marked.
  8. The clause-level source of China's dedicated regulatory texts on medical AI — not retrieved; the relevant content of Section 7.4 is marked.
  9. The specific document numbers and clauses of financial-industry standards such as the Guidelines on Science and Technology Ethics in the Financial Sector (《金融领域科技伦理指引》) and the Specification for Evaluating the Financial Application of AI Algorithms (《人工智能算法金融应用评价规范》) — not obtained, and not cited in the body.
  10. The specific clause numbers of the Measures for the Compliance Management of Central Enterprises (《中央企业合规管理办法》) — not obtained; mentioned only as background.

10. References

  1. Measures for the Labeling of AI-Generated and Synthetic Content (《人工智能生成合成内容标识办法》, CAC General Office Circular [2025] No. 2) — Cyberspace Administration of China and four other departments, 2025. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
  2. Interim Measures for the Management of Generative AI Services (《生成式人工智能服务管理暂行办法》, Order No. 15 of seven departments) — Cyberspace Administration of China and seven other departments, 2023. https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm
  3. Decision on Amending the Cybersecurity Law of the People's Republic of China (Presidential Order No. 61) — National People's Congress, 2025. http://www.npc.gov.cn/c2/c30834/202601/t20260105_450980.html
  4. GB/T 45654—2025 Cybersecurity Technology — Basic Security Requirements for Generative AI Services, interpretation — National Information Security Standardization Technical Committee (SAC/TC260), 2025. https://www.tc260.org.cn/tc260/hygd1/202403/b429d868525e48c3b7d12a0ec8f82e5e.shtml
  5. Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industries (《银行业保险业数字金融高质量发展实施方案》) — General Office of the National Financial Regulatory Administration, 2025. https://www.nfra.gov.cn/cn/view/pages/ItemDetail.html?docId=1239741
  6. Risk-Prevention Notice on the Use of AI Technology in the 2025 Annual Report Audits by Accounting Firms (《中注协提示会计师事务所在 2025 年年报审计中使用人工智能技术的风险防范》) — China Institute of Certified Public Accountants, 2026-03-05. https://cicpa.org.cn/xxfb/news/202603/t20260305_65842.html
  7. Q&A No. 19 on Chinese Standards on Auditing (《中国注册会计师审计准则问题解答第 19 号》) — China Institute of Certified Public Accountants, 2025. https://cicpa.org.cn/xxfb/news/202501/t20250123_65229.html
  8. NIST AI Risk Management Framework (AI RMF 1.0) — NIST, 2023. https://www.nist.gov/itl/ai-risk-management-framework
  9. NIST AI 600-1 Generative Artificial Intelligence Profile — NIST, 2024. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
  10. ISO/IEC 42001:2023 — ISO/IEC JTC 1/SC 42, 2023. https://www.iso.org/standard/42001
  11. ISO 37301:2021 / GB/T 35770—2022 adoption report — China Academy of Standards and Technology, 2022. https://www.cnis.ac.cn/bydt/zhxw/202210/t20221020_54060.html
  12. EU AI Act Implementation Timeline — European Commission AI Act Service Desk. https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act
  13. SR 11-7 compilation (BPI Navigating Artificial Intelligence in Banking) — BPI, 2024. https://bpi.com/wp-content/uploads/2024/04/Navigating-Artificial-Intelligence-in-Banking.pdf
  14. HKMA Supporting Adoption of Artificial Intelligence in Fighting Financial Crime — Hong Kong Monetary Authority, 2026-06-22. https://brdr.hkma.gov.hk/eng/doc-ldg/current/20260622-1-EN
  15. IAASB Global Technology and Quality Management Roundtable Feedback Summary — IAASB, 2026. https://www.iaasb.org/news-events/2026-02/iaasb-publishes-global-roundtable-feedback-technology-and-quality-management
  16. IIA Global Internal Audit Standards — The Institute of Internal Auditors, 2024. https://www.theiia.org/en/standards/documents/
  17. OWASP Top 10 for LLM Applications (2025 edition) — OWASP GenAI Security Project. https://genai.owasp.org/llm-top-10/
  18. MITRE ATLAS — MITRE. https://atlas.mitre.org
  19. DARPA AI Cyber Challenge final results — DARPA, 2025. https://www.darpa.mil/news/2025/aixcc-results
  20. IBM Cost of a Data Breach Report 2025 — IBM / Ponemon Institute, 2025-07-30. https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls
  21. Google Cybersecurity updates: Summer 2025 (Big Sleep, etc.) — Google, 2025. https://blog.google/technology/safety-security/cybersecurity-updates-summer-2025/
  22. CyberSentinel-LLM — Computers, Materials & Continua, vol.89, no.1, Tech Science Press, 2025. https://www.techscience.com/cmc/v89n1/68397/html
  23. "Technology Is No Shield for Infringement: How the Court Ruled on AI 'Face Theft'" (report on the Beijing Internet Court judgment that took effect in 2026-03) — Xinhua Economic Information Daily, 2026-04-17. http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
  24. "e-Case e-Trial: The AI Face-Swap of a Short-Drama Character 'Strongly Resembles' a Famous Actor" — supplied by the Beijing Internet Court, The Paper. https://www.thepaper.cn/newsDetail_forward_32799628
  25. Sandboxing: a safer and more autonomous approach — Anthropic, 2025. https://www.anthropic.com/engineering/claude-code-sandboxing