Finance —— 金融风险与决策支持


覆盖场景:智能投研、信用与风控、量化研究、财报分析、反洗钱(AML)与金融犯罪合规。

取证说明:本文所引案例与数据均标注来源层级(A 级:监管或企业官方原文;B 级:主流财经/行业媒体;C 级:厂商自述,本文不采用)。

1. 介绍

1.1 背景

金融是 AI 应用最早、也是受约束最严的行业之一。中国人民银行科技司司长李伟在中国财富管理 50 人论坛 2025 年会(2025-12-28)上的发言给出了监管侧的判断:AI 在"辅助办公、智能客服、信贷风控、合规审查等领域应用场景探索取得积极成效";同时点出四大挑战——科技伦理治理亟待加强、模型同质化风险不容忽视、算法架构内在的缺陷难以根除、数据治理和安全面临挑战(据中证网、国际金融报报道)。

监管的应对思路同样是明确的。金融监管总局办公厅 2025 年 12 月印发的《银行业保险业数字金融高质量发展实施方案》要求:倡导构建人工智能应用分类分级管理框架流程,建立覆盖重要业务流程和关键节点的人工干预机制,推进企业级模型风险管理平台建设,建立算法模型全生命周期管理体系,持续提升算法透明度和可解释性

这意味着金融机构引入 AI Harness 时,第一个要回答的问题不是"能提多少效",而是"这个 AI 应用属于哪个风险等级、模型风险如何验证、人在哪个节点介入"。

1.2 定义与边界

Finance 方向的 AI Harness,指承载投研、风控、量化、财报分析与反洗钱等任务的智能体运行时框架。它的核心职责是:把多源异构、时效敏感的金融数据组织成模型可判定的上下文,把模型的语义能力编排到确定性的风控工具链上,并把每一次判断固化为可举证的证据链。

与其他方向的边界:

相邻方向分工
ComplianceAML 的检测技术在本方向;AML 的报送义务与治理架构在 Compliance
Audit本方向的模型是被审计对象;审计本方向 AI 系统的方法在 Audit
Security本方向的数据不出域、沙箱与日志基座由 Security 提供
数据科学组数据科学组关注分析方法本身;本方向的分析结论要承担模型风险管理义务

1.3 在 AI Harness 体系中的定位

图 1-1|Finance 方向 AI Harness 六层定位(L1–L6)

Finance 方向 AI Harness 六层定位(L1–L6) 信息截止 2026-06-22 · 示意:基于本文分析绘制 L1 上下文工程 多源金融数据,强制携带时间戳与来源权重 瓶颈层 · 权重最重 L2 工具与执行 只读为主;写操作(下单/放款)需审批 权重 · 中 L3 编排与控制 信号与决策分离;授信/放款强制人在回路 权重 · 高 L4 记忆与状态 会话态 + 模型版本登记 + 客户历史结论 权重 · 中 L5 评估与观测 指标最成熟:误报率/召回率/AUC/回测 权重 · 高 L6 治理与安全 受 SR 11-7 三支柱与人工干预机制约束 权重 · 高 结构解读:权重最高处即瓶颈——L1 上下文工程(时间戳 + 来源权重)决定结论可信度。 L3 信号与决策分离保住判定权;L5 天然 Golden Dataset 使回归验证成本五方向最低。

数据来源:基于本文分析绘制的示意图。

Finance 的具体要求权重
L1 上下文工程行情、公告、研报、财报、舆情、产业链数据、监管文件;必须带时间戳与来源权重;过期财报与过期法规进入上下文是本方向最危险的失效模式最重
L2 工具与执行只读为主:行情接口、XBRL 财报解析、规则引擎、图计算引擎;写操作(下单、放款、额度调整)需审批
L3 编排与控制信号生成与决策分离;授信/放款节点强制人在回路
L4 记忆与状态会话态 + 模型版本与参数登记 + 客户历史结论(供复核)
L5 评估与观测本方向是五个方向中评估指标最成熟的:误报率、召回率、AUC/KS、返回测试、不良生成率
L6 治理与安全模型即"模型",受 SR 11-7 三支柱与分类分级 + 人工干预机制约束

1.4 瓶颈层与价值

瓶颈层:L1 上下文工程层。 金融数据的特征是"多源异构 + 时效敏感 + 口径不统一":同一家公司的财务数据在公告、数据商、内部系统里可能有三套口径;同一条产业链事件在研报、新闻、公告里的时点不同。Context 组装若不做时间戳与来源权重管理,模型会产出"看似合理但基于过期信息"的结论——这在投研与风控里是不可接受的失效。

价值落点

  • 提效但不越界:AI 给出候选与线索(signal),人做决策(decision)。这一"signal vs. decision separation"源自 OCC 对"模型输出不得替代重大决策中的人类判断"的要求,也直接对应《银行业保险业数字金融高质量发展实施方案》的人工干预机制。
  • 验证可负担:本方向的 Golden Dataset 天然存在——历史 SAR 结论、历史不良账户、历史授信审批结果。这是五个方向中回归验证成本最低的一个。
  • 风险相称:OCC《Comptroller's Handbook on Model Risk Management》给出的锚句是:Regardless of how AI is classified (i.e., as a model or not a model), the associated risk management should be commensurate with the level of risk of the function that the AI supports. 换言之,治理强度取决于所支撑职能的风险水平,而不取决于"它是不是模型"这个分类问题。

2. 名词解释

术语英文/缩写释义
模型风险Model Risk因模型本身的错误或模型输出的误用而导致决策失误的可能性。SR 11-7 将其归为两大来源:模型本身的错误(假设缺陷、编码错误、数据不当)与模型输出的误用(超范围使用、无文档化地推翻模型结果、不了解模型局限)
SR 11-7Supervisory Guidance on Model Risk Management美联储与 OCC 于 2011-04-04 发布的模型风险管理监管指引(OCC 配套 Bulletin 2011-12,FDIC 于 2017 采用),确立三支柱框架
三支柱Three PillarsSR 11-7 的模型风险管理框架:模型开发与实施 / 模型验证 / 模型治理
概念合理性评价Conceptual Soundness模型验证三要素之一,评价模型设计假设、理论依据与实现是否合理
持续监测Ongoing Monitoring模型验证三要素之一,对模型上线后的表现做持续跟踪
结果分析Outcomes Analysis模型验证三要素之一,含返回测试(Backtesting),将模型输出与实际结果比对
补偿性控制Compensating Controls当无法在用前完成必要验证活动时,用于缓解结果不确定性的替代措施,包括测试、性能监测、红队、结果分析、基准比对以及对验证局限性的适当文档化
反洗钱Anti-Money Laundering, AML为预防洗钱与恐怖融资而建立的客户识别、交易监测、可疑报告与记录保存制度
可疑活动报告Suspicious Activity Report, SAR金融机构向监管/金融情报机构提交的可疑交易报告;中国语境下常称可疑交易报告(STR)
误报False Positive监测系统将正常交易标记为可疑。规则型 AML 系统的误报率可高至 95%~99.5%,是本方向最主要的成本项
动态风险监测Dynamic Risk Monitoring, DRM持续、实时调整客户风险评级(常见为 0~1 的单一分值)的监测方式,替代静态规则匹配
图智能Graph Intelligence用图计算与图神经网络揭示客户、交易对手、实体之间隐藏关联的技术,用于识别分层洗钱与团伙欺诈
客户尽职调查Know Your Customer / Customer Due Diligence, KYC/CDD识别并核实客户身份、受益所有人、业务关系目的与风险等级的过程
KS / AUCKolmogorov-Smirnov / Area Under Curve二分类风控模型常用区分度指标;KS 衡量正负样本分布的最大分离度,AUC 衡量整体排序能力
预期信用损失Expected Credit Loss, ECLIFRS 9 / CECL 下的前瞻性减值计量方法,其模型需经验证
信号与决策分离Signal vs. Decision SeparationAI 产出风险信号与候选清单,授信、放款、交易等决策由人做出的分工原则
模型风险管理平台Model Risk Management Platform, MRM集中登记模型清单、版本、验证状态、监测指标与审批记录的企业级平台

3. 案例

3.1 案例一:某全球性银行 AI 动态风险评估平台(HKMA 2026 报告)

来源:香港金融管理局《Supporting Adoption of Artificial Intelligence in Fighting Financial Crime》,2026-06-22 致所有认可机构发布。取证等级:A 级

3.1.1 背景

2020 年之前,该行依赖大规模规则型 AML 系统。规则型系统的固有问题是"宁可错杀":为满足监管要求而放宽触发条件,导致误报量巨大,需人工逐条调查。分析师的大量时间被消耗在无效告警上,而真正复杂的分层洗钱反而因为跨步骤、跨产品、跨辖区而难以被单笔规则捕获。

3.1.2 方案
  • 与云服务商共同开发动态 AI 平台,基于交易数据、客户数据与 KYC 数据生成风险评分,并送入案件管理系统。
  • 设置领域训练层:用该行自有的犯罪数据与类型学训练模型,识别"资金快速转移"与异常行为模式。
  • 部署在该行自有云租户内,以满足治理与监管要求(数据不出域)。
  • 每月监控超过 10 亿笔交易

Harness 视角:该方案的关键不在模型选型,而在部署形态(自有租户)+ 领域数据层 + 与既有案件管理系统的编排对接。风险评分是"信号",是否立案仍由案件管理系统中人工作出。

3.1.3 效果
指标结果
误报量下降 60%
可疑活动检出提升 2~4 倍
案件调查速度提速约 50%

3.2 案例二:HSBC × Google Cloud AML AI

来源:Google Cloud 与 HSBC 公开披露,经行业媒体转述。取证等级:B 级

3.2.1 背景

HSBC 零售与商业银行的规则型 AML 交易监测存在严重的告警质量问题:超过 95% 的告警在初审阶段被确认为误报,约 98% 最终未形成 SAR 提交。规则只能逐笔匹配静态模式,无法对客户行为全貌做判断。

3.2.2 方案
  • 采用 Google Cloud 的 AML AI 作为主要 AI 交易监测系统,以机器学习模型替代规则。
  • 基于交易数据、KYC 记录、账户行为历史与既往可疑活动模式,生成整合的客户风险评分
  • 对客户行为全貌做实时风险评估,而非逐笔匹配静态规则。
  • 处理规模:每月约检查 9.8 亿笔交易
3.2.3 效果
指标结果
告警量下降 60%
真阳性检出提升 2~4 倍
调查周期从数周压缩至约 8 天

HSBC 集团金融犯罪风险与合规主管将其称为检测客户及账户异常活动的"根本范式转变"。

Harness 解读:两个法域、两家机构、两个技术栈,得到的数字高度一致(误报 -60%、检出 2~4 倍)。这说明收益主要来自范式替换(规则匹配 → 客户级动态风险评分),而非某个模型的性能优势。对 Harness 建设的启示是:投资应放在"客户全貌数据的组装(L1)"与"评分到案件的编排(L3)",而不是模型参数规模。

3.3 案例三:中国银行业智能风控与反洗钱规模化落地

来源:上市银行 2025 年半年报与业绩说明会(A/B 级)、新华网 2025-09-10 综合报道(B 级)、券商行业研究(B 级)。取证等级:A/B 级

3.3.1 背景

中国大型银行的 AI 应用已从"单点试点"进入"全行规模化"阶段。与境外案例不同,中国银业的路径特征是把 AI 嵌入既有的企业级风控平台与信贷流程,而不是另建一套监测系统。

3.3.2 方案
机构建设内容
工商银行建成企业级千亿金融大模型技术体系"工银智涌",赋能 20 余个主要业务领域、200 余个场景;企业级智能风控平台应用于全部境内分行、130 多个风控决策场景,实现商品、外汇、债券、货币、股票五大市场风险智能化排查预警;推出业内首个信贷 AI 智能体矩阵"智贷通"与信贷评审 AI 数字助手"工小审"
建设银行企业级混合大模型体系,接入 DeepSeek 等主流大模型,整合超 60 类数据源、5 亿条向量数据;"授信审批全流程 AI 应用"1 分钟生成 5 大模块、10 余页审查意见初稿
邮储银行"邮智"大模型开展 230 余项场景建设;每日加工约 1.27 亿笔交易流水,构建近百个可疑预警模型,建立以图智能为核心的新型反洗钱可疑交易监测报告体系,实现可疑分析报告自动化生成、有效呈现团伙洗钱证据链
招商银行"AI First"战略,2025 上半年末落地 184 个场景应用;大模型日均 token 量较 2024 年增长 10.1 倍,应用开发者超 1 万人,领域专精模型 183 个
浦发银行大模型日均调用超 400 万次、日均 token 约 60 亿;自主创设智能体超 2000 个,其中 143 个经严格认证后工程化嵌入业务关键流程;建成 10 亿级规模企业级知识库
中信银行构建智能服务场景超 1600 个,2025 上半年依托智能模型增效超 8600 人年
3.3.3 效果
机构量化效果取证等级
邮储银行人工甄别效率提升 30%;贷后风险监测大模型新增识别约 19% 的原预警盲区风险客户;不动产相关权证 AI 识别准确率超 97%;2025 上半年累计保护潜在受害账户逾 10 万户、防止客户资金损失超 ¥8 亿元A/B
建设银行授信审批效率提升 30% 以上(券商研究口径)B
招商银行AI 场景数存在两个口径(856 个184 个);全年通过 AI 替代 1556 万人工小时;共债识别准确率 95% 以上B
浦发银行"浦慧风控"大模型风险预警准确率 92%,不良生成率同比下降 28%;消费贷不良率同比下降 0.48 个百分点至 1.80%B
兴业银行2025 上半年拦截涉诈资金 ¥5.04 亿元、保护潜在受害者资金 ¥8.03 亿元A/B
浙商银行建成"大模型 + 小模型"双引擎驱动的数智化大监督体系,报告期内新增 120 余个业务风险模型A/B

说明:上述多数数字来自券商行业研究与半年报披露,未逐字核对银行年报原文,引用时应标注来源层级。招商银行场景数的两个口径差异本文并列呈现,不择一。


4. 实践标准

4.1 AGENTS.md 规范

4.1.1. AGENTS.md(风险合规组 · Finance 方向)
# AGENTS.md —— Finance(投研 / 风控 / 量化 / 财报分析 / 反洗钱)

> 继承自:风险合规组组级 AGENTS.md。本文件只可加严,不可放松。

## 角色与边界
- 本文件约束在金融场景工作的智能体:投研助手、风控审查助手、量化研究助手、财报分析助手、反洗钱甄别助手。
- 智能体产出的是**风险信号与候选清单**,不是授信、放款、交易或报送决策。
- 生成者(信号生成)、复核者(独立复核)、模型责任人(模型风险登记)三者身份分离;复核者不得为生成者。
- 可自行完成:行情与公告检索、XBRL 财报解析、规则引擎调用、图计算查询、指标计算、回测脚本运行(只读)、草稿生成。
- 不得自行完成:下单、放款、额度调整、报送提交、对客发送。

## 环境假设
- 已接入:行情与资讯终端、XBRL/公告解析器、风控规则引擎、图计算引擎、模型风险管理平台(MRM)、案件管理系统、数据脱敏网关。
- 客户数据与客户交易明细处理在受控环境内完成,禁止出域。
- 所有外部模型调用须经脱敏网关并记录出域字段清单。

## 上下文加载顺序(Context Budget)
1. 现行有效的监管文件与内部风险政策(含版本与生效日期)
2. 客户/标的的结构化档案(KYC、评级、历史授信、历史告警)
3. 本次任务的时效数据(行情快照、财报期间、流水区间)——**强制携带数据截止时点**
4. 检索命中的外部佐证(研报、新闻、公告、产业链数据)——须标注来源与时间戳
5. 模型通用知识——仅用于语言组织,不得作为事实依据
- 强制校验:数据截止时点晚于报告期结束日的方可用于该期结论;否则标注"数据可能过期"。

## 工具契约
| 工具 | 类型 | 授权 | 约束 |
|---|---|---|---|
| 行情/资讯检索 | 只读 | 开 | 必须返回时间戳与数据源 |
| XBRL 财报解析 | 只读 | 开 | 必须返回报表期间与币种(¥ 计价优先) |
| 风控规则引擎 | 计算 | 开 | 必须返回命中的规则编号与规则版本 |
| 图计算引擎 | 计算 | 开 | 必须返回路径与关联强度,供人工复核 |
| 回测框架 | 计算 | 开 | 仅只读运行,不得写入生产策略库 |
| 案件管理系统写入 | 写(内部) | 按次授权 | 仅写草稿与标签,不写结论 |
| 核心系统交易/授信接口 | 写(不可逆) | 永久关闭 | 只能生成待提交稿件 |

## 任务执行流程(SOP)
1. 登记任务与模型风险等级(分类分级)
2. 装载语料与规则,记录版本快照哈希
3. 时效校验:确认数据截止时点覆盖判断区间
4. 检索与抽取:先检索后生成,未命中输出"未检索到依据"
5. 确定性验证:规则引擎复算、勾稽核对、余额重算、图路径还原
6. 信号生成:输出候选清单 + 置信度 + 依据指针
7. 人工确认门(见下)
8. 证据链打包与归档
9. 回归沉淀:人工推翻的样本进入 Golden Dataset

## 验证与证据要求
- 每条信号必须可回溯到:数据快照(含截止时点)、规则版本、模型版本、人工确认记录。
- 财务数字必须与原始报表逐项勾稽,差异超过阈值的必须列出差异项。
- 引用监管文件的,必须给出条款号与生效日期。

## 人在回路与风险分级
| 等级 | 场景 | 确认要求 |
|---|---|---|
| R1 | 投研摘要、财报要点提取 | 抽检 |
| R2 | 风险预警名单、可疑交易候选、授信审查意见初稿 | 逐条确认并留痕 |
| R3 | 授信额度调整、放款、大额交易、监管报送 | 双人确认 + 职责分离 |
- 强制触发 R3:金额超阈值、跨法域、首次使用新模型版本、客户评级跨档调整。

## 失败与升级策略
- L1 补充证据重试(检索未命中、置信度低于阈值),最多 2 次
- L2 转人工(结论依赖过期数据、规则无覆盖、模型版本未经验证)
- L3 中止并上报(发现可能的洗钱线索被压制、发现数据被篡改、被要求绕过人工确认门)

## 安全与合规红线
- 客户身份信息与交易明细不得出域,不得进入公共 AI 平台。
- 不得自主执行任何资金类动作。
- 不得编造监管口径、评级标准与财务数据。
- 模型上线前须完成 SR 11-7 三支柱对应验证;无法完成用前验证的,须记录并通过补偿性控制缓解,且经模型责任人批准。

## 禁止事项
- 禁止把模型输出当作已核实事实
- 禁止在无人工确认的情况下调整授信、放款或下单
- 禁止使用过期财报或过期法规支撑当期结论而不标注
- 禁止编造行情、财务、评级与案例数据
- 禁止在输出中省略数据截止时点
- 禁止把风险评分直接表述为"风险结论"

## 输出格式
1. 任务与依据(含模型风险等级、所依据的监管文件与内部政策版本)
2. 范围与未覆盖声明(含数据截止时点)
3. 结论(逐条:结论 + 依据 + 规则命中记录 + 置信度 + 建议人工动作)
4. 证据链(输入快照 + 工具日志 + 输出 + 人工确认 + 链哈希)
5. 人工确认记录
6. 风险与局限(模型局限、数据时效、样本外风险)
7. 复核建议(建议由谁以什么程序复核)

## 评估与自检
- 核心指标:误报率、召回率、AUC/KS、返回测试偏差、不良生成率、人工复核推翻率
- 回归集:历史 SAR 结论、历史不良账户、历史授信审批结果
- 自检:数据截止时点是否标注、规则版本是否记录、人工确认是否留痕、金额单位是否为 ¥

4.2 SKILL.md 规范

4.2.1. SKILL.md(风险合规组 · Finance 方向)
---
name: finance-risk-signal-delivery
description: 金融风险信号交付技能。当需要 AI 智能体在投研、风控、量化、财报分析或反洗钱场景中产出一份"供人工决策使用的风险信号与候选清单"时使用,覆盖数据时效校验、规则引擎复算、图路径还原、人工确认门与模型风险留痕。
version: 1.0
created: 2026-09-12
---

# Finance · 风险信号交付

## 适用场景
- 投研:产业链事件归因、公告与研报要点抽取、可比公司筛选
- 风控:授信审查意见初稿、贷后风险预警、共债与欺诈识别
- 量化:因子假设生成、回测脚本辅助编写、异常样本归因
- 财报分析:报表勾稽核对、科目异常波动识别、附注要点抽取
- 反洗钱:可疑交易告警分诊、资金路径还原、可疑报告初稿

## 前置条件
- 已加载组级与本方向 AGENTS.md
- 已接入行情/资讯终端、XBRL 解析器、风控规则引擎、图计算引擎、MRM 平台
- 已确认本次任务所涉模型已完成对应风险等级的验证,或已登记补偿性控制
- 数据环境受控,脱敏网关可用

## 输入
- 任务编号、发起人、模型风险等级、输出用途(内部参考 / 监管 / 对客)
- 标的信息(客户号、证券代码、交易批次)
- 判断区间与数据截止时点
- 判定阈值(金额阈值、置信度阈值、误报容忍度)

## 输出
- 七段式输出(任务与依据 / 范围与未覆盖声明 / 结论 / 证据链 / 人工确认记录 / 风险与局限 / 复核建议)
- `signal-pack.json`:候选清单 + 规则命中记录 + 置信度
- `evidence-pack.json`:证据链机器可读副本

## 执行步骤
1. 登记与定级:写入任务编号、模型风险等级、输出用途
2. 时效校验:确认数据截止时点覆盖判断区间;不满足则标注"数据可能过期"并降级处理
3. 数据装载:结构化档案 → 行情/流水快照 → 外部佐证,全部记录版本与时间戳
4. 检索与抽取:先检索后生成;未命中输出"未检索到依据"
5. 确定性验证:
   - 财报:与原始报表逐项勾稽,列差异项
   - 风控:规则引擎复算,返回规则编号与版本
   - 反洗钱:图计算还原资金路径,返回路径与关联强度
   - 量化:回测框架只读运行,输出绩效与回撤(数值必须由框架计算,不由模型估算)
6. 信号生成:候选清单 + 置信度 + 依据指针;置信度低于阈值进入待确认队列
7. 脱敏:出域前经脱敏网关,记录字段清单与授权
8. 人工确认门:按风险等级触发,确认人不得为生成者
9. 证据链打包:输入快照 + 工具日志 + 输出 + 确认记录 + 链哈希
10. 回归沉淀:人工推翻样本进入 Golden Dataset

## 质量标准(DoD)
- 每条信号可回溯到数据快照、规则版本、模型版本
- 财务数字与原始报表逐项勾稽一致
- 金额单位为 ¥,涨跌幅表述遵循中国习惯(涨红跌绿)
- 数据截止时点已标注,过期数据已标注
- 人工确认留痕完整,确认人与生成人不同
- 无编造的行情、财务、评级与案例数据
- 模型风险等级与验证状态已在输出中声明

## 常见失败与处理
| 失败 | 处置 |
|---|---|
| 数据过期 | 标注并降级;关键结论要求补充当期数据 |
| 规则无覆盖 | 输出"规则未覆盖",转人工判断,不得用模型推测替代 |
| 图路径噪声过大 | 提高关联强度阈值并叠加第二类工具交叉验证 |
| 回测过拟合 | 标注样本内/样本外,样本外绩效单独列出 |
| 人工确认门被绕过 | 中止任务,进入 L3 升级 |

## 示例
任务:对某对公客户生成授信审查意见初稿。
执行:装载客户档案与三期财报(截止时点已校验)→ XBRL 解析并勾稽 → 规则引擎复算命中规则 → 生成 5 大模块意见初稿与风险点清单 → 人工确认门(R2,业务主管逐条确认)→ 证据链打包 → 输出标注"AI 生成初稿,不构成授信结论"。

4.3 落地检查清单

序号检查项判定标准对应层级
1数据时效校验每条结论标注数据截止时点,过期数据显式标注L1
2来源权重管理行情/公告/研报/舆情按来源权重排序,模型通用知识不参与事实判定L1
3只读优先核心系统写接口永久关闭,仅生成待提交稿件L2
4规则版本登记规则引擎返回编号与版本,纳入证据链L2
5信号与决策分离输出统一标注"供人工决策使用",不含决策性表述L3
6人在回路节点授信/放款/大额交易为强制节点,留痕含确认人身份与时间戳L3
7模型版本登记模型名称、版本、参数、生效日期登记入 MRM 平台L4
8证据链留存输入快照 + 工具日志 + 输出 + 确认记录 + 链哈希完整L4/L6
9回归集建设Golden Dataset 至少覆盖历史 SAR、历史不良、历史授信三类L5
10指标监控误报率、召回率、AUC/KS、返回测试偏差持续监测L5
11模型风险分级按《银行业保险业数字金融高质量发展实施方案》完成分类分级,并登记人工干预机制L6
12SR 11-7 三支柱开发实施 / 验证 / 治理三支柱均有责任人与文档L6
13补偿性控制无法完成用前验证的模型已登记补偿性控制并经模型责任人批准L6
14数据出域管控客户身份与交易明细未出域;出域调用经脱敏网关并留痕L6
15可解释性风险评分输出附带主要贡献因子与可读解释L6

5. 总结

Finance 方向给 AI Harness 提供了三个可迁移的结论。

第一,瓶颈在 L1 而不是模型。 三个案例的收益都来自"把客户全貌数据组装起来并做动态评分"这一范式替换,而不是模型能力的跃升。误报下降 60%、检出提升 2~4 倍的数字在两个法域、两套技术栈上重复出现,说明这是工程架构的收益

第二,判定权必须留在确定性工具与人手里。 风险评分是信号,是否立案、是否授信、是否放款由人决定。OCC 的锚句值得反复引用:the associated risk management should be commensurate with the level of risk of the function that the AI supports——治理强度取决于风险,而非取决于"是否算模型"。

第三,本方向的回归验证成本最低,因此最应该先做。 历史 SAR 结论、历史不良账户、历史授信审批结果天然构成 Golden Dataset。在五个方向中,Finance 最适合作为风险合规组"评估与观测层(L5)"的示范场景。

需要警惕的是:模型同质化风险(央行科技司 2025-12 指出)与算法架构内在缺陷难以根除这两条挑战不会因为 Harness 建设而消失。Harness 能做的是让缺陷可被发现、可被定位、可被补偿,而不是消除它。


信息缺口声明

  1. 《金融领域科技伦理指引》的原文文号、发布日期与条款号未取得(仅取得人民银行官员讲话引述与媒体转载),本文仅引用其"最小必要采集""专事专用"等原则性表述。
  2. 中国人民银行《人工智能算法金融应用评价规范》(2021)、《人工智能算法金融应用信息披露指南》(2023)的行业标准编号(JR/T 编号)与条文未取得,本文未引用其条款。
  3. SR 11-7 与 OCC Bulletin 2011-12 的官方 PDF 与逐条条款号未取得,本文引用的三支柱、验证三要素与补偿性控制内容来自 BPI、KPMG 汇编材料(B/C 级,多源一致)。
  4. 中国银行业的多数量化效果指标(除邮储"人工甄别效率提升 30%"等少数披露外)来自券商行业研究,未逐字核对银行年报原文。
  5. 券商 IT 投入存在两套统计口径:证券时报口径为34 家上市券商 2025 年度信息技术支出合计 275.9 亿元、占总营收 6.2%;中国基金报口径为28 家合计约 250 亿元、占比 6%。本文采用前者,两者不得混用。
  6. 招商银行 AI 应用场景数存在 856 个184 个两个口径,本文并列呈现,未择一。
  7. 本方向未取得中国侧"AI 量化投研实盘绩效"的可核实 A 级数据,相关指标均为 [待填写]
  8. 本方向未取得中国侧"AI 财报分析准确率"的 A 级公开数据。

6. 参考资料

  1. 《银行业保险业数字金融高质量发展实施方案》,国家金融监督管理总局办公厅,2025-12。https://www.nfra.gov.cn/cn/view/pages/ItemDetail.html?docId=1239741
  2. 中国人民银行科技司司长李伟在中国财富管理 50 人论坛 2025 年会发言报道,中证网,2025-12。https://www.cs.com.cn/xwzx/hg/202512/t20251229_6530683.html
  3. HKMA《Supporting Adoption of Artificial Intelligence in Fighting Financial Crime》,香港金融管理局,2026-06-22。https://brdr.hkma.gov.hk/eng/doc-ldg/current/20260622-1-EN
  4. 《Navigating Artificial Intelligence in Banking》,Bank Policy Institute(含 SR 11-7 与 OCC Comptroller's Handbook 引用),2024。https://bpi.com/wp-content/uploads/2024/04/Navigating-Artificial-Intelligence-in-Banking.pdf
  5. 《AI and model risk》slipsheet,KPMG US,2024。https://kpmg.com/kpmg-us/content/dam/kpmg/pdf/2024/1a-ai-and-model-risk-slipsheet.pdf
  6. HSBC × Google Cloud AML AI 案例报道,Process Excellence Network。https://www.processexcellencenetwork.com/ai/articles/how-hsbc-turned-its-biggest-compliance-headache-into-an-ai-success-story
  7. 上市银行 AI 应用规模化落地综合报道,新华网,2025-09-10。https://www.xinhua.org/20250910/4e421fdeeb2242fc97299d73b2d62f09/c.html
  8. 银行年报 AI 与风控场景梳理,未央网,2025。https://www.weiyangx.com/462780.html
  9. 券商信息技术投入统计,证券时报,2026-04。https://www.stcn.com/article/detail/3731822.html
  10. 券商 AI 应用与 IT 投入报道,21 世纪经济报道,2026-04-27。https://www.21jingji.com/article/20260427/herald/b0e4a124fc6808e55600d6fcd453da08.html
  11. 《金融领域科技伦理指引》相关解读,中国金融信息网,2025-11-17。https://m.cnfin.com/wx/share?url=//m.cnfin.com/hb-lb//zixun/20251117/4336062_1.html
  12. 《生成式人工智能服务管理暂行办法》,国家互联网信息办公室等七部门,2023。https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm

Finance — Financial Risk & Decision Support

Covered scenarios: intelligent investment research, credit & risk control, quantitative research, financial statement analysis, anti-money laundering (AML) and financial crime compliance.

Evidence note: all cases and data cited in this document carry a source tier (Tier A: original regulatory or corporate official text; Tier B: mainstream financial/industry media; Tier C: vendor self-reports, not used in this document).

1. Introduction

1.1 Background

Finance is one of the earliest industries to adopt AI and one of the most tightly constrained. At the 2025 annual meeting of the China Wealth Management 50 Forum (2025-12-28), Li Wei, Director of the Technology Department of the People's Bank of China, offered the regulator's assessment: AI has achieved positive results in exploring application scenarios in areas such as "administrative assistance, intelligent customer service, credit risk control and compliance review"; he also pointed out four major challenges — science-and-technology ethics governance urgently needs strengthening, model homogeneity risk cannot be ignored, inherent flaws in algorithmic architectures are hard to eliminate, and data governance and security face challenges (as reported by China Securities Journal and International Finance News).

The regulator's response is equally clear. The Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industry, issued by the General Office of the National Financial Regulatory Administration in December 2025, requires: advocating the construction of classified and tiered management framework processes for AI applications, establishing human intervention mechanisms covering important business processes and key nodes, advancing the building of enterprise-grade model risk management platforms, establishing full-lifecycle management systems for algorithmic models, and continuously improving algorithmic transparency and explainability.

This means that when a financial institution introduces an AI Harness, the first question to answer is not "how much efficiency can it gain", but "which risk tier does this AI application belong to, how is model risk validated, and at which node does a human intervene".

1.2 Definition and Boundaries

The AI Harness for the Finance direction refers to the agent runtime framework that carries tasks such as investment research, risk control, quantitative analysis, financial statement analysis and anti-money laundering. Its core responsibilities are: to organize multi-source, heterogeneous, time-sensitive financial data into model-judgable context, to orchestrate the model's semantic capabilities onto deterministic risk-control toolchains, and to solidify every judgment into an auditable chain of evidence.

Boundaries with other directions:

Adjacent directionDivision of labor
ComplianceAML detection technologies belong to this direction; AML reporting obligations and governance architecture belong to Compliance
AuditThis direction's models are the objects being audited; the methods for auditing this direction's AI systems belong to Audit
SecurityThis direction's data never leaves its domain; the sandbox and logging foundation are provided by Security
Data science groupThe data science group focuses on the analytical methods themselves; this direction's analytical conclusions must bear model risk management obligations

1.3 Position within the AI Harness System

图 1-1|Finance 方向 AI Harness 六层定位(L1–L6)

Finance 方向 AI Harness 六层定位(L1–L6) 信息截止 2026-06-22 · 示意:基于本文分析绘制 L1 上下文工程 多源金融数据,强制携带时间戳与来源权重 瓶颈层 · 权重最重 L2 工具与执行 只读为主;写操作(下单/放款)需审批 权重 · 中 L3 编排与控制 信号与决策分离;授信/放款强制人在回路 权重 · 高 L4 记忆与状态 会话态 + 模型版本登记 + 客户历史结论 权重 · 中 L5 评估与观测 指标最成熟:误报率/召回率/AUC/回测 权重 · 高 L6 治理与安全 受 SR 11-7 三支柱与人工干预机制约束 权重 · 高 结构解读:权重最高处即瓶颈——L1 上下文工程(时间戳 + 来源权重)决定结论可信度。 L3 信号与决策分离保住判定权;L5 天然 Golden Dataset 使回归验证成本五方向最低。

数据来源:基于本文分析绘制的示意图。

LayerSpecific Finance requirementsWeight
L1 Context EngineeringMarket data, announcements, research reports, financial statements, market sentiment, supply-chain data, regulatory documents; must carry timestamps and source weights; expired financial statements and expired regulations entering the context is the most dangerous failure mode for this directionHeaviest
L2 Tools & ExecutionRead-only dominant: market data interfaces, XBRL financial statement parsing, rules engines, graph computation engines; write operations (order placement, lending, limit adjustments) require approvalMedium
L3 Orchestration & ControlSignal generation and decision-making are separated; credit/lending nodes mandate human-in-the-loopHigh
L4 Memory & StateSession state + model version and parameter registration + historical customer conclusions (for review)Medium
L5 Evaluation & ObservabilityThis direction has the most mature evaluation metrics of the five directions: false positive rate, recall, AUC/KS, backtesting, non-performing generation rateHigh
L6 Governance & SecurityThe model is a "model", subject to the SR 11-7 three pillars and classification/tiering plus human intervention mechanismHigh

1.4 Bottleneck Layer and Value

Bottleneck layer: the L1 Context Engineering layer. Financial data is characterized by "multi-source heterogeneity + time sensitivity + inconsistent accounting calibers": the same company's financial data may exist in three different calibers across announcements, data vendors and internal systems; the same industry-chain event may appear at different time points across research reports, news and announcements. If Context assembly does not manage timestamps and source weights, the model will produce conclusions that "look reasonable but are based on outdated information" — an unacceptable failure in investment research and risk control.

Where the value lands:

  • Efficiency without overstepping: AI provides candidates and leads (signal), humans make the decisions (decision). This "signal vs. decision separation" derives from the OCC requirement that "model output must not replace human judgment in material decisions", and it directly corresponds to the human intervention mechanism in the Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industry.
  • Affordable validation: a Golden Dataset naturally exists for this direction — historical SAR conclusions, historical non-performing accounts, and historical credit approval results. This has the lowest regression validation cost of the five directions.
  • Risk proportionality: the anchor sentence in the OCC Comptroller's Handbook on Model Risk Management is: Regardless of how AI is classified (i.e., as a model or not a model), the associated risk management should be commensurate with the level of risk of the function that the AI supports. In other words, the intensity of governance depends on the risk level of the function being supported, not on the classification question of "whether it is a model".

2. Glossary

TermEnglish / abbreviationDefinition
Model riskModel RiskThe possibility of decision-making errors caused by errors in the model itself or misuse of model output. SR 11-7 classifies it into two major sources: errors in the model itself (flawed assumptions, coding errors, inappropriate data) and misuse of model output (use beyond scope, undocumented overriding of model results, unawareness of model limitations)
SR 11-7Supervisory Guidance on Model Risk ManagementThe supervisory guidance on model risk management issued by the Federal Reserve and the OCC on 2011-04-04 (with OCC Bulletin 2011-12; adopted by the FDIC in 2017), establishing the three-pillar framework
Three pillarsThree PillarsSR 11-7's model risk management framework: model development and implementation / model validation / model governance
Conceptual soundnessConceptual SoundnessOne of the three elements of model validation; evaluates whether a model's design assumptions, theoretical basis and implementation are reasonable
Ongoing monitoringOngoing MonitoringOne of the three elements of model validation; continuously tracks a model's performance after go-live
Outcomes analysisOutcomes AnalysisOne of the three elements of model validation; includes backtesting, comparing model output against actual results
Compensating controlsCompensating ControlsAlternative measures used to mitigate the uncertainty of results when the necessary validation activities cannot be completed before use, including testing, performance monitoring, red-teaming, outcomes analysis, benchmarking, and appropriate documentation of validation limitations
Anti-money launderingAnti-Money Laundering, AMLThe customer identification, transaction monitoring, suspicious reporting and record-keeping regime established to prevent money laundering and terrorist financing
Suspicious activity reportSuspicious Activity Report, SARA suspicious transaction report submitted by financial institutions to regulators/financial intelligence units; in the Chinese context it is commonly called a Suspicious Transaction Report (STR)
False positiveFalse PositiveA monitoring system flagging a normal transaction as suspicious. Rule-based AML systems can have false positive rates as high as 95%~99.5%, the largest cost item for this direction
Dynamic risk monitoringDynamic Risk Monitoring, DRMA monitoring approach that continuously adjusts customer risk ratings in real time (commonly a single score from 0 to 1), replacing static rule matching
Graph intelligenceGraph IntelligenceTechnology that uses graph computation and graph neural networks to reveal hidden relationships between customers, counterparties and entities, used to identify layering money laundering and organized fraud
Customer due diligenceKnow Your Customer / Customer Due Diligence, KYC/CDDThe process of identifying and verifying customer identity, beneficial owners, the purpose of the business relationship, and risk level
KS / AUCKolmogorov-Smirnov / Area Under CurveCommonly used discrimination metrics for binary-classification risk-control models; KS measures the maximum separation between positive and negative sample distributions, AUC measures overall ranking ability
Expected credit lossExpected Credit Loss, ECLA forward-looking impairment measurement under IFRS 9 / CECL; its models must be validated
Signal vs. decision separationSignal vs. Decision SeparationThe division-of-labor principle whereby AI produces risk signals and candidate lists while decisions on credit, lending, trading, etc. are made by humans
Model risk management platformModel Risk Management Platform, MRMAn enterprise-grade platform that centrally registers the model inventory, versions, validation status, monitoring metrics and approval records

3. Case Studies

3.1 Case 1: A Global Bank's AI Dynamic Risk Assessment Platform (HKMA 2026 Report)

Source: Hong Kong Monetary Authority, Supporting Adoption of Artificial Intelligence in Fighting Financial Crime, issued to all authorized institutions on 2026-06-22. Evidence tier: A.

3.1.1 Background

Before 2020, the bank relied on a large-scale rule-based AML system. The inherent problem of rule-based systems is "rather kill an innocent than let a guilty go": to satisfy regulatory requirements, trigger conditions were loosened, producing a huge volume of false positives that required manual, case-by-case investigation. Analysts' time was largely consumed by ineffective alerts, while truly complex layering money laundering — because it spans steps, products and jurisdictions — was hard for any single-transaction rule to catch.

3.1.2 Solution
  • Developed a dynamic AI platform jointly with a cloud provider, generating risk scores from transaction data, customer data and KYC data, and feeding them into the case management system.
  • Set up a domain training layer: trained the model on the bank's own crime data and typologies to identify "rapid movement of funds" and anomalous behavior patterns.
  • Deployed within the bank's own cloud tenant to satisfy governance and regulatory requirements (data does not leave the domain).
  • Monitors over 1 billion transactions per month.

Harness perspective: the key to this solution is not model selection but the deployment form (own tenant) + domain data layer + orchestration against the existing case management system. The risk score is a "signal"; whether to open a case is still decided by humans within the case management system.

3.1.3 Results
MetricResult
False positivesdown 60%
Suspicious activity detectionup 2~4x
Case investigation speedabout 50% faster

3.2 Case 2: HSBC × Google Cloud AML AI

Source: public disclosures by Google Cloud and HSBC, relayed through industry media. Evidence tier: B.

3.2.1 Background

HSBC's rule-based AML transaction monitoring for retail and commercial banking suffered severe alert-quality problems: over 95% of alerts were confirmed as false positives at the initial review stage, and about 98% ultimately never led to an SAR filing. Rules can only match static patterns transaction by transaction; they cannot judge the full picture of customer behavior.

3.2.2 Solution
  • Adopted Google Cloud's AML AI as the primary AI transaction monitoring system, replacing rules with machine-learning models.
  • Generated an integrated customer risk score based on transaction data, KYC records, account behavior history and past suspicious activity patterns.
  • Performed real-time risk assessment of the full picture of customer behavior, rather than matching static rules per transaction.
  • Processing scale: about 980 million transactions checked per month.
3.2.3 Results
MetricResult
Alert volumedown 60%
True-positive detectionup 2~4x
Investigation cyclecompressed from weeks to about 8 days

HSBC's Group Head of Financial Crime Risk and Compliance called it a "fundamental paradigm shift" in detecting anomalous customer and account activity.

Harness interpretation: two jurisdictions, two institutions, two tech stacks, yet the numbers are highly consistent (false positives -60%, detection 2~4x). This shows the gains come mainly from a paradigm replacement (rule matching → customer-level dynamic risk scoring), not from the performance advantage of any particular model. For Harness construction, the implication is: investment should go into "assembly of the full picture of customer data (L1)" and "orchestration from score to case (L3)", rather than model parameter scale.

3.3 Case 3: Large-Scale Deployment of Intelligent Risk Control and AML in China's Banking Industry

Sources: listed banks' 2025 semi-annual reports and earnings calls (Tiers A/B), Xinhua Net's 2025-09-10 comprehensive report (Tier B), and brokerage industry research (Tier B). Evidence tier: A/B.

3.3.1 Background

AI adoption among China's large banks has moved from "single-point pilots" to "bank-wide scale". Unlike overseas cases, the distinctive path of China's banking industry is to embed AI into existing enterprise-grade risk control platforms and credit processes rather than building a separate monitoring system.

3.3.2 Solution
InstitutionWhat was built
ICBCBuilt the enterprise-grade "ICBC Zhiyong" hundred-billion-parameter financial large-model technology system, empowering 20+ major business domains and 200+ scenarios; the enterprise-grade intelligent risk control platform is applied across all domestic branches and 130+ risk-control decision scenarios, achieving intelligent screening and early warning for five major market risks — commodities, FX, bonds, money market and equities; launched the industry's first credit AI agent matrix "Zhidaifang" and the credit-review AI digital assistant "Gongxiaoshen"
CCBEnterprise-grade hybrid large-model system, integrating mainstream large models such as DeepSeek and aggregating 60+ data sources and 500 million vector records; the "full-process AI credit approval application" generates a first draft of the review opinion covering 5 major modules and 10+ pages in 1 minute
Postal Savings BankThe "Youzhi" large model is used to build 230+ scenarios; processes about 127 million transaction flows per day, builds nearly a hundred suspicious-alert models, and establishes a new AML suspicious-transaction monitoring and reporting system built around graph intelligence, enabling automated generation of suspicious-analysis reports and effective presentation of organized money-laundering evidence chains
China Merchants Bank"AI First" strategy, with 184 scenario applications in place by end of H1 2025; large-model daily token volume up 10.1x versus 2024, over 10,000 application developers, and 183 domain-specialized models
SPD BankLarge model called over 4 million times per day, about 6 billion daily tokens; independently created over 2,000 agents, of which 143 were rigorously certified and engineered into key business processes; built an enterprise-grade knowledge base at the billion-record scale
China CITIC BankBuilt 1,600+ intelligent service scenarios; in H1 2025 relied on intelligent models to achieve efficiency gains exceeding 8,600 person-years
3.3.3 Results
InstitutionQuantified resultsEvidence tier
Postal Savings BankManual screening efficiency up 30%; the post-lending risk monitoring large model newly identifies about 19% of at-risk customers that were blind spots of the previous early warning; AI recognition accuracy for real-estate-related certificates above 97%; in H1 2025 protected more than 100,000 potential victim accounts cumulatively and prevented customer fund losses exceeding ¥800 millionA/B
CCBCredit approval efficiency up 30%+ (brokerage research basis)B
China Merchants BankTwo bases exist for the number of AI scenarios (856 and 184); throughout the year AI replaced 15.56 million person-hours; co-debt identification accuracy above 95%B
SPD Bank"Puhui Risk Control" large-model risk warning accuracy 92%, non-performing generation rate down 28% year on year; consumer-loan non-performing rate down 0.48 percentage points to 1.80%B
Industrial BankIn H1 2025 intercepted ¥504 million of fraud-related funds and protected ¥803 million of potential victims' fundsA/B
CZ BankBuilt a data-intelligent mega-supervision system driven by a "large model + small model" dual-engine, adding 120+ business risk models during the reporting periodA/B

Note: most of the figures above come from brokerage industry research and semi-annual report disclosures, and have not been cross-checked verbatim against the banks' annual report originals; cite them with the source tier noted. The two different bases for China Merchants Bank's scenario count are presented side by side here, without selecting one.


4. Practical Standards

4.1 AGENTS.md Specification

4.1.1. AGENTS.md (Risk & Compliance Group · Finance Direction)
# AGENTS.md —— Finance(投研 / 风控 / 量化 / 财报分析 / 反洗钱)

> 继承自:风险合规组组级 AGENTS.md。本文件只可加严,不可放松。

## 角色与边界
- 本文件约束在金融场景工作的智能体:投研助手、风控审查助手、量化研究助手、财报分析助手、反洗钱甄别助手。
- 智能体产出的是**风险信号与候选清单**,不是授信、放款、交易或报送决策。
- 生成者(信号生成)、复核者(独立复核)、模型责任人(模型风险登记)三者身份分离;复核者不得为生成者。
- 可自行完成:行情与公告检索、XBRL 财报解析、规则引擎调用、图计算查询、指标计算、回测脚本运行(只读)、草稿生成。
- 不得自行完成:下单、放款、额度调整、报送提交、对客发送。

## 环境假设
- 已接入:行情与资讯终端、XBRL/公告解析器、风控规则引擎、图计算引擎、模型风险管理平台(MRM)、案件管理系统、数据脱敏网关。
- 客户数据与客户交易明细处理在受控环境内完成,禁止出域。
- 所有外部模型调用须经脱敏网关并记录出域字段清单。

## 上下文加载顺序(Context Budget)
1. 现行有效的监管文件与内部风险政策(含版本与生效日期)
2. 客户/标的的结构化档案(KYC、评级、历史授信、历史告警)
3. 本次任务的时效数据(行情快照、财报期间、流水区间)——**强制携带数据截止时点**
4. 检索命中的外部佐证(研报、新闻、公告、产业链数据)——须标注来源与时间戳
5. 模型通用知识——仅用于语言组织,不得作为事实依据
- 强制校验:数据截止时点晚于报告期结束日的方可用于该期结论;否则标注"数据可能过期"。

## 工具契约
| 工具 | 类型 | 授权 | 约束 |
|---|---|---|---|
| 行情/资讯检索 | 只读 | 开 | 必须返回时间戳与数据源 |
| XBRL 财报解析 | 只读 | 开 | 必须返回报表期间与币种(¥ 计价优先) |
| 风控规则引擎 | 计算 | 开 | 必须返回命中的规则编号与规则版本 |
| 图计算引擎 | 计算 | 开 | 必须返回路径与关联强度,供人工复核 |
| 回测框架 | 计算 | 开 | 仅只读运行,不得写入生产策略库 |
| 案件管理系统写入 | 写(内部) | 按次授权 | 仅写草稿与标签,不写结论 |
| 核心系统交易/授信接口 | 写(不可逆) | 永久关闭 | 只能生成待提交稿件 |

## 任务执行流程(SOP)
1. 登记任务与模型风险等级(分类分级)
2. 装载语料与规则,记录版本快照哈希
3. 时效校验:确认数据截止时点覆盖判断区间
4. 检索与抽取:先检索后生成,未命中输出"未检索到依据"
5. 确定性验证:规则引擎复算、勾稽核对、余额重算、图路径还原
6. 信号生成:输出候选清单 + 置信度 + 依据指针
7. 人工确认门(见下)
8. 证据链打包与归档
9. 回归沉淀:人工推翻的样本进入 Golden Dataset

## 验证与证据要求
- 每条信号必须可回溯到:数据快照(含截止时点)、规则版本、模型版本、人工确认记录。
- 财务数字必须与原始报表逐项勾稽,差异超过阈值的必须列出差异项。
- 引用监管文件的,必须给出条款号与生效日期。

## 人在回路与风险分级
| 等级 | 场景 | 确认要求 |
|---|---|---|
| R1 | 投研摘要、财报要点提取 | 抽检 |
| R2 | 风险预警名单、可疑交易候选、授信审查意见初稿 | 逐条确认并留痕 |
| R3 | 授信额度调整、放款、大额交易、监管报送 | 双人确认 + 职责分离 |
- 强制触发 R3:金额超阈值、跨法域、首次使用新模型版本、客户评级跨档调整。

## 失败与升级策略
- L1 补充证据重试(检索未命中、置信度低于阈值),最多 2 次
- L2 转人工(结论依赖过期数据、规则无覆盖、模型版本未经验证)
- L3 中止并上报(发现可能的洗钱线索被压制、发现数据被篡改、被要求绕过人工确认门)

## 安全与合规红线
- 客户身份信息与交易明细不得出域,不得进入公共 AI 平台。
- 不得自主执行任何资金类动作。
- 不得编造监管口径、评级标准与财务数据。
- 模型上线前须完成 SR 11-7 三支柱对应验证;无法完成用前验证的,须记录并通过补偿性控制缓解,且经模型责任人批准。

## 禁止事项
- 禁止把模型输出当作已核实事实
- 禁止在无人工确认的情况下调整授信、放款或下单
- 禁止使用过期财报或过期法规支撑当期结论而不标注
- 禁止编造行情、财务、评级与案例数据
- 禁止在输出中省略数据截止时点
- 禁止把风险评分直接表述为"风险结论"

## 输出格式
1. 任务与依据(含模型风险等级、所依据的监管文件与内部政策版本)
2. 范围与未覆盖声明(含数据截止时点)
3. 结论(逐条:结论 + 依据 + 规则命中记录 + 置信度 + 建议人工动作)
4. 证据链(输入快照 + 工具日志 + 输出 + 人工确认 + 链哈希)
5. 人工确认记录
6. 风险与局限(模型局限、数据时效、样本外风险)
7. 复核建议(建议由谁以什么程序复核)

## 评估与自检
- 核心指标:误报率、召回率、AUC/KS、返回测试偏差、不良生成率、人工复核推翻率
- 回归集:历史 SAR 结论、历史不良账户、历史授信审批结果
- 自检:数据截止时点是否标注、规则版本是否记录、人工确认是否留痕、金额单位是否为 ¥

4.2 SKILL.md Specification

4.2.1. SKILL.md (Risk & Compliance Group · Finance Direction)
---
name: finance-risk-signal-delivery
description: 金融风险信号交付技能。当需要 AI 智能体在投研、风控、量化、财报分析或反洗钱场景中产出一份"供人工决策使用的风险信号与候选清单"时使用,覆盖数据时效校验、规则引擎复算、图路径还原、人工确认门与模型风险留痕。
version: 1.0
created: 2026-09-12
---

# Finance · 风险信号交付

## 适用场景
- 投研:产业链事件归因、公告与研报要点抽取、可比公司筛选
- 风控:授信审查意见初稿、贷后风险预警、共债与欺诈识别
- 量化:因子假设生成、回测脚本辅助编写、异常样本归因
- 财报分析:报表勾稽核对、科目异常波动识别、附注要点抽取
- 反洗钱:可疑交易告警分诊、资金路径还原、可疑报告初稿

## 前置条件
- 已加载组级与本方向 AGENTS.md
- 已接入行情/资讯终端、XBRL 解析器、风控规则引擎、图计算引擎、MRM 平台
- 已确认本次任务所涉模型已完成对应风险等级的验证,或已登记补偿性控制
- 数据环境受控,脱敏网关可用

## 输入
- 任务编号、发起人、模型风险等级、输出用途(内部参考 / 监管 / 对客)
- 标的信息(客户号、证券代码、交易批次)
- 判断区间与数据截止时点
- 判定阈值(金额阈值、置信度阈值、误报容忍度)

## 输出
- 七段式输出(任务与依据 / 范围与未覆盖声明 / 结论 / 证据链 / 人工确认记录 / 风险与局限 / 复核建议)
- `signal-pack.json`:候选清单 + 规则命中记录 + 置信度
- `evidence-pack.json`:证据链机器可读副本

## 执行步骤
1. 登记与定级:写入任务编号、模型风险等级、输出用途
2. 时效校验:确认数据截止时点覆盖判断区间;不满足则标注"数据可能过期"并降级处理
3. 数据装载:结构化档案 → 行情/流水快照 → 外部佐证,全部记录版本与时间戳
4. 检索与抽取:先检索后生成;未命中输出"未检索到依据"
5. 确定性验证:
   - 财报:与原始报表逐项勾稽,列差异项
   - 风控:规则引擎复算,返回规则编号与版本
   - 反洗钱:图计算还原资金路径,返回路径与关联强度
   - 量化:回测框架只读运行,输出绩效与回撤(数值必须由框架计算,不由模型估算)
6. 信号生成:候选清单 + 置信度 + 依据指针;置信度低于阈值进入待确认队列
7. 脱敏:出域前经脱敏网关,记录字段清单与授权
8. 人工确认门:按风险等级触发,确认人不得为生成者
9. 证据链打包:输入快照 + 工具日志 + 输出 + 确认记录 + 链哈希
10. 回归沉淀:人工推翻样本进入 Golden Dataset

## 质量标准(DoD)
- 每条信号可回溯到数据快照、规则版本、模型版本
- 财务数字与原始报表逐项勾稽一致
- 金额单位为 ¥,涨跌幅表述遵循中国习惯(涨红跌绿)
- 数据截止时点已标注,过期数据已标注
- 人工确认留痕完整,确认人与生成人不同
- 无编造的行情、财务、评级与案例数据
- 模型风险等级与验证状态已在输出中声明

## 常见失败与处理
| 失败 | 处置 |
|---|---|
| 数据过期 | 标注并降级;关键结论要求补充当期数据 |
| 规则无覆盖 | 输出"规则未覆盖",转人工判断,不得用模型推测替代 |
| 图路径噪声过大 | 提高关联强度阈值并叠加第二类工具交叉验证 |
| 回测过拟合 | 标注样本内/样本外,样本外绩效单独列出 |
| 人工确认门被绕过 | 中止任务,进入 L3 升级 |

## 示例
任务:对某对公客户生成授信审查意见初稿。
执行:装载客户档案与三期财报(截止时点已校验)→ XBRL 解析并勾稽 → 规则引擎复算命中规则 → 生成 5 大模块意见初稿与风险点清单 → 人工确认门(R2,业务主管逐条确认)→ 证据链打包 → 输出标注"AI 生成初稿,不构成授信结论"。

4.3 Landing Checklist

#Check itemAcceptance criteriaCorresponding layer
1Data timeliness checkEach conclusion is marked with its data cutoff time; outdated data is explicitly flaggedL1
2Source weight managementMarket data/announcements/research reports/market sentiment sorted by source weight; model general knowledge does not participate in fact determinationL1
3Read-only priorityCore-system write interfaces are permanently disabled; only drafts for submission are generatedL2
4Rule version registrationRules engine returns IDs and versions, included in the evidence chainL2
5Signal/decision separationOutput is uniformly marked "for human decision-making", containing no decision-style statementsL3
6Human-in-the-loop nodesCredit/lending/large transactions are mandatory nodes; the trail includes the confirmer's identity and timestampL3
7Model version registrationModel name, version, parameters and effective date registered in the MRM platformL4
8Evidence chain retentionInput snapshot + tool logs + output + confirmation records + chain hash are completeL4/L6
9Regression set constructionGolden Dataset covers at least historical SAR, historical non-performing, and historical credit approvalL5
10Metric monitoringFalse positive rate, recall, AUC/KS and backtesting deviation continuously monitoredL5
11Model risk classificationClassification/tiering completed per the Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industry, with the human intervention mechanism registeredL6
12SR 11-7 three pillarsAll three pillars — development/implementation, validation, governance — have owners and documentationL6
13Compensating controlsModels that cannot complete pre-use validation have registered compensating controls approved by the model ownerL6
14Data egress controlCustomer identity and transaction details do not leave the domain; external calls go through the de-identification gateway and leave an audit trailL6
15ExplainabilityRisk score output includes the main contributing factors and a readable explanationL6

5. Conclusion

The Finance direction offers three transferable takeaways for AI Harness.

First, the bottleneck is at L1, not the model. The gains in all three cases come from the paradigm replacement of "assembling the full picture of customer data and scoring it dynamically", not from a leap in model capability. The figures of a 60% drop in false positives and a 2~4x increase in detection recur across two jurisdictions and two tech stacks, showing this is a gain from engineering architecture.

Second, the power to decide must stay in deterministic tools and in human hands. The risk score is a signal; whether to open a case, approve credit, or disburse funds is decided by humans. The OCC anchor sentence is worth quoting repeatedly: the associated risk management should be commensurate with the level of risk of the function that the AI supports — the intensity of governance depends on risk, not on "whether it counts as a model".

Third, this direction has the lowest regression-validation cost, so it should be done first. Historical SAR conclusions, historical non-performing accounts and historical credit approval results naturally form a Golden Dataset. Among the five directions, Finance is best suited as the demonstration scenario for the Risk & Compliance group's "evaluation and observability layer (L5)".

What warrants caution: the two challenges of model homogeneity risk (noted by the PBOC's Tech Department in 2025-12) and the inherent flaws of algorithmic architectures being hard to eliminate will not disappear because of Harness construction. What Harness can do is make flaws discoverable, localizable and compensable, not eliminate them.


Information Gap Statement

  1. The original document number, publication date and article numbers of the Guidelines on Science-and-Technology Ethics in the Financial Sector have not been obtained (only a PBOC official's speech quote and media reprints), so this document cites only its principled formulations such as "minimum necessary collection" and "dedicated use for dedicated purposes".
  2. The industry-standard numbers (JR/T numbers) and provisions of the PBOC's Evaluation Specification for AI Algorithm Financial Applications (2021) and Information Disclosure Guidelines for AI Algorithm Financial Applications (2023) have not been obtained, and this document does not cite their articles.
  3. The official PDFs and per-article numbers of SR 11-7 and OCC Bulletin 2011-12 have not been obtained; the three-pillar, three-validation-element and compensating-control content cited here comes from BPI and KPMG compilations (Tiers B/C, consistent across multiple sources).
  4. Most of the quantitative effect indicators for China's banks (apart from a few disclosures such as Postal Savings Bank's "manual screening efficiency up 30%") come from brokerage industry research and have not been cross-checked verbatim against the banks' annual report originals.
  5. Two statistical bases exist for brokerage IT spending: the Securities Times basis is 34 listed brokers' total 2025 information technology expenditure of ¥27.59 billion, 6.2% of total revenue; the China Fund News basis is 28 brokers totaling about ¥25 billion, 6% of revenue. This document adopts the former; the two must not be mixed.
  6. Two bases exist for China Merchants Bank's AI scenario count — 856 and 184 — and this document presents them side by side without selecting one.
  7. No verifiable Tier-A data has been obtained for China-side "live-trading performance of AI quantitative investment research" in this direction; the relevant metrics are all [To be filled].
  8. No Tier-A public data has been obtained for China-side "AI financial statement analysis accuracy" in this direction.

6. References

  1. Implementation Plan for High-Quality Development of Digital Finance in the Banking and Insurance Industry, General Office of the National Financial Regulatory Administration, 2025-12. https://www.nfra.gov.cn/cn/view/pages/ItemDetail.html?docId=1239741
  2. Report on the speech by Li Wei, Director of the PBOC's Technology Department, at the 2025 annual meeting of the China Wealth Management 50 Forum, China Securities Journal, 2025-12. https://www.cs.com.cn/xwzx/hg/202512/t20251229_6530683.html
  3. HKMA, Supporting Adoption of Artificial Intelligence in Fighting Financial Crime, Hong Kong Monetary Authority, 2026-06-22. https://brdr.hkma.gov.hk/eng/doc-ldg/current/20260622-1-EN
  4. Navigating Artificial Intelligence in Banking, Bank Policy Institute (with references to SR 11-7 and the OCC Comptroller's Handbook), 2024. https://bpi.com/wp-content/uploads/2024/04/Navigating-Artificial-Intelligence-in-Banking.pdf
  5. AI and model risk slipsheet, KPMG US, 2024. https://kpmg.com/kpmg-us/content/dam/kpmg/pdf/2024/1a-ai-and-model-risk-slipsheet.pdf
  6. Report on the HSBC × Google Cloud AML AI case, Process Excellence Network. https://www.processexcellencenetwork.com/ai/articles/how-hsbc-turned-its-biggest-compliance-headache-into-an-ai-success-story
  7. Comprehensive report on large-scale AI deployment by listed banks, Xinhua Net, 2025-09-10. https://www.xinhua.org/20250910/4e421fdeeb2242fc97299d73b2d62f09/c.html
  8. Review of AI and risk-control scenarios in banks' annual reports, Weiyang Net, 2025. https://www.weiyangx.com/462780.html
  9. Brokerage information technology investment statistics, Securities Times, 2026-04. https://www.stcn.com/article/detail/3731822.html
  10. Report on broker AI applications and IT investment, 21st Century Business Herald, 2026-04-27. https://www.21jingji.com/article/20260427/herald/b0e4a124fc6808e55600d6fcd453da08.html
  11. Interpretation related to the Guidelines on Science-and-Technology Ethics in the Financial Sector, China Financial Information Network, 2025-11-17. https://m.cnfin.com/wx/share?url=//m.cnfin.com/hb-lb//zixun/20251117/4336062_1.html
  12. Interim Measures for the Administration of Generative AI Services, Cyberspace Administration of China and six other departments, 2023. https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm