智能制造(Manufacture)中的 AI Harness
1. 介绍
1.1 背景
智能制造是 AI Harness 落地门槛最高的行业方向之一。其根本原因不在于模型能力不足,而在于判定标准的物理性:一次排产失误的代价是库存与交期,一次质检漏检的代价是召回,一次工艺参数误调的代价是批量废品。这些代价无法用"重试一次"来消化,因此智能制造对 Harness 的要求天然偏向可观测(L5)与可约束(L6)。
从标准侧看,智能制造是全球范围内标准化程度最高的制造范式之一,已经形成了三层可引用的规范骨架:
- 成熟度层:《智能制造能力成熟度模型》(GB/T 39116-2020) 与《智能制造能力成熟度评估方法》(GB/T 39117-2020),定义了五级成熟度与评估方法。
- 数字孪生层:ISO 23247 系列《Automation systems and integration — Digital twin framework for manufacturing》,Part 1:2021 概述与一般原则、Part 4:2021 信息交换、Part 5:2026 数字主线、Part 6:2026 数字孪生组合。
- 集成与资产层:IEC 62264(等同 ANSI/ISA-95)企业—控制系统集成五层模型;RAMI 4.0(DIN SPEC 91345:2016-04)参考架构模型与资产管理壳 AAS(IEC 63278-1:2023)。
这套骨架的意义在于:它已经把"制造系统应该长什么样"定义清楚了,AI Harness 不需要另起炉灶,只需要把自己的六层能力对齐到这套骨架上。
1.2 定义与范围
本方向的 AI Harness 指:在智能制造场景下,承接工艺、排产、质检、数字孪生四类任务,把模型输出转化为可下发、可追溯、可回归验证的制造指令与判定结果的工程化承载层。
| 任务域 | 典型决策 | 输出工件 | 错误代价 |
|---|---|---|---|
| 工艺 | 工艺参数寻优、配方推荐、工艺规则沉淀 | 工艺卡、参数建议、规则库条目 | 批量废品、质量波动 |
| 排产 | 多约束下的作业排序与资源分配 | 排产方案、工单序列 | 库存积压、交期违约 |
| 质检 | 缺陷检测与判定、漏检与过杀平衡 | 判定结果、缺陷图谱、SPC 数据 | 召回、客户索赔 |
| 数字孪生 | 虚拟验证、实时控制辅助、在制自适应 | 孪生状态、仿真结论、异常检测结果 | 决策误判、停机损失 |
1.3 在 AI Harness 体系中的定位
图 1-1|智能制造 AI Harness 六层定位与瓶颈
数据来源:基于本文分析绘制的示意图。
主导层:L2 工具与执行层 + L3 编排与控制层。瓶颈层:L5 评估与观测层 + L6 治理与安全层。
| 层 | 在智能制造方向的体现 | 关键约束 |
|---|---|---|
| L1 上下文工程 | 工艺文档、设备手册、检验标准等非结构化知识向量化;多源 OT/IT 数据口径统一。ISO 23247-3 的"可观测制造元素基本信息属性列表"可直接作为上下文 Schema | 上下文必须有 Schema,不能是自由文本堆砌 |
| L2 工具与执行 | MES / APS / 质检设备 / 机器视觉工具化;只读优先,写操作需工单语义(IEC 62264 L3 事务语义) | 工具权限以 L3 事务为单位定义 |
| L3 编排与控制 | 排产—质检—运维多智能体 DAG;GB/T 39116-2020 三级"集成级"(跨业务间数据共享)是编排的前提 | 未达集成级的产线不应上多智能体编排 |
| L4 记忆与状态 | 数字主线(ISO 23247-5:2026)作为跨会话状态载体;设备孪生状态持久化(AAS / IEC 63278-1:2023) | 状态不属于会话,属于工件与主线 |
| L5 评估与观测 | 良率、漏检率、OEE、排产达成率;GB/T 39116-2020 四级要求"对人员、资源、制造进行数据挖掘,形成知识、模型,实现核心业务精准预测与优化" | 本方向的落地瓶颈之一 |
| L6 治理与安全 | 生产数据不出厂;核心工艺参数严禁外传;操作留痕审计 | 本方向的落地瓶颈之一 |
需要特别说明的是,与 Media / Creative / AI 网剧 / AI 动画四个方向不同,智能制造的质量瓶颈不在 L4。制造过程的状态本来就由 MES、SCADA、数字主线持久化,模型不需要"记住"上一炉钢的温度。真正的困难在于:
- 指标体系不统一(L5):良率、OEE、漏检率在不同企业口径不同,缺乏可回归的评估集,导致"上了大模型"与"没上大模型"的差异难以被证明。
- 数据主权与可解释性(L6):产线数据是企业的核心资产,不愿出网;而黑箱模型给出的工艺参数调整建议,工艺工程师不敢采纳。
1.4 产业现状与已公开的量化口径
《智能制造能力成熟度模型》(GB/T 39116-2020) 已细化 228 项具体能力要求。据中国电子技术标准化研究院公开报道,已有 15 万家企业开展线上自评、1500 家企业完成现场评估。
下表为检索到的已公开量化口径,均标注来源与口径层级,引用时不得省略口径标注:
| 场景 | 指标 | 数值 | 来源与口径 |
|---|---|---|---|
| 排产 | 排产准确率 | 96%(青岛思锐卓远,工信部 2025 年人工智能应用典型案例) | 工信部案例 |
| 排产 | 排产准确率 | 50%~70% → 99%+(江苏茗鹤,扬州经开区"数字扬州"建设成果) | 地方遴选成果 |
| 排产 | 单次排产耗时 | 5 天 → 3 分钟(茗鹤某全球电子制造企业落地数据) | 同上 |
| 排产 | 计划员人数 | 10 人 → 2 人(同上) | 同上 |
| 排产 | 设备利用率 / 人力匹配度 | +17% / +15%(同上) | 同上 |
| 工艺 | 高炉炉温异常时间 | 减少 84.8%(兴澄特钢,部署 100 余个垂直模型) | 公开报道 |
| 工艺 | 产品检验不合格率 | 下降 47.3%(兴澄特钢) | 公开报道 |
| 工艺 | 热处理板交货周期 | 缩短 50%(兴澄特钢) | 公开报道 |
| 质检 | 最小可识别瑕疵 | 0.2 mm²(双鹿电池 × 华为云,单颗电池检测面积达 12 万 mm²) | 企业案例 |
| 质检 | 检测准确率 | 接近 100%(双鹿电池 × 华为云) | 企业案例 |
| 质检 | 单件检测时间 | 17 s → 8.5 s(长虹华意加西贝拉压缩机定子 360° 检测,缺陷类型 30 余种) | 公开报道 |
| 质检 | 少样本训练 | 每类缺陷仅需 1~5 个样本;标注时间缩减 80%+、训练周期压缩 80%+、开发成本降 66%(美芝工厂 × 中国联通) | 企业案例 |
| 质检 | 检验准确率 | 50%~60% → 98%+;疵点率 8% → 1.2%(海安市纺织质量数字化试点,投入 19 万元) | China Daily 报道 |
| 运维 | 故障预警提前期 | 7~15 天(先导智能,厂商口径);72 小时(涂装车间案例) | 见 5.2 节说明 |
说明:上述数据中,企业案例类多为厂商自述或地方遴选报道口径,属选择性披露,引用时应采用"据具体披露方(企业或机构名称)披露"表述而非断言。
2. 名词解释
| 术语 | 英文/缩写 | 释义 |
|---|---|---|
| 智能制造能力成熟度模型 | CMMM | GB/T 39116-2020 定义的五级成熟度:一级规划级、二级规范级、三级集成级、四级优化级、五级引领级 |
| 规划级 | Level 1 | CMMM 一级,开始规划智能制造基础条件,对设计、生产、物流、销售、服务等核心业务进行流程化管理 |
| 规范级 | Level 2 | CMMM 二级,采用自动化与信息技术改造核心装备与业务,实现单一业务活动内数据共享 |
| 集成级 | Level 3 | CMMM 三级,对装备与系统开展集成,实现跨业务间数据共享 |
| 优化级 | Level 4 | CMMM 四级,对人员、资源、制造进行数据挖掘,形成知识与模型,实现核心业务精准预测与优化 |
| 引领级 | Level 5 | CMMM 五级,基于模型持续驱动业务优化与创新,实现产业链协同并衍生新的制造模式和商业模式 |
| PTRM 四要素 | Personnel / Technology / Resource / Manufacturing | GB/T 39117-2020 的能力要素划分:人员、技术、资源、制造 |
| 数字孪生框架 | Digital Twin Framework | ISO 23247 系列定义的制造数字孪生框架,覆盖可观测制造元素的人员、设备、物料、工艺、设施、环境、产品与支持文档 |
| 数字主线 | Digital Thread | ISO 23247-5:2026 定义,覆盖数字孪生的创建、连接、管理与维护,贯穿全生命周期 |
| 数字孪生组合 | Digital Twin Composition | ISO 23247-6:2026 定义的三种组合方式:集成式(integrated)、统一式(unified)、联邦式(federated) |
| 可观测制造元素 | Observable Manufacturing Element | ISO 23247-3 定义的制造元素数字化表示,本方向可直接作为上下文 Schema |
| 企业—控制系统集成 | IEC 62264 / ISA-95 | 五层金字塔:L0 物理过程、L1 传感与执行、L2 监控控制、L3 制造运行管理(MES、排产)、L4 企业资源计划与物流 |
| 制造运行管理 | MOM | IEC 62264 L3 的核心内容,实践中常用 B2MML(XML)或 OPC UA 信息模型落地 |
| 资产管理壳 | AAS | 由 Header(资产与 AAS 的 URI 唯一标识)与 Body(Submodels,描述资产功能与属性)构成,现行标准为 IEC 63278-1:2023,由 IDTA 维护 |
| RAMI 4.0 | Reference Architecture Model Industrie 4.0 | DIN SPEC 91345:2016-04 定义的三维模型:Layers 六层、Life Cycle & Value Stream(IEC 62890)、Hierarchy Levels(基于 IEC 62264 与 IEC 61512) |
| 高级计划排程 | APS | 在订单交期、物料齐套、设备产能、人员排班等多重约束下求解生产计划的系统 |
| 约束理论 | Theory of Constraints, TOC | 以瓶颈资源为核心进行排产优化的方法,江苏茗鹤 APS 采用 TOC + AI 算法 |
| 设备综合效率 | OEE | 可用性 × 性能 × 质量的综合指标,本方向 L5 的核心观测指标之一 |
| 漏检率 | Miss Rate | 缺陷未被检出的比例,菲特(天津)案例降至 0.1% 以下 |
| 一次下线合格率 | First Pass Yield | 无需返修即通过检验的比例,菲特(天津)案例提升至 96% 以上 |
3. 案例
3.1 河钢集团唐钢公司 APS 智能排产
3.1.1 背景
河钢集团唐钢公司冷轧产线包含 1 条酸轧线、1 条连退线、3 条镀锌线,汽车钢品种多、工艺复杂。2015 年引进的国外排程系统只能做数据汇总,计划员每天需手工梳理 3~4 小时、制作 10 余张报表,排产方案高度依赖个人经验,难以应对多品种小批量的订单结构。
3.1.2 方案
2022 年 12 月 26 日自主开发的 APS 系统落地,作为"河北太行钢铁大模型"142 个专业模型之一。该大模型由河北省组织华为、阿里、百度等 13 家 AI 企业与 39 家钢铁企业共同攻关,对应原料、炼铁、炼钢、销售等 142 个场景。
从 Harness 视角拆解:
- L1:工艺约束、产线能力、订单结构被结构化为排产上下文,替代了计划员的经验记忆。
- L2:APS 与MES、检化验系统对接,排产结果以工单语义下发。
- L3:多产线(酸轧—连退—镀锌)协同排序,属于典型的多阶段 DAG 编排。
- L5:以准时交付率、库存周转、产线效率作为回归指标持续验证。
3.1.3 效果
- 排程生成时间从 3~4 小时压缩到 0.5 小时。
- 冷轧产线产品库存降低 15%。
- 生产效率提升 20%。
- 订单准时交付率 100%。
来源:河北日报 / 人民网,2025-04-16。口径为地方党媒对企业落地的报道。
3.2 菲特(天津)工业多模态垂类智能检测
3.2.1 背景
汽车零部件与整车漆面检测长期面临三难:缺陷种类多(划痕、颗粒、橘皮、流挂等)、标准不统一(人工目检受疲劳与主观影响)、漏检代价高(流入下工序或客户端)。传统机器视觉方案对每类缺陷都需要大量标注样本,换型成本高。
3.2.2 方案
采用"工业垂类大模型 + 多模态感知"融合架构,云—边—端协同:
- 端侧:高精度光学与结构光传感器采集。
- 边缘侧:专用 AI 推理引擎实时识别分类。
- 云端:训练垂类大模型,采用迁移学习与增量训练持续迭代。
该案例入选工业和信息化部 2025 年人工智能应用典型案例(全国遴选 285 项)。
从 Harness 视角拆解:
- L1:缺陷图谱与检验标准构成判定上下文,避免"模型自己定义什么是缺陷"。
- L2:端—边—云三级推理工具链,边缘侧负责实时判定,云端负责模型迭代。
- L5:漏检率与一次下线合格率构成明确的双指标评估闭环——注意二者存在张力,压低漏检率通常会抬高过杀率,必须双指标同时观测。
3.2.3 效果
- 质检漏检率降至 0.1% 以下。
- 一次下线合格率提升至 96% 以上。
- 单条产线年均节省人工成本 50% 以上。
- 服务比亚迪、大众、华润医药等 300 余家头部企业,国内市场占有率 33%(企业自述口径)。
来源:天津港保税区,2025(工信部典型案例公示)。市场占有率与节省比例为企业自述口径。
3.3 涂装车间智能工业数据分析系统
3.3.1 背景
涂装车间是整车制造中能耗最高、缺陷返工成本最高的环节之一。传统做法依赖人工经验调参,能耗与缺陷率之间存在复杂的耦合关系,单点优化容易顾此失彼。
3.3.2 方案
- 视觉层:2000 万像素工业相机 + YOLOv5 做钢码识别与 12 类喷涂缺陷检测。
- 运维层:振动、温度、电流多维状态监测 + DNN 预测性维护。
- 孪生层:工艺知识图谱与数字孪生,沉淀 300+ 工艺规则、200+ 设备模型。
该系统入选 2025 年"数据要素×"大赛优秀项目案例集。
从 Harness 视角拆解:
- L4:300+ 工艺规则与 200+ 设备模型本质上是持久化的领域状态,不依赖任何模型会话记忆。
- L5:能耗、误报率、缺陷率、返工成本构成多维评估面,而非单一准确率指标。
- L6:已在蔚来汽车、华晨宝马落地,涉及跨企业数据边界,需严格的租户与权限隔离。
3.3.3 效果
- 能耗降低 15%~20%,单车能耗下降 3.5 kWh,年省能源成本 150 万元。
- 故障预警提前 72 小时,误报率低于 3%,非计划停机减少 60%,年维护成本降低 45 万元。
- 钢码识别与 12 类喷涂缺陷检测准确率 98%~99%,喷涂缺陷率降低 80%,年减返工成本 200 万元。
- 已在蔚来汽车、华晨宝马落地,生产效率提升 20%~30%,产品一次合格率提高 5 个百分点,年节省成本超 263 万元,年减碳排放约 800 吨。
来源:吉林省政务服务和数字化建设管理局,2026-05(2025 年"数据要素×"大赛优秀项目案例集)。
4. 实践标准
4.1 AGENTS.md 规范
4.1.1. AGENTS.md(智能制造 · Manufacture 方向)
# AGENTS.md —— 智能制造(Manufacture)
## 角色与边界
- 你运行在制造执行环境之上,负责工艺建议、排产方案、质检判定与孪生查询四类任务。
- 你不得直接操控物理设备。所有作用于产线的指令必须经 MES/APS 以工单语义下发,并留痕。
- 你不得把核心工艺参数、配方、产线节拍数据输出到工厂网络之外。
- 判定标准由工艺工程师与质量工程师定义,你只负责执行与举证,不负责设定阈值。
## 环境假设
- 存在 MES、APS、SCADA、QMS、设备台账与工艺卡系统,通过私有 MCP 服务目录暴露只读接口。
- 存在机器视觉平台(相机 + 缺陷检测模型),可返回缺陷坐标、类型与置信度。
- 存在数字孪生与工艺规则库(本案例为 300+ 工艺规则、200+ 设备模型)。
- 模型推理为私有化部署或混合云边缘部署;生产数据不出厂。
- 存在 scripts/ 目录承载确定性计算:SPC 统计、OEE 计算、报表生成、导出封装。
## 上下文加载顺序(Context Budget)
1. 本文件(方向级)与组级 AGENTS.md
2. 本次任务的判定标准与阈值(不可裁剪)
3. 工艺卡、检验标准、BOM 与工艺路线(ISO 23247-3 可观测制造元素属性列表作为 Schema)
4. 相关设备台账与 AAS 子模型(IEC 63278-1:2023)
5. 历史批次数据与最近一次同类任务的检查点
6. 通用工艺知识(可裁剪)
规则:工艺卡与检验标准属于锚定上下文,任何情况下不得被裁剪。
## 工具契约
- 只读类(自由调用):查询工单、查询设备实时值、查询质检记录、查询孪生状态、检索工艺文档。
- 分析类(调用后必须举证):排产求解、缺陷分类、参数寻优、异常检测。
- 写操作类(工单语义 + 二次确认 + 留痕):下发排产、修改工艺参数、判定合格/不合格、创建工单。
- 写操作必须映射为 IEC 62264 L3 的事务语义,禁止绕过 MES 直接写入设备层。
- 确定性计算一律走 scripts/:OEE、SPC 控制限、良率统计、能耗核算、报表导出。禁止让模型逐 token 生成统计结果。
## 任务执行流程(SOP)
1. 确认任务类型(工艺 / 排产 / 质检 / 孪生)与判定阈值,写入任务台账。
2. 加载工艺卡与检验标准,锁定版本号。
3. 取数:明确时间窗、产线、批次、设备范围;记录取数 SQL 或接口调用。
4. 分析:调用分析类工具,输出候选方案与置信度。
5. 校验:对照判定阈值逐项比对;多目标冲突时(如良率 vs 节拍)显式列出权衡。
6. 人工签核:提交候选 + 证据,由工艺或质量工程师确认。
7. 下发:以工单语义写入 MES/APS,记录操作人与时间。
8. 归档:写入数字主线,保存版本、轨迹、成本与耗时。
## 验证与证据要求
- 每条结论必须附:数据来源(系统 + 时间窗 + 口径)、计算脚本路径、数值结果、与阈值的差值。
- 禁止只给"建议提高温度 5℃"这类无证据结论;必须给出依据批次与预期影响区间。
- 引用外部效果数据必须标注口径层级(官方 / 权威媒体 / 协会 / 厂商案例 / 研报 / 第三方估算)。
- 漏检率与过杀率必须成对报告,不得只报漏检率。
## 失败与升级策略
- 数据缺失或口径不一致:停止分析,输出《数据缺口清单》,升级到数据工程师。
- 分析结果落在阈值边界(±5%)内:标注为"边界结论",强制人工复核。
- 写操作二次确认未通过:立即中止,保持现场,生成待办工单。
- 同一问题连续 2 次求解不达标:升级到工艺工程师,并附上失败样本。
- 升级必须携带任务 ID、阶段、检查点、证据路径、已尝试处理。
## 安全与合规红线
- 生产数据不出厂;核心工艺参数与配方严禁外传。
- 不得打通 IT/OT 边界;与设备层的交互必须经 L3 及以上。
- 操作全量留痕,留痕记录纳入审计。
- 涉及安全联锁(SIS)的参数,智能体不得给出建议,一律转人工。
- 单次任务算力与调用预算设上限,超出即中断并上报。
## 禁止事项
- 禁止绕过 MES/APS 直接下发设备指令。
- 禁止把工艺参数、配方、客户订单信息写入外部服务或日志。
- 禁止在没有判定阈值的情况下给出合格/不合格结论。
- 禁止用模型记忆替代数字主线中的状态。
- 禁止编造标准编号与条款;不确定写 [待填写]。
## 输出格式
任务编号 / 任务类型 / 数据范围 / 判定阈值 / 分析结果(含置信度)/ 权衡说明 / 证据路径 / 待签核项 / 遗留问题
## 评估与自检
- 本次结论是否可复现(相同输入重跑是否得相同结果)?
- 漏检率与过杀率是否成对报告?
- 写操作是否都有工单号与操作人?
- 引用的每个数值是否标注了口径来源?
- 本次任务是否已纳入评估集(含失败样本)? 4.2 SKILL.md 规范
4.2.1. SKILL.md(智能制造 · 排产方案交付)
---
name: manufacture-aps-schedule-delivery
description: 智能制造排产交付技能。当需要在订单交期、物料齐套、设备产能、人员排班等多重约束下生成产线排产方案,并完成约束校验、冲突权衡说明、工单语义下发与归档时使用。适用于钢铁冷轧、离散制造混线、电子制造等多品种小批量场景。
version: 1.0
created: 2026-09-12
---
# 智能制造 · 排产方案交付
## 适用场景
- 在多重约束下生成覆盖全工序、每日、每班、每台设备的排产计划。
- 已有排产系统但准确率不足,需要 AI 重排并给出可解释的约束冲突说明。
- 需要把排产结果以工单语义下发到 MES/APS,并留痕归档。
## 前置条件
- 已加载方向级 AGENTS.md 与组级 AGENTS.md。
- 可访问 MES/APS 只读接口,能获取订单、BOM、工艺路线、设备状态、人员班次。
- 已确定排产目标优先级(交期优先 / 产能优先 / 库存优先),并经计划主管确认。
- 已确定硬约束清单(物料齐套、设备模具、安全库存、交期承诺)与软约束权重。
- 已确定具名签核人(计划主管)。
## 输入
| 输入项 | 说明 | 必需 |
|---|---|---|
| 排产时间窗 | 起止日期与班次粒度 | 是 |
| 订单清单 | 订单号、数量、交期、优先级、客户 | 是 |
| 工艺路线 | 工序序列、标准工时、换型时间 | 是 |
| 资源清单 | 设备、模具、人员、可用产能 | 是 |
| 约束配置 | 硬约束清单与软约束权重 | 是 |
| 目标优先级 | 交期 / 产能 / 库存的排序 | 是 |
## 输出
- 排产方案(工单序列 + 设备 + 班次 + 起止时间)
- 约束冲突说明(哪些约束被满足、哪些被妥协、代价是什么)
- 关键指标(计划准确率、产能利用率、预计库存、交期达成率)
- 校验证据(脚本路径、指标计算过程)
- 签核记录与归档记录
## 执行步骤
1. 读取订单与资源数据,记录取数时间窗与接口调用。
2. 校验数据完整性:缺物料齐套信息或缺设备状态则停止,输出缺口清单。
3. 加载硬约束与软约束权重,构建求解输入。
4. 调用求解器生成候选方案 2~3 套(不同目标优先级)。
5. 用 scripts/ 中的计算脚本核算指标:产能利用率、交期达成率、预计库存、换型次数。
6. 生成约束冲突说明:明确指出被妥协的约束及其代价。
7. 提交计划主管签核。
8. 签核通过后,以工单语义写入 APS/MES,记录工单号与操作人。
9. 归档到数字主线,保存方案版本、求解输入、指标与轨迹。
## 质量标准(DoD)
- 全部硬约束被满足,或每一条被妥协的硬约束都有显式说明与签核。
- 指标由脚本计算而非模型生成,脚本路径可追溯。
- 方案可复现:相同输入重跑得相同结果。
- 工单下发有工单号、操作人、时间戳。
- 引用基准数据(如"排产准确率 50%~70% → 99%+")时标注来源层级。
## 常见失败与处理
- 数据口径不一致(订单单位与 BOM 单位不符):停止求解,回到数据治理环节。
- 求解超时:缩小时间窗或降低求解精度,并在输出中标注精度下降。
- 多目标冲突无解:不要强行给方案,输出冲突矩阵交人工决策。
- 换型时间被忽略导致方案不可执行:把换型时间列入硬约束后重求解。
- 计划准确率提升但库存上升:说明存在目标冲突,需人工确认优先级。
## 示例
任务:某冷轧产线(1 条酸轧线、1 条连退线、3 条镀锌线)三日排产。
输入:订单 87 张、交期分布 3~15 天、设备可用率 92%、班次 3 班。
执行:求解 2 套方案(交期优先 / 库存优先),脚本核算指标,输出冲突矩阵。
结果:交期优先方案准时交付率 100%、库存 +3%;库存优先方案准时交付率 94%、库存 -12%。
签核:计划主管选择交期优先方案,工单号 APS-2026xxxx,归档至数字主线。 4.3 Deployment Checklist
| No. | Check Item | Layer | Judgment | Note |
|---|---|---|---|---|
| M-01 | Enterprise maturity has been assessed per GB/T 39116-2020 and is not lower than level 3 "integration level" | L3 | Required | Lines that have not reached the integration level (cross-business data sharing) should not adopt multi-agent orchestration |
| M-02 | Process cards, inspection standards, and BOM are structured and carry version numbers | L1 | Required | Process documents without a version number are treated as unusable |
| M-03 | Context Schema is aligned with the ISO 23247-3 observable manufacturing element attribute list | L1 | Recommended | The international standard's attribute list can be reused directly |
| M-04 | Tool permissions are divided per IEC 62264 L3 transaction semantics | L2 | Required | Write operations must have a transaction boundary |
| M-05 | Read-only / analysis / write three tool categories are clearly graded | L2 | Required | — |
| M-06 | The DAG orchestration for scheduling, inspection, and operations has explicit rollback points | L3 | Required | One checkpoint per stage |
| M-07 | Cross-session state is stored in the digital thread (ISO 23247-5:2026) rather than in the model session | L4 | Required | Digital twin composition can be chosen per ISO 23247-6:2026 as integrated / unified / federated |
| M-08 | Equipment assets are expressed and persisted as AAS (IEC 63278-1:2023) | L4 | Recommended | Note that IEC PAS 63088 was withdrawn in 2024-09 and must no longer be cited as a current standard |
| M-09 | The indicator system includes yield, miss rate, over-reject rate, OEE, and schedule attainment | L5 | Required | Miss rate and over-reject rate must be observed as a pair |
| M-10 | An evaluation set has been established and includes failure samples | L5 | Required | An evaluation set containing only success samples is meaningless |
| M-11 | All effectiveness data is annotated with its calibration level | L5 | Required | Official / authoritative media / association / vendor case / research report / third-party estimate |
| M-12 | Production data never leaves the factory, with the model deployed privately or on hybrid cloud | L6 | Required | — |
| M-13 | Egress channels for core process parameters and recipes are sealed off | L6 | Required | Including logging and telemetry channels |
| M-14 | All operations leave full audit trails | L6 | Required | Including query-type operations |
| M-15 | Safety-interlock (SIS) related parameters are listed as restricted zones for agents | L6 | Required | Always route to human |
| M-16 | A cap is set on per-task compute and invocation budget | L6 | Recommended | — |
| M-17 | Deterministic computation (OEE, SPC, yield, energy) is scripted | L2 | Required | Prohibit the model from generating statistical results token by token |
| M-18 | The GB/T 39116 capability-subdomain calibration is unified internally | L5 | Recommended | Note two calibrations coexist; see the information-gap statement |
5. 总结
智能制造方向的 AI Harness 有三个不同于内容侧方向的特征:
第一,它不需要解决一致性问题,但需要解决可信性问题。制造过程的状态本来就由 MES、SCADA 与数字主线持久化,L4 不是瓶颈。真正卡住落地的是 L5——指标体系不统一、评估集缺失,导致"AI 到底带来了多少收益"无法被证明;以及 L6——数据主权与可解释性,导致工艺工程师不敢采纳黑箱建议。
第二,它拥有最完整的标准骨架可供对齐。GB/T 39116-2020 提供了成熟度分级,ISO 23247 系列提供了数字孪生与数字主线框架,IEC 62264 提供了 IT/OT 边界与事务语义,IEC 63278-1:2023 提供了资产表达方式。Harness 的设计应当是"对齐"而不是"发明"。特别地,GB/T 39116-2020 三级"集成级"是上多智能体编排的前置条件,这一条在工程上经常被忽略。
第三,它的收益是真实的但口径复杂。唐钢 APS 把排程时间从 3~4 小时压到 0.5 小时、库存降低 15%、效率提升 20%、准时交付率 100%;菲特(天津)把漏检率压到 0.1% 以下、一次下线合格率提到 96% 以上。这些数字来自地方党媒报道、工信部典型案例与企业自述,口径层级不同,引用时必须标注,不得统一提升为"行业普遍水平"。
信息缺口声明
| 缺口项 | 处理方式 |
|---|---|
| GB/T 39116-2020 / GB/T 39117-2020 的能力子域口径 | 检索到两种口径并存:「PTRM 四要素 → 8 个能力域 → 20 个能力子域」与「10 大类核心能力 → 27 个要素域」。本文统一采用前者,并在此注明另一口径存在 |
| ISO 23247-2 的正式出版年份 | 已确认 Part 1:2021、Part 4:2021、Part 5:2026、Part 6:2026;Part 2 出版年份未核实 → |
| ISO 23247-3 的出版年份 | 中国采标计划(计划号 20240667-T-604)显示等同采用 ISO 23247-3:2021,可据此推定为 2021,但未见 ISO 官方页面直接确认 → 标注为"推定 2021" |
| 宁德时代、海尔、美的等企业 AI 排产数据 | 检索到的数值来源为个人知识库聚合页面,非企业公告或权威媒体,可信度不足 → 本文不采用 |
| 涂装车间案例的"生产效率提升 20%~30%"是否含其他技改贡献 | 案例集未拆分 AI 与其他技改的贡献度 → 引用时不得单独归因于 AI |
| 先导智能 7~15 天预警提前期 | 属厂商口径,见 02-industry.md 的处理 |
6. 参考资料
- 智能制造能力成熟度评估介绍(GB/T 39116-2020、GB/T 39117-2020)— 中国电子技术标准化研究院。https://www.cc.cesi.cn/service/show-2478.aspx
- ISO 23247-1:2021 Automation systems and integration — Digital twin framework for manufacturing — Part 1: Overview and general principles — ISO。https://www.iso.org/standard/75066.html
- ISO 23247-5:2026 Digital twin framework for manufacturing — Part 5: Digital thread — ISO。https://www.iso.org/standard/87425.html
- ISO 23247-6:2026 Digital twin framework for manufacturing — Part 6: Digital twin composition — ISO。https://www.iso.org/standard/87426.html
- 自动化系统与集成 面向制造的数字孪生框架 第3部分:制造元素的数字表示(国家标准计划,等同采用 ISO 23247-3:2021)— 全国标准信息公共服务平台。https://std.samr.gov.cn/gb/search/gbDetailed?id=E18C68D8D2A34FFDE05397BE0A0AE2CE
- ISA-95 架构与层 — 西门子。https://www.siemens.com/zh-tw/technology/isa-95-framework-layers
- White Paper: Terminal Automation(IEC 62264 分部结构)— TIC4.0,2025。https://tic40.org/wp-content/uploads/2025/12/White-Paper-Terminal-Automation-1.pdf
- RAMI 4.0: The Reference Architecture Model for Industrie 4.0 — ERP Information。https://www.erp-information.com/rami-4-0
- Structure of the Administration Shell(资产管理壳结构)— Plattform Industrie 4.0 / ZVEI,2016。https://industrialdigitaltwin.org/wp-content/uploads/2021/09/01_structure_of_the_administration_shell_en_2016.pdf
- 钢企"最强大脑"是怎样炼成的(河钢唐钢 APS)— 河北日报 / 人民网,2025。https://he.people.com.cn/n2/2025/0416/c192235-41197850.html
- 本市2项案例入选工信部人工智能应用典型案例(菲特天津)— 天津港保税区,2025。https://www.tjftz.gov.cn/contents/6302/381807.html
- 2025年"数据要素×"大赛优秀项目案例集 · 涂装车间智能工业数据分析系统 — 吉林省政务服务和数字化建设管理局,2026。https://zsj.jl.gov.cn/ztzl/2026sjysdsjlsfs/yxal/202605/t20260519_3631673.html
- Hai'an launches AI-powered quality inspection pilot — China Daily,2025。https://subsites.chinadaily.com.cn/nantong/2025-12/16/c_1149318.htm
- 青岛经开区两案例跻身工信部典型案例名单(思锐卓远、青岛港)— 青岛经济技术开发区,2026。http://qda.qingdao.gov.cn/dtyw/dtxx_114/202607/t20260727_10685848.shtml
- AGENTS.md 官方站 — Agentic AI Foundation(Linux Foundation)。https://agents.md/
- Agent Skills Specification — agentskills.io。https://agentskills.io/specification
AI Harness in Smart Manufacturing (Manufacture)
1. Introduction
1.1 Background
Smart manufacturing is one of the industry directions with the highest barrier to entry for AI Harness adoption. The root cause is not insufficient model capability, but the physicality of the acceptance criteria: a single scheduling error costs inventory and delivery dates, a single missed inspection costs a recall, and a single mis-tuned process parameter costs a batch of scrap. These costs cannot be absorbed by "trying again," so smart manufacturing's requirements on Harness naturally lean toward observability (L5) and constraint-governance (L6).
From a standards perspective, smart manufacturing is one of the most heavily standardized manufacturing paradigms globally, and it has already formed a three-tier, citable specification skeleton:
- Maturity tier: 《Smart Manufacturing Capability Maturity Model》(GB/T 39116-2020) and 《Smart Manufacturing Capability Maturity Assessment Methods》(GB/T 39117-2020), which define the five maturity levels and assessment methods.
- Digital twin tier: the ISO 23247 series 《Automation systems and integration — Digital twin framework for manufacturing》, Part 1:2021 overview and general principles, Part 4:2021 information exchange, Part 5:2026 digital thread, Part 6:2026 digital twin composition.
- Integration and asset tier: IEC 62264 (equivalent to ANSI/ISA-95) enterprise—control system integration five-layer model; RAMI 4.0 (DIN SPEC 91345:2016-04) reference architecture model and the Asset Administration Shell AAS (IEC 63278-1:2023).
The significance of this skeleton is that it has already defined what a manufacturing system should look like. AI Harness does not need to reinvent the wheel; it only needs to align its own six-layer capabilities to this skeleton.
1.2 Definition and Scope
AI Harness in this direction refers to the engineered delivery layer that, in smart manufacturing scenarios, undertakes the four task types of process, scheduling, quality inspection, and digital twin, transforming model output into manufacturing instructions and judgment results that can be issued, traced, and regression-verified.
| Task Domain | Typical Decisions | Output Artifacts | Cost of Error |
|---|---|---|---|
| Process | Process parameter optimization, recipe recommendation, process rule accumulation | Process cards, parameter suggestions, rule-base entries | Batch scrap, quality fluctuation |
| Scheduling | Job sequencing and resource allocation under multiple constraints | Scheduling plans, work-order sequences | Inventory backlog, delivery-date breaches |
| Quality inspection | Defect detection and judgment, balancing misses against over-rejects | Judgment results, defect maps, SPC data | Recalls, customer claims |
| Digital twin | Virtual verification, real-time control support, in-process adaptation | Twin state, simulation conclusions, anomaly-detection results | Decision errors, downtime losses |
1.3 Positioning within the AI Harness System
图 1-1|智能制造 AI Harness 六层定位与瓶颈
数据来源:基于本文分析绘制的示意图。
Leading layers: L2 Tools & Execution + L3 Orchestration & Control. Bottleneck layers: L5 Evaluation & Observability + L6 Governance & Security.
| Layer | Expression in Smart Manufacturing | Key Constraints |
|---|---|---|
| L1 Context Engineering | Vectorize unstructured knowledge such as process documents, equipment manuals, and inspection standards; unify the calibration of multi-source OT/IT data. ISO 23247-3's "basic information attribute list of observable manufacturing elements" can serve directly as the context Schema | Context must have a Schema; it cannot be a pile of free text |
| L2 Tools & Execution | Toolize MES / APS / inspection equipment / machine vision; read-only first, write operations require work-order semantics (IEC 62264 L3 transaction semantics) | Tool permissions are defined per L3 transaction |
| L3 Orchestration & Control | Multi-agent DAG for scheduling—inspection—operations; the GB/T 39116-2020 level-3 "integration level" (data sharing across businesses) is the prerequisite for orchestration | Lines that have not reached the integration level should not adopt multi-agent orchestration |
| L4 Memory & State | Digital thread (ISO 23247-5:2026) as the cross-session state carrier; persist equipment twin state (AAS / IEC 63278-1:2023) | State belongs not to the session but to artifacts and the thread |
| L5 Evaluation & Observability | Yield, miss rate, OEE, schedule attainment; the GB/T 39116-2020 level-4 requirement reads "conduct data mining on personnel, resources, and manufacturing to form knowledge and models, achieving accurate prediction and optimization of core business" | One of the adoption bottlenecks in this direction |
| L6 Governance & Security | Production data never leaves the factory; core process parameters must not be transmitted out; operations leave audit trails | One of the adoption bottlenecks in this direction |
It should be particularly noted that, unlike the four directions of Media / Creative / AI Web Drama / AI Animation, smart manufacturing's quality bottleneck is not at L4. The state of the manufacturing process is already persisted by MES, SCADA, and the digital thread; the model does not need to "remember" the temperature of the previous batch of steel. The real difficulty lies in:
- Inconsistent indicator systems (L5): Yield, OEE, and miss rate are measured differently across enterprises, and there is no regression-verifiable evaluation set, which makes it hard to prove the difference between "with a large model" and "without a large model."
- Data sovereignty and explainability (L6): Production-line data is the enterprise's core asset and companies are unwilling to let it leave the network; meanwhile, process engineers are reluctant to adopt process-parameter adjustment suggestions from black-box models.
1.4 Industry Status and Published Quantitative Metrics
《Smart Manufacturing Capability Maturity Model》(GB/T 39116-2020) has broken the requirements down into 228 specific capability requirements. According to public reports from the China Electronics Standardization Institute, more than 150,000 enterprises have conducted online self-assessment and 1,500 enterprises have completed on-site assessment.
The table below lists the published quantitative metrics found; all are annotated with source and calibration level, and citation must not omit the calibration note:
| Scenario | Metric | Value | Source & Calibration |
|---|---|---|---|
| Scheduling | Scheduling accuracy | 96% (Qingdao Siruizhuoyuan, MIIT 2025 typical AI application cases) | MIIT case |
| Scheduling | Scheduling accuracy | 50%~70% → 99%+ (Jiangsu Minghe, Yangzhou Economic Development Zone "Digital Yangzhou" construction results) | Local selection results |
| Scheduling | Per-scheduling time | 5 days → 3 minutes (data from Minghe's deployment at a global electronics manufacturer) | Same as above |
| Scheduling | Number of planners | 10 people → 2 people (same as above) | Same as above |
| Scheduling | Equipment utilization / manpower matching | +17% / +15% (same as above) | Same as above |
| Process | Blast-furnace temperature anomaly time | Reduced 84.8% (Xingcheng Special Steel, deploying 100+ vertical models) | Public report |
| Process | Product inspection rejection rate | Down 47.3% (Xingcheng Special Steel) | Public report |
| Process | Heat-treated plate delivery cycle | Shortened 50% (Xingcheng Special Steel) | Public report |
| Quality inspection | Smallest detectable defect | 0.2 mm² (Shuanglu Battery × Huawei Cloud, with a single battery's inspection area reaching 120,000 mm²) | Enterprise case |
| Quality inspection | Inspection accuracy | Near 100% (Shuanglu Battery × Huawei Cloud) | Enterprise case |
| Quality inspection | Per-unit inspection time | 17 s → 8.5 s (Changhong Huayi & Jiaxipera compressor stator 360° inspection, 30+ defect types) | Public report |
| Quality inspection | Few-shot training | Only 1~5 samples per defect type; annotation time cut 80%+, training cycle compressed 80%+, development cost down 66% (Meizhi factory × China Unicom) | Enterprise case |
| Quality inspection | Inspection accuracy | 50%~60% → 98%+; defect rate 8% → 1.2% (Hai'an textile quality digitalization pilot, investing RMB 190,000) | China Daily report |
| Operations | Fault-warning lead time | 7~15 days (Lead Intelligent, vendor calibration); 72 hours (paint-shop case) | See Section 5.2 |
Note: Among the data above, enterprise-case items are mostly vendor self-reported or local-selection report calibration, which constitutes selective disclosure; citation should use the phrasing "as disclosed by the specific disclosing party (enterprise or institution name)" rather than assertion.
2. Glossary
| Term | English/Abbreviation | Definition |
|---|---|---|
| Smart Manufacturing Capability Maturity Model | CMMM | The five maturity levels defined by GB/T 39116-2020: Level 1 Planning, Level 2 Standardized, Level 3 Integrated, Level 4 Optimizing, Level 5 Leading |
| Planning level | Level 1 | CMMM level 1: begin planning the foundational conditions of smart manufacturing and carry out process-based management of core businesses such as design, production, logistics, sales, and service |
| Standardized level | Level 2 | CMMM level 2: adopt automation and information technology to upgrade core equipment and businesses, achieving data sharing within a single business activity |
| Integration level | Level 3 | CMMM level 3: integrate equipment and systems to achieve data sharing across businesses |
| Optimizing level | Level 4 | CMMM level 4: conduct data mining on personnel, resources, and manufacturing to form knowledge and models, achieving accurate prediction and optimization of core business |
| Leading level | Level 5 | CMMM level 5: continuously drive business optimization and innovation with models, achieving industry-chain collaboration and deriving new manufacturing and business models |
| PTRM four elements | Personnel / Technology / Resource / Manufacturing | Capability-element division in GB/T 39117-2020: personnel, technology, resource, manufacturing |
| Digital twin framework | Digital Twin Framework | The manufacturing digital twin framework defined by the ISO 23247 series, covering the personnel, equipment, materials, process, facilities, environment, products, and supporting documents of observable manufacturing elements |
| Digital thread | Digital Thread | Defined by ISO 23247-5:2026, covering the creation, connection, management, and maintenance of digital twins throughout the full life cycle |
| Digital twin composition | Digital Twin Composition | The three composition modes defined by ISO 23247-6:2026: integrated, unified, and federated |
| Observable manufacturing element | Observable Manufacturing Element | The digital representation of manufacturing elements defined by ISO 23247-3, which can be used directly as the context Schema in this direction |
| Enterprise—control system integration | IEC 62264 / ISA-95 | Five-layer pyramid: L0 physical process, L1 sensing and actuation, L2 monitoring and control, L3 manufacturing operations management (MES, scheduling), L4 enterprise resource planning and logistics |
| Manufacturing operations management | MOM | Core content of IEC 62264 L3, commonly implemented in practice via B2MML (XML) or OPC UA information models |
| Asset Administration Shell | AAS | Composed of a Header (unique URI identification of the asset and AAS) and a Body (Submodels describing asset functions and properties); the current standard is IEC 63278-1:2023, maintained by IDTA |
| RAMI 4.0 | Reference Architecture Model Industrie 4.0 | The three-dimensional model defined by DIN SPEC 91345:2016-04: Layers (six layers), Life Cycle & Value Stream (IEC 62890), Hierarchy Levels (based on IEC 62264 and IEC 61512) |
| Advanced Planning and Scheduling | APS | A system that solves production plans under multiple constraints such as order delivery dates, material availability, equipment capacity, and staffing schedules |
| Theory of Constraints | Theory of Constraints, TOC | A method that optimizes scheduling around the bottleneck resource; Jiangsu Minghe's APS uses TOC + AI algorithms |
| Overall Equipment Effectiveness | OEE | A composite metric of availability × performance × quality, one of the core observation indicators at L5 in this direction |
| Miss rate | Miss Rate | The proportion of defects not detected; reduced to below 0.1% in the Feite (Tianjin) case |
| First-pass yield | First Pass Yield | The proportion of units passing inspection without rework; raised to above 96% in the Feite (Tianjin) case |
3. Case Studies
3.1 HBIS Group Tangsteel APS Intelligent Scheduling
3.1.1 Background
HBIS Tangsteel's cold-rolling lines comprise 1 pickling-rolling line, 1 continuous annealing line, and 3 galvanizing lines, with many automotive-steel varieties and complex processes. The foreign scheduling system introduced in 2015 could only aggregate data; planners had to manually sort through 3~4 hours of work each day and produce more than 10 reports, and scheduling plans relied heavily on personal experience, making it hard to cope with a high-mix, small-lot order structure.
3.1.2 Solution
On December 26, 2022, the independently developed APS system went live as one of the 142 specialized models of the "Hebei Taihang Iron & Steel Large Model." This large model was jointly developed by Hebei Province, organizing 13 AI companies including Huawei, Alibaba, and Baidu together with 39 steel enterprises, covering 142 scenarios across raw materials, ironmaking, steelmaking, sales, and more.
Breakdown from the Harness perspective:
- L1: Process constraints, line capabilities, and order structure are structured into the scheduling context, replacing the planner's experiential memory.
- L2: APS integrates with MES and inspection/test systems, and scheduling results are issued with work-order semantics.
- L3: Multi-line (pickling-rolling—continuous annealing—galvanizing) collaborative sequencing is a typical multi-stage DAG orchestration.
- L5: On-time delivery rate, inventory turnover, and line efficiency are used as regression metrics for continuous verification.
3.1.3 Results
- Scheduling generation time compressed from 3~4 hours to 0.5 hours.
- Cold-rolling line product inventory reduced 15%.
- Production efficiency up 20%.
- On-time order delivery rate 100%.
Source: Hebei Daily / People's Daily Online, 2025-04-16. Calibration reflects local party-media reporting of the enterprise's deployment.
3.2 Feite (Tianjin) Industrial Multimodal Vertical Intelligent Inspection
3.2.1 Background
Inspection of automotive parts and finished-vehicle paint surfaces has long faced three difficulties: many defect types (scratches, particles, orange peel, sags, etc.), non-uniform standards (manual visual inspection is affected by fatigue and subjectivity), and high cost of misses (defects passing to downstream processes or to the customer). Traditional machine-vision approaches require a large number of annotated samples per defect type and incur high changeover costs.
3.2.2 Solution
It adopts a fused architecture of "industrial vertical large model + multimodal perception" with cloud—edge—terminal collaboration:
- Terminal side: high-precision optical and structured-light sensors for acquisition.
- Edge side: a dedicated AI inference engine performs real-time recognition and classification.
- Cloud side: trains the vertical large model, using transfer learning and incremental training for continuous iteration.
This case was selected as a typical AI application case of the Ministry of Industry and Information Technology (MIIT) in 2025 (285 cases selected nationwide).
Breakdown from the Harness perspective:
- L1: The defect map and inspection standards constitute the judgment context, avoiding "the model defining what a defect is by itself."
- L2: A terminal—edge—cloud three-tier inference toolchain, where the edge side handles real-time judgment and the cloud side handles model iteration.
- L5: The miss rate and first-pass yield form a clear two-metric evaluation loop — note the tension between them: lowering the miss rate usually raises the over-reject rate, so both metrics must be observed simultaneously.
3.2.3 Results
- Inspection miss rate reduced to below 0.1%.
- First-pass yield raised to above 96%.
- Annual labor-cost savings of more than 50% per production line.
- Serving leading companies such as BYD, Volkswagen, China Resources Pharmaceutical, and over 300 others, with a 33% domestic market share (vendor self-reported calibration).
Source: Tianjin Port Free Trade Zone, 2025 (publicity of MIIT typical cases). Market share and savings ratios are vendor self-reported calibration.
3.3 Intelligent Industrial Data Analysis System for Paint Shops
3.3.1 Background
The paint shop is one of the stages in complete-vehicle manufacturing with the highest energy consumption and the highest defect-rework cost. Traditional practice relies on manual, experience-based parameter tuning; energy consumption and defect rate are coupled in a complex way, so point optimization often sacrifices one for the other.
3.3.2 Solution
- Vision layer: a 20-megapixel industrial camera + YOLOv5 for steel-code recognition and detection of 12 types of spray defects.
- Operations layer: multi-dimensional condition monitoring of vibration, temperature, and current + DNN-based predictive maintenance.
- Twin layer: process knowledge graph and digital twin, accumulating 300+ process rules and 200+ equipment models.
The system was selected for the 2025 "Data Elements ×" competition's collection of excellent project cases.
Breakdown from the Harness perspective:
- L4: The 300+ process rules and 200+ equipment models are essentially persisted domain state, depending on no model session memory.
- L5: Energy consumption, false-alarm rate, defect rate, and rework cost form a multi-dimensional evaluation surface, rather than a single accuracy metric.
- L6: Already deployed at NIO and BMW Brilliance, it involves cross-enterprise data boundaries and requires strict tenant and permission isolation.
3.3.3 Results
- Energy consumption reduced 15%~20%, per-vehicle energy down 3.5 kWh, annual energy-cost savings of RMB 1.5 million.
- Fault warnings issued 72 hours in advance, with a false-alarm rate below 3%, unplanned downtime reduced 60%, and annual maintenance-cost savings of RMB 450,000.
- Steel-code recognition and 12-type spray-defect detection accuracy of 98%~99%, spray-defect rate reduced 80%, and annual rework-cost savings of RMB 2 million.
- Already deployed at NIO and BMW Brilliance, production efficiency up 20%~30%, product first-pass yield up 5 percentage points, annual savings exceeding RMB 2.63 million, and annual carbon-emission reduction of about 800 tonnes.
Source: Jilin Provincial Administration of Government Services and Digitalization, 2026-05 (2025 "Data Elements ×" competition's collection of excellent project cases).
4. Practice Standards
4.1 AGENTS.md Specification
4.1.1. AGENTS.md (Smart Manufacturing · Manufacture Direction)
# AGENTS.md —— 智能制造(Manufacture)
## 角色与边界
- 你运行在制造执行环境之上,负责工艺建议、排产方案、质检判定与孪生查询四类任务。
- 你不得直接操控物理设备。所有作用于产线的指令必须经 MES/APS 以工单语义下发,并留痕。
- 你不得把核心工艺参数、配方、产线节拍数据输出到工厂网络之外。
- 判定标准由工艺工程师与质量工程师定义,你只负责执行与举证,不负责设定阈值。
## 环境假设
- 存在 MES、APS、SCADA、QMS、设备台账与工艺卡系统,通过私有 MCP 服务目录暴露只读接口。
- 存在机器视觉平台(相机 + 缺陷检测模型),可返回缺陷坐标、类型与置信度。
- 存在数字孪生与工艺规则库(本案例为 300+ 工艺规则、200+ 设备模型)。
- 模型推理为私有化部署或混合云边缘部署;生产数据不出厂。
- 存在 scripts/ 目录承载确定性计算:SPC 统计、OEE 计算、报表生成、导出封装。
## 上下文加载顺序(Context Budget)
1. 本文件(方向级)与组级 AGENTS.md
2. 本次任务的判定标准与阈值(不可裁剪)
3. 工艺卡、检验标准、BOM 与工艺路线(ISO 23247-3 可观测制造元素属性列表作为 Schema)
4. 相关设备台账与 AAS 子模型(IEC 63278-1:2023)
5. 历史批次数据与最近一次同类任务的检查点
6. 通用工艺知识(可裁剪)
规则:工艺卡与检验标准属于锚定上下文,任何情况下不得被裁剪。
## 工具契约
- 只读类(自由调用):查询工单、查询设备实时值、查询质检记录、查询孪生状态、检索工艺文档。
- 分析类(调用后必须举证):排产求解、缺陷分类、参数寻优、异常检测。
- 写操作类(工单语义 + 二次确认 + 留痕):下发排产、修改工艺参数、判定合格/不合格、创建工单。
- 写操作必须映射为 IEC 62264 L3 的事务语义,禁止绕过 MES 直接写入设备层。
- 确定性计算一律走 scripts/:OEE、SPC 控制限、良率统计、能耗核算、报表导出。禁止让模型逐 token 生成统计结果。
## 任务执行流程(SOP)
1. 确认任务类型(工艺 / 排产 / 质检 / 孪生)与判定阈值,写入任务台账。
2. 加载工艺卡与检验标准,锁定版本号。
3. 取数:明确时间窗、产线、批次、设备范围;记录取数 SQL 或接口调用。
4. 分析:调用分析类工具,输出候选方案与置信度。
5. 校验:对照判定阈值逐项比对;多目标冲突时(如良率 vs 节拍)显式列出权衡。
6. 人工签核:提交候选 + 证据,由工艺或质量工程师确认。
7. 下发:以工单语义写入 MES/APS,记录操作人与时间。
8. 归档:写入数字主线,保存版本、轨迹、成本与耗时。
## 验证与证据要求
- 每条结论必须附:数据来源(系统 + 时间窗 + 口径)、计算脚本路径、数值结果、与阈值的差值。
- 禁止只给"建议提高温度 5℃"这类无证据结论;必须给出依据批次与预期影响区间。
- 引用外部效果数据必须标注口径层级(官方 / 权威媒体 / 协会 / 厂商案例 / 研报 / 第三方估算)。
- 漏检率与过杀率必须成对报告,不得只报漏检率。
## 失败与升级策略
- 数据缺失或口径不一致:停止分析,输出《数据缺口清单》,升级到数据工程师。
- 分析结果落在阈值边界(±5%)内:标注为"边界结论",强制人工复核。
- 写操作二次确认未通过:立即中止,保持现场,生成待办工单。
- 同一问题连续 2 次求解不达标:升级到工艺工程师,并附上失败样本。
- 升级必须携带任务 ID、阶段、检查点、证据路径、已尝试处理。
## 安全与合规红线
- 生产数据不出厂;核心工艺参数与配方严禁外传。
- 不得打通 IT/OT 边界;与设备层的交互必须经 L3 及以上。
- 操作全量留痕,留痕记录纳入审计。
- 涉及安全联锁(SIS)的参数,智能体不得给出建议,一律转人工。
- 单次任务算力与调用预算设上限,超出即中断并上报。
## 禁止事项
- 禁止绕过 MES/APS 直接下发设备指令。
- 禁止把工艺参数、配方、客户订单信息写入外部服务或日志。
- 禁止在没有判定阈值的情况下给出合格/不合格结论。
- 禁止用模型记忆替代数字主线中的状态。
- 禁止编造标准编号与条款;不确定写 [待填写]。
## 输出格式
任务编号 / 任务类型 / 数据范围 / 判定阈值 / 分析结果(含置信度)/ 权衡说明 / 证据路径 / 待签核项 / 遗留问题
## 评估与自检
- 本次结论是否可复现(相同输入重跑是否得相同结果)?
- 漏检率与过杀率是否成对报告?
- 写操作是否都有工单号与操作人?
- 引用的每个数值是否标注了口径来源?
- 本次任务是否已纳入评估集(含失败样本)? 4.2 SKILL.md Specification
4.2.1. SKILL.md (Smart Manufacturing · Scheduling Plan Delivery)
---
name: manufacture-aps-schedule-delivery
description: 智能制造排产交付技能。当需要在订单交期、物料齐套、设备产能、人员排班等多重约束下生成产线排产方案,并完成约束校验、冲突权衡说明、工单语义下发与归档时使用。适用于钢铁冷轧、离散制造混线、电子制造等多品种小批量场景。
version: 1.0
created: 2026-09-12
---
# 智能制造 · 排产方案交付
## 适用场景
- 在多重约束下生成覆盖全工序、每日、每班、每台设备的排产计划。
- 已有排产系统但准确率不足,需要 AI 重排并给出可解释的约束冲突说明。
- 需要把排产结果以工单语义下发到 MES/APS,并留痕归档。
## 前置条件
- 已加载方向级 AGENTS.md 与组级 AGENTS.md。
- 可访问 MES/APS 只读接口,能获取订单、BOM、工艺路线、设备状态、人员班次。
- 已确定排产目标优先级(交期优先 / 产能优先 / 库存优先),并经计划主管确认。
- 已确定硬约束清单(物料齐套、设备模具、安全库存、交期承诺)与软约束权重。
- 已确定具名签核人(计划主管)。
## 输入
| 输入项 | 说明 | 必需 |
|---|---|---|
| 排产时间窗 | 起止日期与班次粒度 | 是 |
| 订单清单 | 订单号、数量、交期、优先级、客户 | 是 |
| 工艺路线 | 工序序列、标准工时、换型时间 | 是 |
| 资源清单 | 设备、模具、人员、可用产能 | 是 |
| 约束配置 | 硬约束清单与软约束权重 | 是 |
| 目标优先级 | 交期 / 产能 / 库存的排序 | 是 |
## 输出
- 排产方案(工单序列 + 设备 + 班次 + 起止时间)
- 约束冲突说明(哪些约束被满足、哪些被妥协、代价是什么)
- 关键指标(计划准确率、产能利用率、预计库存、交期达成率)
- 校验证据(脚本路径、指标计算过程)
- 签核记录与归档记录
## 执行步骤
1. 读取订单与资源数据,记录取数时间窗与接口调用。
2. 校验数据完整性:缺物料齐套信息或缺设备状态则停止,输出缺口清单。
3. 加载硬约束与软约束权重,构建求解输入。
4. 调用求解器生成候选方案 2~3 套(不同目标优先级)。
5. 用 scripts/ 中的计算脚本核算指标:产能利用率、交期达成率、预计库存、换型次数。
6. 生成约束冲突说明:明确指出被妥协的约束及其代价。
7. 提交计划主管签核。
8. 签核通过后,以工单语义写入 APS/MES,记录工单号与操作人。
9. 归档到数字主线,保存方案版本、求解输入、指标与轨迹。
## 质量标准(DoD)
- 全部硬约束被满足,或每一条被妥协的硬约束都有显式说明与签核。
- 指标由脚本计算而非模型生成,脚本路径可追溯。
- 方案可复现:相同输入重跑得相同结果。
- 工单下发有工单号、操作人、时间戳。
- 引用基准数据(如"排产准确率 50%~70% → 99%+")时标注来源层级。
## 常见失败与处理
- 数据口径不一致(订单单位与 BOM 单位不符):停止求解,回到数据治理环节。
- 求解超时:缩小时间窗或降低求解精度,并在输出中标注精度下降。
- 多目标冲突无解:不要强行给方案,输出冲突矩阵交人工决策。
- 换型时间被忽略导致方案不可执行:把换型时间列入硬约束后重求解。
- 计划准确率提升但库存上升:说明存在目标冲突,需人工确认优先级。
## 示例
任务:某冷轧产线(1 条酸轧线、1 条连退线、3 条镀锌线)三日排产。
输入:订单 87 张、交期分布 3~15 天、设备可用率 92%、班次 3 班。
执行:求解 2 套方案(交期优先 / 库存优先),脚本核算指标,输出冲突矩阵。
结果:交期优先方案准时交付率 100%、库存 +3%;库存优先方案准时交付率 94%、库存 -12%。
签核:计划主管选择交期优先方案,工单号 APS-2026xxxx,归档至数字主线。 4.3 Deployment Checklist
| 编号 | 检查项 | 层级 | 判定 | 说明 |
|---|---|---|---|---|
| M-01 | 已按 GB/T 39116-2020 评估企业成熟度等级,且不低于三级"集成级" | L3 | 必备 | 未达集成级(跨业务数据共享)不应上多智能体编排 |
| M-02 | 工艺卡、检验标准、BOM 已结构化,且带版本号 | L1 | 必备 | 无版本号的工艺文件视为不可用 |
| M-03 | 上下文 Schema 已对齐 ISO 23247-3 可观测制造元素属性列表 | L1 | 建议 | 可直接复用国际标准属性清单 |
| M-04 | 工具权限已按 IEC 62264 L3 事务语义划分 | L2 | 必备 | 写操作必须有事务边界 |
| M-05 | 只读 / 分析 / 写操作三类工具分级明确 | L2 | 必备 | — |
| M-06 | 排产、质检、运维的 DAG 编排有明确回滚点 | L3 | 必备 | 每阶段一个检查点 |
| M-07 | 跨会话状态存放在数字主线(ISO 23247-5:2026)而非模型会话 | L4 | 必备 | 数字孪生组合方式可按 ISO 23247-6:2026 选择集成式 / 统一式 / 联邦式 |
| M-08 | 设备资产以 AAS(IEC 63278-1:2023)表达并持久化 | L4 | 建议 | 注意 IEC PAS 63088 已于 2024-09 撤回,不得再引用为现行标准 |
| M-09 | 指标体系包含良率、漏检率、过杀率、OEE、排产达成率 | L5 | 必备 | 漏检率与过杀率必须成对观测 |
| M-10 | 已建立评估集并包含失败样本 | L5 | 必备 | 只放成功样本的评估集无意义 |
| M-11 | 所有效果数据标注口径层级 | L5 | 必备 | 官方 / 权威媒体 / 协会 / 厂商案例 / 研报 / 第三方估算 |
| M-12 | 生产数据不出厂,模型为私有化或混合云部署 | L6 | 必备 | — |
| M-13 | 核心工艺参数与配方的外传通道已封堵 | L6 | 必备 | 含日志与遥测通道 |
| M-14 | 操作全量留痕且可审计 | L6 | 必备 | 含查询类操作 |
| M-15 | 安全联锁(SIS)相关参数已列入智能体禁区 | L6 | 必备 | 一律转人工 |
| M-16 | 单次任务算力与调用预算已设上限 | L6 | 建议 | — |
| M-17 | 确定性计算(OEE、SPC、良率、能耗)已脚本化 | L2 | 必备 | 禁止模型逐 token 生成统计结果 |
| M-18 | GB/T 39116 能力子域口径已在内部统一 | L5 | 建议 | 注意两种口径并存,见信息缺口声明 |
5. Summary
AI Harness in smart manufacturing has three characteristics that differ from the content-side directions:
First, it does not need to solve the consistency problem, but it does need to solve the trustworthiness problem. The state of the manufacturing process is already persisted by MES, SCADA, and the digital thread, so L4 is not the bottleneck. What actually blocks deployment is L5 — inconsistent indicator systems and a missing evaluation set, making it impossible to prove "how much benefit AI actually brought" — and L6 — data sovereignty and explainability, which make process engineers reluctant to adopt black-box suggestions.
Second, it has the most complete standard skeleton to align to. GB/T 39116-2020 provides the maturity grading, the ISO 23247 series provides the digital twin and digital thread framework, IEC 62264 provides IT/OT boundaries and transaction semantics, and IEC 63278-1:2023 provides the asset representation. Harness's design should be about "alignment" rather than "invention." In particular, the GB/T 39116-2020 level-3 "integration level" is a prerequisite for adopting multi-agent orchestration, a point that is often overlooked in engineering practice.
Third, its benefits are real but the calibration is complex. Tangsteel's APS compressed scheduling time from 3~4 hours to 0.5 hours, cut inventory 15%, raised efficiency 20%, and achieved 100% on-time delivery; Feite (Tianjin) reduced the miss rate to below 0.1% and raised first-pass yield to above 96%. These figures come from local party-media reports, MIIT typical cases, and vendor self-reports with different calibration levels; citation must annotate them and must not elevate them uniformly to "industry-wide average levels."
Information-Gap Statement
| Gap Item | Handling |
|---|---|
| Calibration of capability subdomains in GB/T 39116-2020 / GB/T 39117-2020 | Two calibrations were found to coexist: 「PTRM four elements → 8 capability domains → 20 capability subdomains」 and 「10 major capability categories → 27 element domains」. This document uniformly adopts the former, and notes here that the other calibration also exists |
| Official publication year of ISO 23247-2 | Part 1:2021, Part 4:2021, Part 5:2026, Part 6:2026 confirmed; Part 2's publication year not verified → |
| Publication year of ISO 23247-3 | China's adoption plan (plan no. 20240667-T-604) shows identical adoption of ISO 23247-3:2021, so 2021 can be inferred, but no ISO official page directly confirms it → annotated as "inferred 2021" |
| AI scheduling data from CATL, Haier, Midea, and other enterprises | The retrieved numeric sources are personal-knowledge-base aggregation pages, not corporate announcements or authoritative media, with insufficient credibility → not adopted in this document |
| Whether the paint-shop case's "production efficiency up 20%~30%" includes other technical-reform contributions | The case collection did not split AI vs. other technical-reform contributions → citation must not attribute it solely to AI |
| Lead Intelligent's 7~15-day warning lead time | Vendor calibration; see the handling in 02-industry.md |
6. References
- Introduction to Smart Manufacturing Capability Maturity Assessment (GB/T 39116-2020, GB/T 39117-2020) — China Electronics Standardization Institute. https://www.cc.cesi.cn/service/show-2478.aspx
- ISO 23247-1:2021 Automation systems and integration — Digital twin framework for manufacturing — Part 1: Overview and general principles — ISO. https://www.iso.org/standard/75066.html
- ISO 23247-5:2026 Digital twin framework for manufacturing — Part 5: Digital thread — ISO. https://www.iso.org/standard/87425.html
- ISO 23247-6:2026 Digital twin framework for manufacturing — Part 6: Digital twin composition — ISO. https://www.iso.org/standard/87426.html
- Automation systems and integration — Digital twin framework for manufacturing Part 3: Digital representation of manufacturing elements (national standard plan, identical adoption of ISO 23247-3:2021) — National Public Service Platform for Standards Information. https://std.samr.gov.cn/gb/search/gbDetailed?id=E18C68D8D2A34FFDE05397BE0A0AE2CE
- ISA-95 architecture and layers — Siemens. https://www.siemens.com/zh-tw/technology/isa-95-framework-layers
- White Paper: Terminal Automation (IEC 62264 part structure) — TIC4.0, 2025. https://tic40.org/wp-content/uploads/2025/12/White-Paper-Terminal-Automation-1.pdf
- RAMI 4.0: The Reference Architecture Model for Industrie 4.0 — ERP Information. https://www.erp-information.com/rami-4-0
- Structure of the Administration Shell — Plattform Industrie 4.0 / ZVEI, 2016. https://industrialdigitaltwin.org/wp-content/uploads/2021/09/01_structure_of_the_administration_shell_en_2016.pdf
- How a steel company's "strongest brain" was forged (HBIS Tangsteel APS) — Hebei Daily / People's Daily Online, 2025. https://he.people.com.cn/n2/2025/0416/c192235-41197850.html
- This city's 2 cases selected into MIIT typical AI application cases (Feite Tianjin) — Tianjin Port Free Trade Zone, 2025. https://www.tjftz.gov.cn/contents/6302/381807.html
- 2025 "Data Elements ×" competition collection of excellent project cases · paint-shop intelligent industrial data analysis system — Jilin Provincial Administration of Government Services and Digitalization, 2026. https://zsj.jl.gov.cn/ztzl/2026sjysdsjlsfs/yxal/202605/t20260519_3631673.html
- Hai'an launches AI-powered quality inspection pilot — China Daily, 2025. https://subsites.chinadaily.com.cn/nantong/2025-12/16/c_1149318.htm
- Two cases from Qingdao Economic Development Zone on the MIIT typical-case list (Siruizhuoyuan, Qingdao Port) — Qingdao Economic and Technological Development Zone, 2026. http://qda.qingdao.gov.cn/dtyw/dtxx_114/202607/t20260727_10685848.shtml
- AGENTS.md official site — Agentic AI Foundation (Linux Foundation). https://agents.md/
- Agent Skills Specification — agentskills.io. https://agentskills.io/specification