可灵 AI 图像(快手)
1. 介绍
1.1. 平台概况
可灵 AI 由快手科技开发,图像线与视频线共用同一产品入口与同一套 Credits 计量体系。在 AI Harness 六层能力模型中,可灵的独特价值在于 L4 记忆与状态层的产品化做得最直白:它把"人物 / 商品 / 物体"抽象为主体创建(Element Creation)这一等公民资产,并用会员档位直接规定可持有数量(30 / 50 / 150 / 500)——这是本组中少见的"资产层被显式计价、显式限量"的设计。
同时,可灵的 Kling O1 Image 走的是"生成即对话"路线:不区分生成与编辑,上传参考 → 生成 → 用文本迭代修正,形成多轮会话式闭环。这让它在 L3 编排层上比 Midjourney 更进一步,但仍未达到工作流工件级别。
| 项 | 内容 | 置信度 |
|---|---|---|
| 开发商 | 快手科技(Kuaishou Technology) | 高 |
| 图像模型线 | 可图 Kolors(2024-05 发布,开源)、Kling O1 Image、Kolors 2.1;会员页显示 IMAGE 3.0 Omni & O1 | 高(官方会员页)+ 中(券商研报) |
| 开放形态 | Web(klingai.com)+ App;企业侧 API | 高 |
| 免费档商用 | 不可商用(官方会员页明确) | 高(官方) |
| 参考图上限 | Kling O1 Image 最多 10 张 | 中(渠道口径) |
1.2. 模型线谱系
| 模型 | 时间 / 状态 | 定位 | 开放性 |
|---|---|---|---|
| 可图 Kolors | 2024-05 发布 | 文生图 + 图生图,中文理解强 | 开源 |
| Kolors 2.1 | 后续迭代 | 写实渲染、色彩科学、人像皮肤纹理、图内文字 | 未检索到明确开源状态, |
| Kling O1 Image | 当前主力 | 统一多模态图像引擎:文生图 + 自然语言指令编辑 + 最多 10 张参考图的身份一致性保持 | 闭源 |
| IMAGE 3.0 Omni | 会员页显示 | 与 O1 并列提供 | 闭源 |
Kling O1 Image 官方与渠道口径:支持 18+ 任务类型、4K 输出、最多 10 张参考图。该组参数来自第三方聚合站营销页,非官方,。
1.3. 定价体系
1.3.1. 国际站会员(官方页,高置信)
| 档位 | 月付 | 原价 | Credits/月 | 年付 |
|---|---|---|---|---|
| Basic | $0 | — | — | — |
| Standard | $6.99/月 | $10 | 660 | $72/年 |
| Pro | $25.99/月 | $37 | 3,000 | $269/年 |
| Premier | $64.99/月 | $92 | 8,000 | $659/年 |
| Ultra | $127.99/月 | $180 | 26,000 | $1,296/年 |
官方换算口径:660 Credits ≈ 3,300 张图 或 33 个 720p 视频。由此反推,1 张图约 0.2 Credit 量级。
Credits 单购:3,500 Credits / $50;7,500 Credits / $100;16,000 Credits / $200。
1.3.2. 国内站会员(二次来源,中置信)
| 档位 | 下月续费价 | 灵感值 | 备注 |
|---|---|---|---|
| 黄金 | ¥58/月 | 660 | 首月 6 折起 |
| 铂金 | ¥234/月 | 3,000 | — |
| 钻石 | ¥586/月 | 8,000 | — |
| 黑金 | ¥1,149/月 | 26,000 | — |
来源为国盛证券研报截图。国内站与国际站的 Credits / 灵感值档位设计对齐(660 / 3,000 / 8,000 / 26,000),但价格体系独立。国内站价格以 App 内实时价为准。
口径说明(跨组并列):本篇 1.3.2 所记国内站会员价为黄金 ¥58 / 铂金 ¥234 / 钻石 ¥586 / 黑金 ¥1,149(灵感值 660 / 3,000 / 8,000 / 26,000);另在 03-市场研究/04-AI-漫剧组/02-kling-drama.md 中记录了另一套国内站口径:黄金 ¥19 / 铂金 ¥58 / 钻石 ¥198(灵感值 660 / 2,000 / 8,000)。国际站档位上限亦存在两个数字:本篇 1.3.1 的 Ultra 为 $127.99/月,02-kling-drama.md 的 Ultra 续费价为 $159.99/月。两套国内站口径均来自公开渠道转述,未获官方定价页直接确认,;差异可能对应不同时段或不同套餐体系,本篇不择一采信,引用时请以官方页实时价为准。
2. 名词解释
2.1. AI 图像通用术语
| 术语 | 英文 / 缩写 | 释义 | 可灵的对应实现 |
|---|---|---|---|
| 文生图 | Text-to-Image(T2I) | 仅由文本提示词生成图像 | 支持,O1 Image 主能力之一 |
| 图生图 | Image-to-Image(I2I) | 以一张或多张图像为条件生成新图像 | 支持,最多 10 张参考图 |
| 局部重绘 | Inpainting | 对指定区域重新生成,区域外保持不变 | 被自然语言指令编辑取代:描述区域即可,无需遮罩 |
| 外扩 | Outpainting | 在画布外扩区域继续生成 | 平台层未见独立外扩产品说明,[待填写] |
| 可控生成 | ControlNet | 以 Canny、Depth、Pose、Mask 等视觉信号控制生成结构 | 未见原生 ControlNet 说明;结构控制由参考图与指令承担 |
| 参考图 | Reference Image | 作为身份、风格、结构约束输入的图像 | 最多 10 张,经 Identity Map 处理 |
| 随机种子 | Seed | 固定后可在同参数下复现同一张图 | 平台层未公开种子锁定与快照端点说明,[待填写] |
| 引导强度 | CFG | 提示词对生成结果的约束强度 | 平台层未公开暴露方式,[待填写] |
| 低秩适配 | LoRA | 小参数量微调模块,用于固化人物、风格、服装资产 | 未见用户侧 LoRA 训练与加载 |
| 图像提示适配 | IP-Adapter | 用图像编码器特征注入注意力,实现"以图为提示词" | 未公开使用;功能由 Identity Map 承担 |
| 零样本身份注入 | InstantID | 单张参考图、无需微调即可迁移身份 | 未公开使用;O1 采用多图(5~10 张)构建稳定身份标记 |
2.2. 可灵特有术语
| 术语 | 英文 / 缩写 | 释义 |
|---|---|---|
| 灵感值 / Credits | Credits | 可灵统一计量单位,图片与视频共用;官方口径 660 Credits ≈ 3,300 张图或 33 个 720p 视频 |
| 主体创建 | Element Creation | 把人物、物体等创建为可复用主体的资产能力;会员档位限制数量为 30 / 50 / 150 / 500 |
| 身份映射 | Identity Map | 从 5~10 张不同角度、光线、距离的参考照中提取稳定身份标记(骨相、显著特征、比例),并与表情、穿搭等瞬时属性解耦 |
| 稳定身份标记 | Stable Identity Marker | 身份映射抽取的持久特征,与目标人物的表情、光线、穿搭等瞬时属性分离 |
| 瞬时属性 | Transient Attribute | 表情、穿搭、光线、妆造等随场景变化的特征,在身份映射中被显式解耦 |
| Kling O1 Image | — | 统一多模态图像引擎:文生图 + 自然语言指令编辑 + 最多 10 张参考图的身份一致性保持 |
| 主体智能补全 | Element Smart Completion | 视频 O1 系列的主体补全能力,各档位每天 3 次免费 |
| 图片画质增强 | Image Quality Enhancement | 会员档权益,Pro 及以上 |
| 可图 Kolors | Kolors | 2024-05 发布的开源图像模型,文生图 + 图生图 |
| 同时不限任务并发数 | Unlimited Concurrent Tasks | 会员档位权益:不限任务并发数 + 专享快速生成通道 |
2.3. 换装与换脸方向通用术语
说明:可灵在 1.5 版本新增 AI 模特能力,并通过"主体创建"支持商品与人物资产复用,因此在换装方向有实际触点;平台未提供独立的换脸产品能力。
| 术语 | 英文 / 缩写 | 释义 | 可灵相关性 |
|---|---|---|---|
| 虚拟试穿 | VTON(Virtual Try-On) | 将目标服装"穿"到指定人物图像上并生成视觉可信结果 | AI 模特 + 主体创建可承载,未见独立 VTON 产品线 |
| 服装形变 | Garment Warping | 用 TPS(薄板样条)等几何变换把平铺服装对齐到人体姿态,再送入生成 | 未见公开说明 |
| 服装掩码 | Cloth Mask | 人体解析得到的上衣、下装、外套区域二值图,用于限定重绘范围 | 不适用(自然语言指令编辑无需掩码) |
| 试穿扩散 | Try-on Diffusion | 以扩散模型端到端完成服装与人体融合 | 未见公开说明 |
| 换脸 | Face Swap | 把 A 的脸替换到 B 的面部位置 | 未提供独立换脸能力 |
| 人脸重演 | Face Reenactment | 保留身份、迁移表情、口型与头部姿态 | 视频侧 O1 系列有相关表述, |
| 身份保持 | Identity Preservation | 生成结果在多大程度上仍"是那个人" | 核心能力:Identity Map 即为此设计 |
3. 功能说明
3.1. 生成与编辑能力
| 能力 | 支持 | 说明 |
|---|---|---|
| 文生图 | 支持 | O1 Image 主能力 |
| 图生图 | 支持 | 最多 10 张参考图 |
| 自然语言指令编辑 | 支持 | "去掉背景""把外套改成酒红色""在窗边加雾",无需遮罩与选区 |
| 风格迁移 | 支持 | — |
| 局部编辑 | 支持 | 以文本描述区域,无需涂抹 |
| 图片超分 / 画质增强 | 支持 | Pro 及以上档位 |
| AI 模特 | 支持 | 1.5 版本新增 |
| 主体创建 | 支持 | 档位限制 30 / 50 / 150 / 500 |
| 视频生成 | 支持 | 与图像共用 Credits;O1 系列含主体智能补全 |
3.2. 自然语言指令编辑
可灵的编辑范式是"用说话代替涂抹":用户描述"把外套改成酒红色",模型自行定位目标区域并完成修改。这与 Midjourney 的 Vary Region(需先框选)、即梦智能画布(需先涂抹)形成三代交互差异:
| 平台 | 编辑交互 | 需要遮罩 | 需要重新上传参考 |
|---|---|---|---|
| Midjourney | Vary Region(框选 + 参数剥离) | 是 | 是 |
| 即梦 | 智能画布(涂抹 + 图层) | 是 | 否(画布内保留) |
| 可灵 O1 | 自然语言指令 | 否 | 否(会话内保留) |
工程含义:遮罩环节是人工错误与流程中断的高发点。取消遮罩意味着一次生成可被程序化连续修正,这是 L3 编排层向"可自动化"迈出的实质一步。但可灵仍未把这条链路固化为可导出的工作流工件。
3.3. 主体创建与 AI 模特
- 主体创建(Element):把人物、商品、物体固化为可复用主体,档位决定数量上限(30 / 50 / 150 / 500)。
- AI 模特(1.5 版本新增):结合主体创建,可支撑电商场景的虚拟模特复用。
- 每日灵感值赠送:会员每日赠送灵感值(具体数值 [待填写])。
3.4. 商用授权与并发
- 免费档(Basic)生成内容不可商用——这是官方会员页的明确条款,是本组中少见的、被官方明文写出的免费层商用限制,选型时须特别注意。
- Standard 及以上含商用权(官方会员页明确)。
- 并发:会员档位提供"同时不限任务并发数" + 专享快速生成通道。
- 水印:付费档去品牌水印。
4. 平台架构
图 4-1|可灵在 AI Harness 六层能力模型中的实现
数据来源:基于本文分析绘制的示意图。
4.1. 双模型并行
| 模型 | 侧重点 |
|---|---|
| Kling O1 Image | 统一生成 + 编辑 + 一致性;自然语言指令编辑;10 参考图 |
| Kolors 2.1 | 写实渲染、色彩科学、人像皮肤纹理、图内文字 |
两条线并行,说明快手并未把所有能力压到单一模型上,而是按"一致性 / 对话式编辑"与"写实渲染 / 文字"分工。这与 FLUX.2 的"模型选择即编排"(klein 实时 / pro 生产 / flex 排版 / max 最高质)在思路上相近,但可灵未把模型选择上升为可编排的路由策略——用户只能手动切换。
4.2. 文本编码器取舍
Kolors 采用 GLM 作为文本编码器,区别于 Imagen / SD3 常用的 T5。用中文原生大语言模型做编码器,是其强化中英文理解的底层原因。
该事实来自第三方技术梳理,未定位到原始论文,。
4.3. 开放形态
- Web(klingai.com)+ App。
- 企业侧 API(快手开放平台)。
- 被第三方平台(如 vivago.ai)与聚合平台集成,聚合生态对其分发有实质贡献。
5. Harness 设计
5.1. 六层能力总览
| 层 | 名称 | 可灵的实现 | 成熟度 | 证据强度 |
|---|---|---|---|---|
| L1 | 上下文工程 | 最多 10 张参考图 → Identity Map,把稳定身份标记与瞬时属性解耦后锚定 | 强 | 中(建议以官方文档复核) |
| L2 | 工具与执行 | 生成 / 编辑 / 风格迁移 / 超分 / 视频延长 / 主体创建,同一会话内以自然语言调用 | 中 | 中高 |
| L3 | 编排与控制 | "生成即对话":生成与编辑不区分,多轮会话式迭代闭环,无需重新上传、无需遮罩、无需重抽 | 中强 | 中 |
| L4 | 记忆与状态 | 主体创建(Element)即资产持久化,档位决定数量上限;会员每日赠送灵感值 | 中强 | 高(官方页) |
| L5 | 评估与观测 | Kolors 曾在智源 FlagEval 主观总分全球第二、主观画质第一(榜单时点值);无公开官方 Eval Set | 弱 | 中 |
| L6 | 治理与安全 | 免费档不可商用;付费档去水印;真人素材与身份一致性风险由会员协议约束 | 中 | 中 |
5.2. L1 上下文工程层
可灵的 L1 设计是本组中最具"上下文分层"自觉的一个:
- 输入层:最多 10 张参考图(渠道口径;建议以官方文档复核)。官方建议从 5~10 张不同角度、不同光线、不同距离的参考照中提取特征。
- 分层机制(Identity Map):把特征分为两类——
- 稳定身份标记:骨相、显著特征、比例等持久属性;
- 瞬时属性:表情、穿搭、光线、妆造等随场景变化的属性。
- 锚定逻辑:生成时只锚定前者,允许后者自由变化。
这正是参数卡中 L1 定义的"检索、压缩、缓存、优先级排序"里的最后一项——优先级排序。可灵把"什么必须保持一致、什么可以改变"做成了显式的产品语义,而不是让用户自己调权重(对比 Midjourney 的 --ow 单一数值旋钮)。
待补强项:无结构化提示词契约(对比 FLUX.2 的 JSON 字段)、无外部知识检索(对比 Seedream 5.0 的联网检索与 Nano Banana 的 Search Grounding)。
5.3. L2 工具与执行层
- 工具集:生成、编辑、风格迁移、超分、视频延长、主体创建。
- 调用形态:统一在同一会话内以自然语言调用,工具边界对用户不可见。
- 无开放工具注册:不支持 Function Calling、无 MCP Server(对比 Runway 于 2025-06 发布 MCP Server)。
- 成本口径:Credits 统一计量,图片与视频共用。官方给出 660 Credits ≈ 3,300 张图的换算,成本可测算性较好;但无生成前成本预估端点(对比 Leonardo.ai 的 Pricing Calculator)。
5.4. L3 编排与控制层
"生成即对话"是可灵在 L3 上的核心贡献:
| 传统流程 | 可灵 O1 流程 |
|---|---|
| 上传参考 → 生成 → 不满意 → 重新上传 → 调整参数 → 重抽 | 上传参考 → 生成 → 用文本批评 → 定向修改 → 继续批评 |
取消了三个环节:重新上传、遮罩涂抹、整图重抽。这带来两个直接收益:
- 上下文不丢失:参考图与已生成结果都留在会话内,避免"重开一局"导致的一致性断裂。
- 迭代可收敛:每次只改一处,多轮收敛到目标,而不是每轮重新掷骰子。
仍缺:无 DAG、无子智能体、无中断恢复、无工作流导出与版本控制。对比 ComfyUI 把工作流做成可版本控制的 JSON 图、美图设计室用 Agent Teams 做多智能体分工,可灵的 L3 停留在"人驱动的多轮对话",尚未进入"系统驱动的自动编排"。
5.5. L4 记忆与状态层
主体创建(Element)是本组中最直白的 L4 产品化设计:
| 平台 | L4 一等资产 | 计价 / 限量方式 |
|---|---|---|
| 可灵 | 主体创建(Element) | 档位限量 30 / 50 / 150 / 500 |
| 妙鸭相机 | 数字分身 | 单次制作(¥9.9 / ¥29.9),锁定在单一产品内、不可迁移 |
| Runway | Brand Kits / Voice Clones | Max 档最多各 3 个 |
| Nano Banana | 无专门抽象 | 依赖 Google 生态(Drive / NotebookLM) |
| 美图设计室 | 创作资产(官方明示可复用) | 未公开限量口径 |
可灵把"资产数量"直接写进会员权益表,意味着资产即产品价值。这在商业上成立,但也带来与妙鸭相同的风险:资产锁定在平台内。区别是可灵的资产形态(人物 / 商品 / 物体主体)比妙鸭的"数字分身"更通用,且平台仍在持续运营。
缺口:无 Checkpoint、无工件版本管理、无模型版本锁定端点、无资产导出机制说明。
5.6. L5 评估与观测层
- Kolors 曾在智源 FlagEval 主观总分全球第二、主观画质第一。该结果为特定榜单时点值,须标注时点,不应作为持续有效结论。
- 无公开官方 Eval Set、Golden Dataset 或回归集。
- 观测口径:Credits 消耗。无生成前成本预估,无 A/B 机制。
5.7. L6 治理与安全层
- 商用权分层:免费档不可商用(官方会员页明确),Standard 及以上含商用权。这是本组中最清晰的商用权分层设计之一。
- 水印:付费档去品牌水印。
- 肖像与身份风险:真人素材与身份一致性(Identity Map 明确以真人为参考对象)带来的肖像权风险,由平台会员协议约束,未见结构化的身份授权与护栏机制(对比 Nano Banana 的公众人物肖像限制、SynthID 强制水印)。
- 《标识办法》合规:未见官方公开的显式 / 隐式标识实现说明。[待填写]
5.8. 成熟度判断
可灵是"中强模型 + 中强 Harness"形态,其最大特色是 L4 资产化最直白(主体创建限量计价)与 L3 生成即对话(取消遮罩与重抽)。它比 Midjourney 更接近可交付形态,但距美图设计室的 Agent Teams 与 ComfyUI 的可版本控制工作流仍有明显差距。若把妙鸭视为"有资产但不可迁移"的失败样本,可灵则是"有资产且平台活跃"的对照组——资产是否可迁移与平台是否持续运营,共同决定 L4 投入的实际价值。
6. 实际案例
6.1. 官方客户案例检索结果
本次检索未找到可灵官方发布的、带量化效果数据的品牌客户案例。 按本组统一口径标注「无结果」,不使用"行业广泛使用"等模糊表述替代,也不虚构案例。
6.2. 渠道方工作流示例(不作为案例引用)
第三方渠道方(vivago.ai)给出的工作流示例:AI 虚拟网红——5 张参考图 → 120 场景/月,无摄影师与模特档期,商用权含在 Plus 计划内。
该内容属渠道营销材料,非可灵官方发布,无第三方验证,本报告不将其作为案例引用,仅作为"典型用法"参考,且其中的量化数字(120 场景/月)不应被转引为效果数据。
6.3. 可核实的用法与配比数据
可核实的事实型材料(非效果数据):
- 官方换算口径:660 Credits ≈ 3,300 张图或 33 个 720p 视频(官方会员页,高置信)——可用于单位成本测算。
- 主体创建档位限量:30 / 50 / 150 / 500(官方页,高置信)——可用于容量规划。
- 免费档不可商用(官方页,高置信)——可用于合规判定。
- Kolors 开源(2024-05)——可用于本地部署与二次开发评估。
7. 总结
7.1. 优势
- Identity Map 的上下文分层设计:显式区分稳定身份标记与瞬时属性,是本组 L1 层最清晰的语义设计。
- 生成即对话:取消遮罩与重抽,迭代可收敛,人工错误点显著减少。
- 主体创建(Element):资产一等公民化且限量计价,容量规划可预期。
- 商用权分层明确:免费档不可商用被官方明文写出,合规判定简单。
- Credits 换算口径官方给定:660 Credits ≈ 3,300 张图,单位成本可测算。
- Kolors 开源:为本地部署与二次开发留出通路。
7.2. 局限与适用边界
| 局限 | 影响 |
|---|---|
| 无工作流工件与版本管理 | 生成即对话的会话历史不可导出、不可回归 |
| 无模型版本锁定端点 | 模型升级后同参考同指令输出漂移 |
| L5 仅有时点榜单 | 质量改进不可量化 |
| 无开放工具注册 / MCP | 不能嵌入外部 Agent 工具链 |
| 主体资产不可迁移 | 与妙鸭同类风险,依赖平台持续运营 |
| 免费档不可商用 | 试用结果不能直接商用,须先升级 |
| 10 参考图 / 4K / 18+ 任务为渠道口径 | 参数承诺不可作为合同依据 |
适用边界:适合需要"同一人物 / 同一商品在多场景复用"的内容生产(虚拟网红、电商模特、IP 形象延展);不适合需要可编程编排、可复现回归、跨系统审计的企业级流水线。
7.3. 选型建议
- 先做容量规划再选档:按"每月主体数 + 每月图片数"两项反推档位,主体数是硬约束(30 / 50 / 150 / 500),图片数可用 Credits 单购补足。
- 商用前必须升级:免费档输出不可商用,不得用于任何对外素材。
- 自建可复现记录:记录"参考图集合哈希 + 指令序列 + 模型版本 + 主体 ID"四元组,作为回归与举证材料。
- 参数承诺须以官方为准:10 参考图、4K、18+ 任务类型等参数来自渠道营销页,签约前须向快手官方确认。
- 涉真人参考图:必须取得肖像权人书面授权,并保存授权链路,平台协议不能替代权利人对权利人的授权。
7.4. 合规提示
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) 第四条、第五条:服务提供者提供下载、复制、导出功能时应当确保显式标识,并应当在文件元数据中添加隐式标识。可灵侧未见公开实现说明,企业使用时应自行补建标识注入环节。
- 《中华人民共和国民法典》第一千零一十九条:不得以利用信息技术手段伪造等方式侵害他人肖像权。Identity Map 以真人参考照构建稳定身份标记,属于高度敏感的身份复用行为。
- 北京互联网法院 2026-03 生效判决确立"可识别性 + 举证责任转移":不要求与原肖像完全一致,公众能识别即构成使用特定自然人肖像;被告主张"偶然撞脸"须复现创作过程。使用多参考图身份锚定时,一旦生成结果与某自然人高度相似,举证责任在使用方。
信息缺口声明
- Kolors 使用 GLM 文本编码器的原始出处:仅第三方技术梳理提及,未定位到论文。
- Kling O1 Image 的"18+ 任务类型 / 4K 输出 / 10 张参考图":来自第三方聚合站营销页,非官方。
- 国内站会员价(¥58 / ¥234 / ¥586 / ¥1,149):来自券商研报截图,非官方页。
- 每日赠送灵感值数值:未检索到。[待填写]
- 官方客户案例与量化效果数据:未检索到。[无结果]
- 在中国《标识办法》下的显式与隐式标识实现细节:未见官方公开说明。[待填写]
- 种子锁定、CFG 暴露、模型版本锁定端点:未检索到官方说明。[待填写]
- Kolors 2.1 是否开源:未检索到明确说明。
- FlagEval 榜单的具体时点:未检索到榜单发布时点,,引用时须标注为时点值。
- 企业侧 API 的定价与并发规格:未检索到官方 API 定价页。[待填写]
8. 参考资料
- Kling AI 会员计划(官方页,含档位、Credits 换算与商用权条款) — 快手,2026。https://app.klingai.com/global/membership/membership-plan
- 可灵 AI 官网 — 快手。https://klingai.com/
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — 国家网信办、工业和信息化部、公安部、国家广播电视总局,2025-03-14 发布,2025-09-01 施行。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- 《中华人民共和国民法典》第一千零一十八条、第一千零一十九条 — 全国人民代表大会,2020。(正文引用条文,无官方链接)
- 经济参考报 ·《技术不是侵权"挡箭牌" 法院这样认定 AI"盗脸"》 — 新华社《经济参考报》,2026-04-17。http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
- InstantID 官方项目页(零样本身份注入,作为 Identity Map 的对照基准) — InstantX Team / 小红书 / 北京大学。https://instantid.github.io/
- DeepWiki · SDXL_EcomID_ComfyUI 5.4 对比(EcomID / InstantID / PuLID 控制机制对比) — 阿里妈妈。https://deepwiki.com/alimama-creative/SDXL_EcomID_ComfyUI/5.4-comparison-with-other-methods
- ComfyUI 官方工作流「虚拟角色试穿 - 四合一」(可版本控制工作流的对照基准) — Comfy Org。https://comfy.org/zh/workflows/templates_rob_fashion_shoot_vton-4in1.app/
- CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models — arXiv 2407.15886(虚拟试穿技术路线与评测指标)。https://arxiv.org/pdf/2407.15886
- Black Forest Labs · FLUX.2 Overview(模型路由与固定快照端点的对照基准) — BFL,2026。https://docs.bfl.ai/flux_2/flux2_overview
- Leonardo.Ai Pricing(官方页,Pricing Calculator 与容量约束的对照基准) — Leonardo Interactive。https://www.leonardo.ai/pricing
Kling AI Image (Kuaishou)
1. Introduction
1.1. Platform Overview
Kling AI is developed by Kuaishou Technology. The image line and the video line share the same product entry point and the same Credits metering system. In the AI Harness six-layer capability model, Kling's unique value lies in the fact that the L4 Memory & State layer has been productized in the most explicit way: it abstracts "people / products / objects" into Element Creation as first-class assets, and directly specifies the number that can be held per membership tier (30 / 50 / 150 / 500) — a rare design in this group where the "asset layer is explicitly priced and explicitly limited."
At the same time, Kling O1 Image follows the "generation as conversation" route: it does not distinguish generation from editing — upload reference → generate → iterate with text corrections, forming a multi-turn conversational loop. This pushes it one step further than Midjourney on the L3 orchestration layer, though it still falls short of workflow-level artifacts.
| Item | Content | Confidence |
|---|---|---|
| Developer | Kuaishou Technology | High |
| Image model line | 可图 Kolors (released 2024-05, open source), Kling O1 Image, Kolors 2.1; membership page shows IMAGE 3.0 Omni & O1 | High (official membership page) + Medium (broker research report) |
| Open format | Web (klingai.com) + App; enterprise-side API | High |
| Free-tier commercial use | Not permitted for commercial use (stated explicitly on the official membership page) | High (official) |
| Reference image limit | Kling O1 Image up to 10 images | Medium (channel source) |
1.2. Model Lineage
| Model | Time / Status | Positioning | Openness |
|---|---|---|---|
| 可图 Kolors | Released 2024-05 | Text-to-image + image-to-image, strong Chinese comprehension | Open source |
| Kolors 2.1 | Subsequent iteration | Photorealistic rendering, color science, portrait skin texture, in-image text | Open-source status not clearly found |
| Kling O1 Image | Currently the main model | Unified multimodal image engine: text-to-image + natural-language instruction editing + identity consistency preservation across up to 10 reference images | Closed source |
| IMAGE 3.0 Omni | Shown on membership page | Offered in parallel with O1 | Closed source |
Official and channel statements for Kling O1 Image: supports 18+ task types, 4K output, and up to 10 reference images. These parameters come from a third-party aggregator marketing page, not official.
1.3. Pricing System
1.3.1. International-Site Membership (official page, high confidence)
| Tier | Monthly | Original price | Credits/month | Annual |
|---|---|---|---|---|
| Basic | $0 | — | — | — |
| Standard | $6.99/month | $10 | 660 | $72/year |
| Pro | $25.99/month | $37 | 3,000 | $269/year |
| Premier | $64.99/month | $92 | 8,000 | $659/year |
| Ultra | $127.99/month | $180 | 26,000 | $1,296/year |
Official conversion basis: 660 Credits ≈ 3,300 images or 33 720p videos. By reverse inference, one image is on the order of about 0.2 Credit.
Credits top-ups: 3,500 Credits / $50; 7,500 Credits / $100; 16,000 Credits / $200.
1.3.2. China-Site Membership (secondary source, medium confidence)
| Tier | Next-month renewal price | Inspiration points | Notes |
|---|---|---|---|
| Gold | ¥58/month | 660 | First month from 40% off |
| Platinum | ¥234/month | 3,000 | — |
| Diamond | ¥586/month | 8,000 | — |
| Black Gold | ¥1,149/month | 26,000 | — |
Source: a research report screenshot from Guosheng Securities. The China site and the international site align in their Credits / inspiration-point tier design (660 / 3,000 / 8,000 / 26,000), but the pricing systems are independent. China-site prices [To be verified — refer to the real-time price in the App].
Basis note (listed in parallel across the group): Section 1.3.2 records the China-site membership prices as Gold ¥58 / Platinum ¥234 / Diamond ¥586 / Black Gold ¥1,149 (inspiration points 660 / 3,000 / 8,000 / 26,000); another China-site basis is recorded in 03-市场研究/04-AI-漫剧组/02-kling-drama.md: Gold ¥19 / Platinum ¥58 / Diamond ¥198 (inspiration points 660 / 2,000 / 8,000). The international-site tier ceiling also has two figures: the Ultra in this section 1.3.1 is $127.99/month, while the Ultra renewal price in 02-kling-drama.md is $159.99/month. Both China-site figures come from relayed public channels and have not been directly confirmed against the official pricing page; the discrepancy may correspond to different time periods or different package systems. This section does not adopt either one; please refer to the real-time price on the official page when citing.
2. Glossary
2.1. General AI Image Terms
| Term | English / Abbreviation | Definition | Kling's corresponding implementation |
|---|---|---|---|
| Text-to-image | Text-to-Image (T2I) | Generates an image solely from a text prompt | Supported, one of O1 Image's main capabilities |
| Image-to-image | Image-to-Image (I2I) | Generates a new image conditioned on one or more images | Supported, up to 10 reference images |
| Inpainting | Inpainting | Regenerates a specified region while leaving the rest unchanged | Superseded by natural-language instruction editing: describe the region, no mask needed |
| Outpainting | Outpainting | Continues generating beyond the canvas edges | No standalone outpainting product description found at the platform layer, [To be filled] |
| Controllable generation | ControlNet | Controls generation structure via visual signals such as Canny, Depth, Pose, Mask | No native ControlNet description; structural control is handled by reference images and instructions |
| Reference image | Reference Image | An image used as an input for identity, style, and structural constraints | Up to 10 images, processed through the Identity Map |
| Random seed | Seed | When fixed, reproduces the same image under the same parameters | Seed locking and snapshot endpoint documentation not made public at the platform layer, [To be filled] |
| Guidance strength | CFG | The strength with which the prompt constrains the generation result | No publicly exposed mechanism at the platform layer, [To be filled] |
| Low-rank adaptation | LoRA | A low-parameter fine-tuning module used to fix characters, styles, and garment assets | No user-side LoRA training or loading |
| Image prompt adapter | IP-Adapter | Injects image-encoder features into attention to achieve "image as prompt" | Not publicly used; the function is served by the Identity Map |
| Zero-shot identity injection | InstantID | Transfers identity from a single reference image without fine-tuning | Not publicly used; O1 uses multiple images (5–10) to build stable identity markers |
2.2. Kling-Specific Terms
| Term | English / Abbreviation | Definition |
|---|---|---|
| Inspiration points / Credits | Credits | Kling's unified metering unit, shared by images and video; official basis: 660 Credits ≈ 3,300 images or 33 720p videos |
| Element Creation | Element Creation | An asset capability that creates people, objects, etc. as reusable subjects; membership tiers cap the count at 30 / 50 / 150 / 500 |
| Identity Map | Identity Map | Extracts stable identity markers (bone structure, salient features, proportions) from 5–10 reference photos at different angles, lighting, and distances, and decouples them from transient attributes such as expression and outfit |
| Stable identity marker | Stable Identity Marker | Persistent features extracted by identity mapping, separated from transient attributes such as the subject's expression, lighting, and outfit |
| Transient attribute | Transient Attribute | Features that change with the scene, such as expression, outfit, lighting, and makeup, explicitly decoupled in identity mapping |
| Kling O1 Image | — | Unified multimodal image engine: text-to-image + natural-language instruction editing + identity consistency preservation across up to 10 reference images |
| Element Smart Completion | Element Smart Completion | Subject-completion capability of the O1 video series; free 3 times per day across tiers |
| Image quality enhancement | Image Quality Enhancement | Membership benefit, Pro and above |
| 可图 Kolors | Kolors | Open-source image model released 2024-05, text-to-image + image-to-image |
| Unlimited concurrent tasks | Unlimited Concurrent Tasks | Membership benefit: unlimited concurrent tasks + dedicated fast-generation channel |
2.3. General Terms for Virtual Try-On and Face Swap
Note: In version 1.5, Kling added the AI Model capability, and supports the reuse of product and person assets through "Element Creation", so it has practical touchpoints for virtual try-on; the platform does not offer a standalone face-swap product capability.
| Term | English / Abbreviation | Definition | Kling relevance |
|---|---|---|---|
| Virtual try-on | VTON (Virtual Try-On) | "Wears" the target garment onto a specified person's image and produces a visually plausible result | Can be supported by AI Model + Element Creation; no standalone VTON product line found |
| Garment warping | Garment Warping | Uses geometric transforms such as TPS (thin-plate spline) to align a flat garment to the body pose, then feeds it into generation | No public description |
| Cloth mask | Cloth Mask | Binary map of top, bottom, and outerwear regions derived from human parsing, used to limit the repaint area | Not applicable (natural-language instruction editing needs no mask) |
| Try-on diffusion | Try-on Diffusion | Uses a diffusion model to fuse garment and body end-to-end | No public description |
| Face swap | Face Swap | Replaces person A's face onto person B's facial position | No standalone face-swap capability |
| Face reenactment | Face Reenactment | Preserves identity while transferring expression, lip sync, and head pose | Related statements on the O1 video series |
| Identity preservation | Identity Preservation | The extent to which the generation result is still "that person" | Core capability: the Identity Map is designed precisely for this |
3. Feature Description
3.1. Generation and Editing Capabilities
| Capability | Supported | Description |
|---|---|---|
| Text-to-image | Supported | O1 Image's main capability |
| Image-to-image | Supported | Up to 10 reference images |
| Natural-language instruction editing | Supported | "Remove the background", "change the coat to burgundy", "add fog by the window", without masks or selections |
| Style transfer | Supported | — |
| Local editing | Supported | Describe the region in text, no painting needed |
| Image super-resolution / quality enhancement | Supported | Tiers Pro and above |
| AI Model | Supported | Added in version 1.5 |
| Element Creation | Supported | Tier limits of 30 / 50 / 150 / 500 |
| Video generation | Supported | Shares Credits with images; the O1 series includes Element Smart Completion |
3.2. Natural-Language Instruction Editing
Kling's editing paradigm is "use words instead of painting": the user describes "change the coat to burgundy", and the model locates the target region itself and completes the change. This forms a three-generation interaction difference versus Midjourney's Vary Region (requires first selecting a region) and Jimeng's intelligent canvas (requires first painting):
| Platform | Editing interaction | Requires mask | Requires re-uploading reference |
|---|---|---|---|
| Midjourney | Vary Region (region selection + parameter isolation) | Yes | Yes |
| Jimeng | Intelligent canvas (painting + layers) | Yes | No (kept in the canvas) |
| Kling O1 | Natural-language instruction | No | No (kept in the session) |
Engineering implication: the masking step is a frequent source of human error and workflow interruption. Removing the mask means a single generation can be programmatically and continuously corrected, a substantive step by the L3 orchestration layer toward "automation". However, Kling has not yet solidified this chain into an exportable workflow artifact.
3.3. Element Creation and AI Model
- Element Creation (Element): fixes people, products, and objects as reusable subjects; the tier determines the count cap (30 / 50 / 150 / 500).
- AI Model (added in version 1.5): combined with Element Creation, it can support the reuse of virtual models in e-commerce scenarios.
- Daily inspiration points: members receive daily inspiration points (specific value
[To be filled]).
3.4. Commercial Licensing and Concurrency
- Content generated on the free tier (Basic) is not permitted for commercial use — this is an explicit clause on the official membership page and is one of the rare free-tier commercial-use restrictions in this group that are explicitly written down by the vendor; pay special attention during selection.
- Standard and above include commercial-use rights (stated explicitly on the official membership page).
- Concurrency: membership tiers offer "unlimited concurrent tasks" + a dedicated fast-generation channel.
- Watermark: paid tiers remove the brand watermark.
4. Platform Architecture
图 4-1|可灵在 AI Harness 六层能力模型中的实现
数据来源:基于本文分析绘制的示意图。
4.1. Dual-Model Parallel
| Model | Focus |
|---|---|
| Kling O1 Image | Unified generation + editing + consistency; natural-language instruction editing; 10 reference images |
| Kolors 2.1 | Photorealistic rendering, color science, portrait skin texture, in-image text |
The two lines run in parallel, showing that Kuaishou has not pushed all capabilities onto a single model but instead divides them by "consistency / conversational editing" and "photorealistic rendering / text". This is conceptually similar to FLUX.2's "model selection as orchestration" (klein realtime / pro production / flex layout / max highest quality), but Kling has not elevated model selection into an orchestratable routing strategy — users can only switch manually.
4.2. Text Encoder Trade-offs
Kolors uses GLM as its text encoder, differing from the T5 commonly used by Imagen / SD3. Using a natively Chinese large language model as the encoder is the underlying reason for its strengthened Chinese-English comprehension.
This fact comes from a third-party technical write-up; the original paper has not been located.
4.3. Open Format
- Web (klingai.com) + App.
- Enterprise-side API (Kuaishou Open Platform).
- Integrated into third-party platforms (e.g., vivago.ai) and aggregator platforms; the aggregator ecosystem contributes meaningfully to its distribution.
5. Harness Design
5.1. Six-Layer Capability Overview
| Layer | Name | Kling's implementation | Maturity | Evidence strength |
|---|---|---|---|---|
| L1 | Context engineering | Up to 10 reference images → Identity Map, decoupling stable identity markers from transient attributes before anchoring | Strong | Medium (recommend rechecking against official documentation) |
| L2 | Tools & execution | Generation / editing / style transfer / super-resolution / video extension / Element Creation, invoked via natural language within the same session | Medium | Medium-high |
| L3 | Orchestration & control | "Generation as conversation": generation and editing are not distinguished; multi-turn conversational iterative loop with no re-uploading, no masking, no re-rolling | Medium-strong | Medium |
| L4 | Memory & state | Element Creation is asset persistence, with the tier determining the count cap; members receive daily inspiration points | Medium-strong | High (official page) |
| L5 | Evaluation & observation | Kolors ranked 2nd globally in subjective overall score and 1st in subjective image quality on BAAI FlagEval (a leaderboard point-in-time value); no public official Eval Set | Weak | Medium |
| L6 | Governance & security | Free tier not for commercial use; paid tiers remove watermark; risks around real-person footage and identity consistency are governed by the membership agreement | Medium | Medium |
5.2. L1 Context Engineering Layer
Kling's L1 design is the one in this group with the greatest "context layering" awareness:
- Input layer: up to 10 reference images (channel source; recommended to recheck against official documentation). Officially it is recommended to extract features from 5–10 reference photos at different angles, different lighting, and different distances.
- Layering mechanism (Identity Map): divides features into two categories —
- Stable identity markers: persistent attributes such as bone structure, salient features, and proportions;
- Transient attributes: attributes that change with the scene, such as expression, outfit, lighting, and makeup.
- Anchoring logic: only the former is anchored during generation, while the latter is allowed to vary freely.
This is exactly the last item — priority ordering — in the L1 definition of "retrieval, compression, caching, priority ordering" from the parameter card. Kling turns "what must stay consistent, what may change" into an explicit product semantic, rather than letting users tune weights themselves (contrast Midjourney's single---ow numeric knob).
Areas to strengthen: no structured prompt contract (contrast FLUX.2's JSON fields), no external knowledge retrieval (contrast Seedream 5.0's web retrieval and Nano Banana's Search Grounding).
5.3. L2 Tools & Execution Layer
- Toolset: generation, editing, style transfer, super-resolution, video extension, Element Creation.
- Invocation form: uniformly invoked via natural language within the same session; tool boundaries are invisible to the user.
- No open tool registration: no Function Calling support, no MCP Server (contrast Runway, which released an MCP Server in 2025-06).
- Cost basis: Credits are a unified metering unit shared by images and video. The official conversion of 660 Credits ≈ 3,300 images gives decent cost estimability, but there is no pre-generation cost-estimation endpoint (contrast Leonardo.ai's Pricing Calculator).
5.4. L3 Orchestration & Control Layer
"Generation as conversation" is Kling's core contribution at L3:
| Traditional workflow | Kling O1 workflow |
|---|---|
| Upload reference → generate → unsatisfied → re-upload → adjust parameters → re-roll | Upload reference → generate → critique in text → targeted fix → continue critiquing |
Three steps are removed: re-uploading, masking/painting, and whole-image re-rolling. This brings two direct benefits:
- Context is not lost: the reference images and the generated result both remain in the session, avoiding consistency breaks caused by "starting a new round".
- Iteration can converge: each change touches only one spot, converging to the target over multiple rounds rather than rolling the dice again each round.
Still missing: no DAG, no sub-agents, no interruption recovery, no workflow export or version control. Compared with ComfyUI turning workflows into version-controllable JSON graphs and Meitu Design Studio using Agent Teams for multi-agent division of labor, Kling's L3 remains at "human-driven multi-turn conversation" and has not yet entered "system-driven automatic orchestration".
5.5. L4 Memory & State Layer
Element Creation is the most straightforward L4 productization design in this group:
| Platform | L4 first-class asset | Pricing / limit method |
|---|---|---|
| Kling | Element Creation | Tier limits of 30 / 50 / 150 / 500 |
| Miaoya Camera | Digital avatar | Per-use creation (¥9.9 / ¥29.9), locked inside a single product, not transferable |
| Runway | Brand Kits / Voice Clones | Up to 3 each on the Max tier |
| Nano Banana | No dedicated abstraction | Relies on the Google ecosystem (Drive / NotebookLM) |
| Meitu Design Studio | Creative assets (officially stated as reusable) | Limit basis not disclosed |
Kling writes "asset count" directly into the membership entitlements table, meaning assets are product value. This makes commercial sense, but it also carries the same risk as Miaoya: assets are locked inside the platform. The difference is that Kling's asset form (person / product / object subjects) is more general than Miaoya's "digital avatar", and the platform remains in continuous operation.
Gaps: no Checkpoint, no artifact version management, no model-version locking endpoint, no documentation of an asset export mechanism.
5.6. L5 Evaluation & Observation Layer
- Kolors ranked 2nd globally in subjective overall score and 1st in subjective image quality on BAAI FlagEval. This result is a point-in-time leaderboard value; it must be timestamped and should not be treated as a continuously valid conclusion.
- No public official Eval Set, Golden Dataset, or regression set.
- Observation basis: Credits consumption. No pre-generation cost estimation, no A/B mechanism.
5.7. L6 Governance & Security Layer
- Tiered commercial-use rights: the free tier is not for commercial use (stated explicitly on the official membership page), and Standard and above include commercial-use rights. This is one of the clearest tiered commercial-use designs in this group.
- Watermark: paid tiers remove the brand watermark.
- Portrait and identity risks: the portrait-rights risks posed by real-person footage and identity consistency (Identity Map explicitly uses real people as reference subjects) are governed by the platform's membership agreement; no structured identity authorization or guardrail mechanism was found (contrast Nano Banana's public-figure portrait restrictions and SynthID's mandatory watermarking).
- Compliance with the Labeling Measures: no official public description of explicit / implicit labeling implementation was found.
[To be filled]
5.8. Maturity Assessment
Kling is a "medium-strong model + medium-strong Harness" form, whose greatest features are the most straightforward L4 assetization (Element Creation limited-quantity pricing) and L3 generation-as-conversation (removing masking and re-rolling). It is closer to a deliverable form than Midjourney, but still has a clear gap from Meitu Design Studio's Agent Teams and ComfyUI's version-controllable workflows. If Miaoya is taken as a failed sample of "having assets but not transferable", then Kling is the control group of "having assets and an active platform" — whether assets are transferable and whether the platform stays operational together determine the real value of L4 investment.
6. Practical Cases
6.1. Official Customer-Case Search Results
This search did not find any brand customer case officially published by Kling with quantified performance data. Per this group's uniform basis, it is marked as [No results], without substituting vague phrasing such as "widely used in the industry" and without fabricating cases.
6.2. Channel-Partner Workflow Example (not cited as a case)
Workflow example provided by a third-party channel partner (vivago.ai): an AI virtual influencer — 5 reference images → 120 scenes/month, no photographer or model scheduling, with commercial-use rights included in the Plus plan.
This content is channel marketing material, not an official Kling release, with no third-party verification; this report does not cite it as a case, and treats it only as a "typical usage" reference. The quantified figure within it (120 scenes/month) should not be re-quoted as performance data.
6.3. Verifiable Usage and Allocation Data
Verifiable factual material (not performance data):
- Official conversion basis: 660 Credits ≈ 3,300 images or 33 720p videos (official membership page, high confidence) — usable for per-unit cost estimation.
- Element Creation tier limits: 30 / 50 / 150 / 500 (official page, high confidence) — usable for capacity planning.
- Free tier not for commercial use (official page, high confidence) — usable for compliance determination.
- Kolors open source (2024-05) — usable for local deployment and secondary-development evaluation.
7. Summary
7.1. Strengths
- Identity Map's context-layering design: explicitly distinguishes stable identity markers from transient attributes, the clearest semantic design at the L1 layer in this group.
- Generation as conversation: removes masking and re-rolling; iteration can converge and human error points are notably reduced.
- Element Creation: assets become first-class and are limited-priced, making capacity planning predictable.
- Clear tiered commercial-use rights: free tier not for commercial use is written down explicitly by the vendor, making compliance determination simple.
- Officially given Credits conversion basis: 660 Credits ≈ 3,300 images, so per-unit cost is estimable.
- Kolors open source: leaves a path for local deployment and secondary development.
7.2. Limitations and Applicability Boundaries
| Limitation | Impact |
|---|---|
| No workflow artifacts or version management | Session history from generation-as-conversation cannot be exported or regressed |
| No model-version locking endpoint | Output drifts for the same reference and instruction after a model upgrade |
| L5 has only point-in-time leaderboards | Quality improvements cannot be quantified |
| No open tool registration / MCP | Cannot be embedded into external agent toolchains |
| Element assets not transferable | Risk of the same kind as Miaoya; depends on the platform's continued operation |
| Free tier not for commercial use | Trial results cannot be used commercially directly; must upgrade first |
| 10 reference images / 4K / 18+ tasks are channel sources | Parameter promises cannot serve as contractual basis |
Applicability boundaries: suited to content production that needs "reusing the same person / the same product across multiple scenarios" (virtual influencers, e-commerce models, IP character extension); not suited to enterprise-grade pipelines requiring programmable orchestration, reproducible regression, and cross-system auditing.
7.3. Selection Recommendations
- Do capacity planning before choosing a tier: back out the tier from two metrics — "monthly subject count + monthly image count". Subject count is a hard constraint (30 / 50 / 150 / 500), and image count can be topped up with single Credits purchases.
- Must upgrade before commercial use: free-tier output is not for commercial use and must not be used in any external-facing material.
- Build your own reproducible record: record the four-tuple of "reference-set hash + instruction sequence + model version + Element ID" as regression and evidence material.
- Parameter promises must defer to the official source: parameters such as 10 reference images, 4K, and 18+ task types come from channel marketing pages; confirm with Kuaishou officially before signing.
- For reference images of real people: must obtain written authorization from the rights holder and keep the authorization chain; platform agreements cannot replace rights-holder-to-rights-holder authorization.
7.4. Compliance Notes
- Articles 4 and 5 of the Measures for the Labeling of AI-Generated Synthetic Content (Office of the Central Cyberspace Affairs Commission [2025] No. 2): when service providers offer download, copy, and export functions, they shall ensure explicit labeling and shall add implicit labeling to the file metadata. No public implementation description was found on Kling's side; when used by enterprises, you should build the labeling-injection step yourself.
- Article 1019 of the Civil Code of the People's Republic of China: it is prohibited to infringe on another person's right to portrait by means such as forgery using information-technology methods. Identity Map builds stable identity markers from real-person reference photos, a highly sensitive form of identity reuse.
- A judgment effective in 2026-03 at the Beijing Internet Court established "identifiability + burden-of-proof transfer": the result need not be identical to the original portrait; as long as the public can identify it, it constitutes use of a specific natural person's portrait; a defendant claiming "coincidental resemblance" must reproduce the creation process. When anchoring identity with multiple reference images, once a generated result closely resembles a natural person, the burden of proof lies with the user.
Information Gap Statement
- Original source of Kolors using the GLM text encoder: only mentioned in third-party technical write-ups; the paper has not been located.
- Kling O1 Image's "18+ task types / 4K output / 10 reference images": from a third-party aggregator marketing page, not official.
- China-site membership prices (¥58 / ¥234 / ¥586 / ¥1,149): from a broker research report screenshot, not the official page.
- Daily inspiration-point amount: not found.
[To be filled] - Official customer cases and quantified performance data: not found.
[No results] - Details of explicit and implicit labeling implementation under China's Labeling Measures: no official public description found.
[To be filled] - Seed locking, CFG exposure, model-version locking endpoint: no official description found.
[To be filled] - Whether Kolors 2.1 is open source: no clear statement found.
- Specific date of the FlagEval leaderboard: the leaderboard release date was not found; when citing, it must be labeled as a point-in-time value.
- Pricing and concurrency specifications for the enterprise-side API: no official API pricing page found.
[To be filled]
8. References
- Kling AI Membership Plans (official page, including tiers, Credits conversion, and commercial-use terms) — Kuaishou, 2026. https://app.klingai.com/global/membership/membership-plan
- Kling AI Official Website — Kuaishou. https://klingai.com/
- Measures for the Labeling of AI-Generated Synthetic Content (Office of the Central Cyberspace Affairs Commission [2025] No. 2) — Cyberspace Administration of China, Ministry of Industry and Information Technology, Ministry of Public Security, National Radio and Television Administration; issued 2025-03-14, effective 2025-09-01. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- Civil Code of the People's Republic of China, Articles 1018 and 1019 — National People's Congress, 2020. (Articles cited in the body text; no official link)
- Economic Information Daily · "Technology Is Not a Shield against Infringement — How the Court Found AI 'Face Theft'" — Xinhua News Agency / Economic Information Daily, 2026-04-17. http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
- InstantID official project page (zero-shot identity injection, as the comparison baseline for Identity Map) — InstantX Team / Xiaohongshu / Peking University. https://instantid.github.io/
- DeepWiki · SDXL_EcomID_ComfyUI 5.4 Comparison (comparison of EcomID / InstantID / PuLID control mechanisms) — Alimama. https://deepwiki.com/alimama-creative/SDXL_EcomID_ComfyUI/5.4-comparison-with-other-methods
- ComfyUI official workflow "Virtual Character Try-On - All-in-One" (comparison baseline for version-controllable workflows) — Comfy Org. https://comfy.org/zh/workflows/templates_rob_fashion_shoot_vton-4in1.app/
- CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models — arXiv 2407.15886 (virtual try-on technical approach and evaluation metrics). https://arxiv.org/pdf/2407.15886
- Black Forest Labs · FLUX.2 Overview (comparison baseline for model routing and fixed-snapshot endpoints) — BFL, 2026. https://docs.bfl.ai/flux_2/flux2_overview
- Leonardo.Ai Pricing (official page; comparison baseline for the Pricing Calculator and capacity constraints) — Leonardo Interactive. https://www.leonardo.ai/pricing