AI 图像组组概述与横向对比


1. 组概述

1.1. 研究问题与统一框架

本组对 AI 图像领域 16 个平台(含 1 个开源生态)进行了市场研究,统一使用项目参数卡定义的 AI Harness 六层能力模型作为分析框架:

名称职责本组的典型观察对象
L1上下文工程层决定模型"看到什么"参考图上限、长提示词、结构化提示契约
L2工具与执行层决定模型"能做什么"生成/编辑/换装/换脸工具集、API、节点
L3编排与控制层决定"按什么顺序做"工作流、Agent Teams、多轮会话迭代
L4记忆与状态层决定"记住什么"资产库、LoRA、品牌包、可复现工件
L5评估与观测层决定"做得好不好"Eval Set、一致性 Gate、成本观测
L6治理与安全层决定"不能做什么"水印、标识、审核、许可、审计

本组的核心研究结论可以用一句话概括:AI 图像行业的竞争已经从"模型好不好看"转入"生成过程是否可预期"——参考图槽位、可复现端点、工作流版本控制、内容标识这些 Harness 能力,正在成为平台间真正的分水岭。

1.2. 数据时点与证据分级

  • 所有平台数据为 2026-09-11 至 2026-09-12 检索时点快照。AI 图像领域迭代极快(检索窗口内即出现 Midjourney V8.2 上线、Qwen-Image 3.0 闭源、可灵图像线进入 O1 / 3.0 Omni),引用前应对版本号、定价、榜单排名做实时复核。
  • 证据分级沿用检索报告口径:官方来源(官网、官方文档、财报、法院判决)> 权威媒体 > 第三方评测/社区(低—中置信,正文中均已标注)。
  • 各篇文档末尾的「信息缺口声明」列明未检索到或置信不足的数据;无公开量化案例的平台一律如实写"未检索到",不以模糊表述替代。

2. AI 图像市场全景

按产品命题与用户群划分,本组 16 个样本覆盖八条细分赛道:

赛道核心命题组内样本
通用文生图 / 图生图(创意)美学质量与风格控制Midjourney、Leonardo.ai、FLUX.2
指令式生成与编辑(多模态)"说一句话改一张图"的多轮迭代Gemini / Nano Banana、即梦 AI、Seedream
视频驱动的图像线图像与视频生成同栈可灵 AI 图像(Kolors / O1)、Runway
人像写真单点爆款身份生成妙鸭相机(已收缩运营,反面样本)
电商换装 / 商拍SKU 一致性 + 成本替代商拍美图设计室、WeShop、PicCopilot
换脸 / 娱乐社交身份迁移的娱乐化开源生态(ReActor / InstantID 线)+ 各平台参考图能力
设计中台 / Agent 化设计简报进、成套资产出LOVART / 星流
开源自建权重与编排自主权Qwen-Image(1.0/2.0)、FLUX klein、开源换脸换装生态、LiblibAI(云化聚合)

三条结构性趋势贯穿全部赛道:

  1. 参考图成为主战场:从 Midjourney 的 --oref 到 Nano Banana Pro 的 14 张参考 + 5 人物一致性,"把品牌、人物、商品装进上下文"的能力决定商业价值上限;
  2. 编排层快速上移:美图 Agent Teams、LOVART 设计 Agent、ComfyUI 工作流三种形态并行,共同把"多步生产链"做成产品;
  3. 开放与闭源的张力公开化:Qwen-Image 3.0 收回权重(2026-07)与 FLUX.2 klein 4B 坚持 Apache 2.0 形成同期对照,开放路线的可持续性成为选型必答题。

3. 16 平台横向对比矩阵

平台开发商定位开放形态定价(时点价,详见各篇)L1L2L3L4L5L6适用边界
MidjourneyMidjourney Inc.(美)美学驱动的通用文生图Web + Discord$10/$30/$60/$120 月品牌创意、概念视觉
即梦 AI字节跳动(剪映团队)C 端一站式创作Web/App/小程序 + 火山 APIC 端 ;API 0.2~0.22 元/图中(曾被查处)中文创作者、短视频生态
可灵 AI 图像快手图像 + 视频统一创作Web/App + API国际站 $6.99~$127.99 月中强中强中高主体一致性创作、视频前置
Nano BananaGoogle DeepMind多模态指令式生成编辑Gemini 全家桶 + API/VertexNB2 1K $0.067;Pro $0.134最强企业生产、合规敏感场景
美图设计室美图公司电商设计工作流 + AgentWeb/桌面/App + 海外版¥30/¥88 起 最强电商卖家、批量商拍
妙鸭相机未序网络 / 阿里大文娱人像写真单点产品小程序/App¥9.9 / ¥29.9(2023 价)历史样本:单点爆款难留存
Leonardo.aiLeonardo Interactive(Canva 旗下)多模型聚合创作者平台Web + PAYG API$12/$30/$60;团队 $72/$144中强创作者规模化生产、API 集成
RunwayRunway AI(美)视频为核心的多模态平台Web + API + MCP$12/$28/$76 月;$0.01/credit中强影视级图像视频生产
FLUX.2Black Forest Labs(德)开源 + API 双轨图像基座API + HF 权重 + ComfyUIklein 4B $0.014 起(Apache 2.0);pro $0.03/MP最强排版文字、自建管线、可复现生产
Seedream / SeedEdit字节 Seed 团队统一生成编辑模型字节生态内嵌 + 火山 API0.2~0.22 元/张生态内生产、虚拟试穿多图参考
通义万相 / Qwen-Image阿里通义团队开源权重 + 云 API 双线HF 权重 + 百炼 API + Qwen Chat0.2 元/张(生成)、0.14 元/张(编辑)中→弱(3.0 断裂)中强图内有字场景、私有化部署
开源换脸换装生态学术界 + 社区换装(VTON)+ 换脸(Identity)方案集ComfyUI 节点 + 权重免费(许可分层)最强最强最弱自建管线、成本敏感批量场景
LiblibAI 2.0北京奇点星宇开源模型聚合 + 云端创作社区Web/App + 云 ComfyUI¥35/¥70/¥199 月 强(社区资产库)弱—中(曾遭央视点名)无 GPU 的开源生态用户
LOVART / 星流LiblibAI 体系设计垂类 AgentWeb最强最强(Brand Kit)品牌全案设计、成套资产交付
WeShop蘑菇街团队孵化服饰换装商拍Web(国际站 + 国内站)中强中强强(SKU Gate)服饰电商换装、固定模特
PicCopilot阿里国际跨境电商营销素材Web框架强约束框架高风险跨境素材广度、多语言本地化

注:L1~L6 为六层 Harness 成熟度的定性评级(弱 / 中 / 中强 / 强 / 最强),依据各篇正文的逐层分析汇总;LiblibAI、LOVART、WeShop 三行依据组内对应文档的六层小结;PicCopilot 行证据最弱(其官方能力未获专题检索支持,仅框架性判断成立)。

4. 六层 Harness 成熟度总评

图 4-1|AI 图像平台的六层 Harness 能力模型

AI 图像平台的六层 Harness 能力模型 信息截止 2026-09-12 · 示意:基于本文分析绘制 L1 · 上下文工程层|决定模型“看到什么” 参考图上限、长提示词、结构化提示契约 L2 · 工具与执行层|决定模型“能做什么” 生成 / 编辑 / 换装 / 换脸工具集、API、节点 L3 · 编排与控制层|决定“按什么顺序做” 工作流、Agent Teams、多轮会话迭代 L4 · 记忆与状态层|决定“记住什么” 资产库、LoRA、品牌包、可复现工件 全行业短板 L5 · 评估与观测层|决定“做得好不好” Eval Set、一致性 Gate、成本观测 全行业短板 L6 · 治理与安全层|决定“不能做什么” 水印、标识、审核、许可、审计 结构解读:竞争已从“模型好不好看”转入“生成过程是否可预期”,六层成熟度构成统一分析框架。 全行业普遍短板在 L4(资产持久化)与 L5(评估标准)——正是平台间真正的分水岭。

数据来源:基于本文分析绘制的示意图。

4.1. 全行业的两个短板:L4 与 L5

把 16 个样本横向放在一起,最清晰的行业结论是:全行业普遍弱在 L4(可复现性与资产持久化)与 L5(效果评估缺乏标准)

  • L4 的普遍缺口:多数平台没有把"用户的生成历史、参数、参考图、模型版本"作为可迁移的一等资产。妙鸭相机是极端案例——数字分身锁定在单一产品内不可迁移,被认为是其 2025-09 团队解散的工程根因之一。做得最好的三家(美图创作资产复用、Runway Brand Kits、ComfyUI 工作流 + LoRA 全部归用户)恰好也是商业化最健康或生态最活跃的样本。
  • L5 的普遍缺口:几乎没有平台公开官方 Eval Set 或回归集。行业评估主要依赖三种代偿:厂商自评榜单(可灵 FlagEval、Qwen-Image 9 项基准第一——且 Qwen 自评榜单中出现对己不利的第五名,是罕见例外)、业务指标反推(美图的算力点消费与 ARPPU 倍数)、以及人眼主观评测(Midjourney)。电商垂直路线是唯一把评估做成硬 Gate 的场景(SKU 一致性三层抽检),这恰说明:没有外部强约束时,厂商没有动力建评估层。
  • 可复现性正在补 L4 的课:FLUX.2 的固定快照端点、ComfyUI 的 JSON 工作流版本控制、Midjourney 的 seed 参数体系,共同把"同参数复现"从奢望变成工程开关——这是本组观察到的最有价值的行业趋同。

4.2. 六层中的极端样本

最强样本最弱样本对照含义
L1FLUX.2(JSON 结构化提示词 + hex 品牌色 + Grounding)妙鸭相机(模板即全部)上下文结构化程度决定可控性上限
L2ComfyUI 开源生态 / LiblibAI(节点即工具)Midjourney / 妙鸭(工具菜单极短)工具开放度决定生态密度
L3美图 Agent Teams / LOVART / ComfyUI 工作流妙鸭(原子操作)编排是把生成变生产的关键
L4LOVART Brand Kit / Runway Brand Kits / ComfyUI(资产归用户)妙鸭(资产不可迁移)资产可迁移性决定留存
L5美图(业务指标闭环) / 电商路线(SKU 硬 Gate)多数平台(无 Eval Set)无标准则无法回归
L6Nano Banana(SynthID + C2PA + 公众人物限制)开源生态(无强制标识)治理是商业化的前置条件

4.3. 参考图能力对比(L1 核心指标)

平台参考图上限人物一致性机制名称
Midjourney可堆叠Omni Reference(--oref + --ow)/ Style Reference
即梦 / Seedream十余张(第三方 API 口径 1~10)多图参考 + 视觉信号(Canny/Depth/Mask)
可灵 O1 Image10 张跨 100+ 次生成(渠道口径)Identity Map(稳定身份标记 vs 瞬时属性解耦)
Nano Banana Pro14 张5 人多参考融合 + Search Grounding
Runway Gen-4 Image3 张单参考即可锁定Reference Images(subject/scene/style 分工)
FLUX.28(API)/ 10(Playground);klein 48 个一致角色Multi-Reference + Structured Prompting
开源(ComfyUI)由工作流决定由 LoRA / InstantID / PuLID 决定节点编排

4.4. 定价速查

平台付费形态关键单价置信度
Midjourney订阅$10 / $30 / $60 / $120 每月
即梦(火山引擎)按量0.2 元/图(3.0 系)、0.22 元/张(4.0/4.6)官方,极高
可灵订阅 + Credits$6.99~$127.99 月;660 Credits ≈ 3,300 张图官方,高
Nano Banana订阅 + APINB2 1K $0.067;NB Pro $0.134中高
美图设计室订阅 + 算力点¥30 / ¥88 起(第三方口径)低,
妙鸭相机单次付费¥9.9 / ¥29.9(2023 时点)
Leonardo.ai订阅 + PAYG API$12 / $30 / $60;API $5 起始、10 并发官方,高
Runway订阅 + API$12 / $28 / $76;$0.01/credit官方,高
FLUX.2API + 开源权重klein 4B $0.014 起(Apache 2.0);pro $0.03/MP官方,极高
Seedream(火山引擎)按量0.2 / 0.22 元官方,极高
通义万相 / Qwen-Image按量 + 开源0.2 元/张(生成);0.14 元/张(编辑)官方,极高
开源生态免费(许可分层)Apache 2.0 / NCL / 非商用论文 + HF,高
LiblibAI订阅 + 算力点¥35 / ¥70 / ¥199 月低—中,
LOVART / 星流订阅
WeShop免费 + 订阅
PicCopilot

5. 选型决策树

按需求场景给出一棵可直接执行的决策树(各分支终点的平台详见对应篇目):

你的核心场景是什么?
│
├── 电商换装 / 商拍
│   ├── 服饰类目、要固定模特与换装深度 → WeShop(15 篇)
│   ├── 跨境营销素材广度(多语言、模板、视频本地化)→ PicCopilot(16 篇)
│   ├── 批量设计工作流 + Agent 化 + 数据可核验 → 美图设计室(05 篇)
│   └── 成本敏感、有 GPU 运维能力、可自建合规 → 开源生态 CatVTON / Re-CatVTON(12 篇)
│
├── 品牌创意 / 营销视觉
│   ├── 美学优先、可接受弱编排 → Midjourney(01 篇)
│   ├── 需要联网事实与最强治理 → Nano Banana / Gemini(04 篇)
│   └── 中文语境、与短视频生态协同 → 即梦(02 篇)/ Seedream(10 篇)
│
├── 人像写真
│   ├── C 端单次写真 → 妙鸭相机(06 篇,注意其运营收缩状态)
│   ├── 需要主体/身份长期一致 → 可灵 O1 Identity Map(03 篇)/ Nano Banana 多参考(04 篇)
│   └── 有技术团队、要零样本身份注入 → 开源 InstantID / PuLID / EcomID(12 篇)
│
├── 换脸娱乐 / 影视后期
│   ├── 低分辨率快速迭代、自带 NSFW 检测 → ReActor(12 篇)
│   └── 高分辨率身份保持 → PuLID / EcomID(12 篇);视频侧 → Runway(08 篇)
│
├── 开源自建 / 私有化
│   ├── 图内有字(中文排版、信息图)→ Qwen-Image 1.0/2.0 权重(11 篇)
│   ├── 排版文字 + 可复现生产 → FLUX.2 flex / 固定快照端点(09 篇)
│   ├── 换装批量管线 → CatVTON(2.26GB 显存)(12 篇)
│   └── 无 GPU、要社区模型与工作流 → LiblibAI 云端(13 篇)
│
└── 设计中台 / 成套资产交付
    ├── 品牌全案(Logo 到短视频一揽子)→ LOVART(14 篇;注意星流国内版 2026-10-10 停服)
    └── 生成 API 深度集成进自家产品 → Leonardo PAYG(07 篇)/ 火山引擎 / 百炼 API

三条通用的选型校验项(无论选哪个分支):① 参考图上限与人物一致性机制是否满足业务资产规模;② 是否提供可复现手段(seed / 固定端点 / 工作流文件);③ L5(一致性评估)与 L6(标识合规)由平台承担还是由你承担——这是商业平台与开源生态的真正差价所在。

6. 合规专章

AI 图像的合规约束在中国已经完成"立法 → 监管执法 → 司法裁判"的闭环,全部组内文档共用本章基线。

6.1. 监管框架:标识办法与深度合成规定

  • 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号,国家网信办、工信部、公安部、广电总局联合发布,2025-09-01 施行):
  • 第四条:提供生成合成内容下载、复制、导出功能时,应当确保文件中含有满足要求的显式标识(图片为适当位置添加显著的提示标识);
  • 第五条:应当在文件元数据中添加隐式标识(生成合成属性信息、服务提供者名称或编码、内容编号等),鼓励添加数字水印;
  • 第六条:传播平台应当核验元数据隐式标识并分三档处理(已标识 / 用户声明 / 检测到痕迹),并追加传播要素信息;
  • 第十条(红线):任何组织和个人不得恶意删除、篡改、伪造、隐匿标识,不得为他人实施上述行为提供工具或服务。
  • 配套强制性国家标准《网络安全技术 人工智能生成合成内容标识方法》同步发布。
  • 《互联网信息服务深度合成管理规定》:第十六条(隐式标识)、第十七条(人脸替换、人脸生成等深度合成服务,可能造成公众混淆误认的,应当进行显著标识)。
  • 《生成式人工智能服务管理暂行办法》第十二条:要求对生成内容进行标识。
  • 组内合规实现现状:除 Nano Banana(SynthID 不可见水印 + C2PA + 可见水印分级)外,其余平台均未见公开的标识实现说明;开源生态(12 篇)则完全需要使用方自建标识组件。这是选型时的必查项。

6.2. 民事边界:民法典肖像权条款

  • 第一千零一十八条:肖像是"在一定载体上所反映的特定自然人可以被识别的外部形象"——不要求完全一致,"可以被识别"即受保护;
  • 第一千零一十九条:任何组织或者个人不得以丑化、污损,或者利用信息技术手段伪造等方式侵害他人的肖像权;未经同意不得制作、使用、公开肖像权人的肖像——AI 换脸直接落入"利用信息技术手段伪造"
  • 第一千零二十条:合理使用情形(个人学习、新闻报道、公务履职、公共环境等)不包含商业性换脸
  • 第一千零二十三条:自然人声音参照肖像权规则保护(换声 / 语音合成同受约束)。

6.3. 司法落地:北京互联网法院 2026-03 判决

某知名演员诉短剧制作方与播出方 AI 换脸肖像权侵权案(2026-03 生效),四项裁判要点已成为行业合规基线:

  1. "可识别性"是核心判定标准:AI 换脸形象与原肖像无需完全一致,只要面部轮廓、五官特征高度相似、社会一般公众能够识别,即构成使用特定自然人肖像;
  2. 举证责任转移:被告主张"AI 偶然撞脸"的,须复现创作过程;无法复现即承担举证不能的不利后果——对使用开源工作流(参数多、过程难复现)的企业尤为不利,应主动留存生成日志与参数快照
  3. 著作权授权不能吸收肖像权:获得信息网络传播权授权不等于获得肖像使用授权,播出方未尽合理审查义务同样构成侵权;
  4. "技术中立"不是免责事由,争议片段"时长极短"亦不构成抗辩。

6.4. 监管执法:即梦 AI 被查处与央视点名

  • 即梦 AI(2026-04-28):因未有效落实人工智能生成合成内容标识规定要求被网信部门依法查处(百度百科词条口径,建议以官方通报复核)。这是本组调研时点内唯一的头部平台被查处案例,证明《标识办法》的执法已经触及一线大厂产品。
  • LiblibAI(2026-04):被央视点名多款 AI 软件存在严重监管漏洞,"哩布哩布 AI"被指可用隐晦提示词生成不当内容且仍可在应用商店正常下载(建议以监管通报复核)。
  • 两个案例共同说明:L6 治理层不再是可以后补的工程项,而是平台存续的前置条件。

6.5. 刑事风险样本

  • 江苏张家港王某 AI 换脸诽谤案:诽谤罪,有期徒刑一年三个月;
  • 山东宁阳孙某 AI 换脸结合语音合成网恋诈骗案:诈骗罪,有期徒刑三年缓刑四年。

6.6. 组内 16 平台合规要点对照

平台涉及的图像合规风险面公开的标识 / 水印实现组内文档的合规结论
Midjourney换脸 / 真人素材治理以社区准则为主未检索到官方说明显式+隐式标识实现
即梦 AI生成合成内容标识(曾被查处);真人素材限制未检索到官方说明2026-04-28 被网信部门查处(建议复核官方通报)
可灵 AI 图像免费档商用限制;真人身份一致性未检索到官方说明以会员协议约束肖像风险
Nano Banana公众人物肖像限制SynthID 强制不可见水印 + C2PA + 可见水印分级全组唯一公开、完整的标识体系
美图设计室换装涉真人形象EXIF 元数据 AI 声明(第三方口径,需实测复核)国内工具中合规较规范(待复核)
妙鸭相机生物特征(人脸)信息收集与用户协议未检索到官方说明初期用户协议争议 → 道歉修订(重要合规范例)
Leonardo.ai免费档 IP 归属;训练数据使用未检索到官方说明输入/输出均不用于训练;权利让渡清晰
Runway训练数据版权诉讼水印策略按档位治理以政策为主,诉讼在身
FLUX.2底座许可分层Apache 2.0 / NCL / 闭源 API 许可边界最清晰
Seedream生成合成内容标识(载体即梦曾被查处)未检索到官方说明模型侧无公开标识机制说明
通义万相 / Qwen-Image开源部署路径的标识义务落在部署者未提供标识组件去水印能力使用需自证合法性
开源换脸换装生态肖像权、深度合成标识、刑事滥用风险无强制标识(ReActor 仅 NSFW 检测)护栏需使用方自建(全组最大治理落差)
LiblibAI社区内容监管漏洞(曾遭央视点名)区块链创作溯源存证(官方口径)平台治理是软肋
LOVART / 星流生成内容标识与品牌资产版权星流停服带来资产迁移合规问题
WeShop虚假宣传(图文不符)、肖像权平台侧实现 [待填写]商家须自建"结构—属性—合规"三层抽检
PicCopilot虚假宣传(跨境双重法规)一致性与标识均需商家逐项确认

6.7. 合规自查清单

任何使用 AI 图像能力对外交付内容的团队,上线前应逐项核对:

  1. 生成物是否已添加显式标识(图内显著提示)?
  2. 文件元数据是否含隐式标识(属性、服务方编码、内容编号)?平台不提供时是否已自建注入?
  3. 涉及真人形象的:是否取得肖像授权?生成形象是否满足"可识别性"排查(宁严勿宽)?
  4. 换脸 / 人脸生成用途是否按《深度合成管理规定》第十七条做显著标识?
  5. 是否留有可复现创作过程的完整记录(应对举证责任转移)?
  6. 电商场景:生成图与实物 SKU 是否通过"结构—属性—合规"三层一致性抽检(防虚假宣传)?
  7. 是否存在去标识、去水印等触碰《标识办法》第十条红线的能力滥用?

7. 文档导航索引

本组共 17 篇文档(16 篇平台研究 + 1 篇组概述):

序号文件主题一句话定位
README.md组概述与横向对比本篇:16 平台对比矩阵、成熟度总评、决策树、合规专章
0101-midjourney.mdMidjourney美学驱动的通用文生图标杆,"强模型 + 弱 Harness"形态
0202-jimeng.md即梦 AI(字节)中文一站式创作平台,曾被查处的合规反面样本
0303-kling-image.md可灵 AI 图像(快手)Kolors + O1 双线,Identity Map 主体一致性
0404-nano-banana.mdGemini / Nano Banana(Google)多模态指令式生成编辑,L6 治理全组最强
0505-meitu-design-kit.md美图设计室(美图公司)电商设计工作流 + Agent Teams,数据最扎实的官方案例
0606-miaoya.md妙鸭相机(阿里)人像写真单点爆款,2025-09 团队解散的完整生命周期反面样本
0707-leonardo.mdLeonardo.ai多模型聚合 + PAYG API 的创作者平台
0808-runway.mdRunway视频为核心的多模态平台,MCP Server 唯一样本
0909-flux.mdFLUX.2(Black Forest Labs)开源 + API 双轨基座,结构化提示词与固定快照端点
1010-seedream.mdSeedream / SeedEdit(字节)统一生成编辑架构,多图参考虚拟试穿
1111-qwen-image.md通义万相 / Qwen-Image(阿里)开源权重 + 国产大厂路线,3.0 闭源断裂的 L4 风险案例
1212-faceswap-oss.md开源换脸换装生态VTON 与 Identity 两条技术路线对比,L6 护栏需自建
1313-liblibai.mdLiblibAI 2.0开源模型聚合 + 云端创作社区,社区资产库形态
1414-lovart.mdLOVART / 星流设计垂类 Agent,Brand Kit 为一等公民(星流将停服)
1515-weshop.mdWeShop(唯象)服饰换装商拍,SKU 一致性硬 Gate 的垂直样本
1616-piccopilot.mdPicCopilot(阿里国际)跨境营销素材广度路线(证据最弱,需官方复核)
阅读建议

首次进入本组,按 01(标杆)→ 04(治理最强)→ 05(编排最强)→ 12(开源对照)→ 本 README 的顺序阅读,可最短路径建立"六层 Harness"的分析直觉;选型场景可直接跳到第 5 节决策树。

7.1. 方法论边界与使用须知

为便于引用者正确使用本组结论,特别说明以下方法论边界:

  1. 评级是定性判断,不是量化评分:第 3 节矩阵中的 L1~L6 评级(弱 / 中 / 中强 / 强 / 最强)由各篇正文逐层分析汇总而来,评级依据在对应篇目的 5.x 节;跨平台比较时允许 ±1 级的主观弹性,不建议把矩阵直接转成数值打分。
  2. 证据强度不对称:16 个样本中,美图(财报)、Google(官方定价与案例)、火山引擎 / 阿里云 / Runway / Leonardo / FLUX(官方文档)证据较强;PicCopilot、LOVART、WeShop 的部分关键数据缺失(各篇信息缺口声明列明)。引用 PicCopilot 行结论前必须复核官方渠道。
  3. 时点敏感性:版本号与定价两类数据的半衰期极短。本组在检索窗口内即观察到三处重大变动(Midjourney V8.2、Qwen-Image 3.0 闭源、可灵图像线进入 O1),引用时应注明"2026-09 时点"。
  4. Harness 视角的取舍:本组刻意把"模型质量"(美学、画质)降为背景变量,聚焦工程承载层。若读者的决策变量主要是出图审美,则应补充阅读各篇之外的第三方盲测榜单,本组不提供该维度结论。
  5. 案例口径:除美图(财报数据)与 Google(THG Ingenuity / Max Fashion / L'agence)外,其余平台均未检索到带量化效果数据的官方客户案例;各篇正文如实标注,引用时不得以"行业广泛使用"等模糊表述替代。

AI Image Group Overview and Cross-Comparison

1. Group Overview

1.1. Research Question and Unified Framework

This group conducted market research on 16 platforms in the AI image domain (including 1 open-source ecosystem), using the AI Harness six-layer capability model defined in the project parameter sheet as the analytical framework:

LayerNameResponsibilityTypical observation targets in this group
L1Context engineering layerDetermines what the model "sees"Reference-image limit, long prompts, structured prompt contracts
L2Tools & execution layerDetermines what the model "can do"Generate/edit/outfit/face-swap tool sets, APIs, nodes
L3Orchestration & control layerDetermines "in what order to do things"Workflows, Agent Teams, multi-turn session iteration
L4Memory & state layerDetermines "what to remember"Asset libraries, LoRA, brand kits, reproducible artifacts
L5Evaluation & observability layerDetermines "how well it is done"Eval Sets, consistency gates, cost observability
L6Governance & safety layerDetermines "what cannot be done"Watermarks, labeling, moderation, licensing, audit

The core research conclusion of this group can be summarized in one sentence: the competition in the AI image industry has shifted from "whether the model produces good-looking results" to "whether the generation process is predictable" — Harness capabilities such as reference-image slots, reproducible endpoints, workflow version control, and content labeling are becoming the real differentiators among platforms.

1.2. Data Timeline and Evidence Grading

  • All platform data are snapshots as of the retrieval window of 2026-09-11 to 2026-09-12. The AI image domain iterates extremely quickly (within the retrieval window alone, Midjourney V8.2 launched, Qwen-Image 3.0 went closed-source, and the Kling image line entered O1 / 3.0 Omni), so version numbers, pricing, and ranking positions should be re-verified in real time before citing.
  • Evidence grading follows the retrieval-report standard: official sources (official websites, official documentation, financial reports, court rulings) > authoritative media > third-party reviews/communities (low-to-medium confidence, all annotated in the main text).
  • The "information gap statement" at the end of each document lists data that was not retrieved or lacks sufficient confidence; platforms with no public quantified case studies are truthfully marked as "not retrieved," never replaced by vague wording.

2. AI Image Market Panorama

Divided by product proposition and user segment, this group's 16 samples cover eight niches:

TrackCore propositionSamples in group
General text-to-image / image-to-image (creative)Aesthetic quality and style controlMidjourney, Leonardo.ai, FLUX.2
Instruction-driven generation & editing (multimodal)Multi-turn iteration that "changes an image with a single sentence"Gemini / Nano Banana, Jimeng AI, Seedream
Video-driven image linesImage and video generation on the same stackKling AI Image (Kolors / O1), Runway
Portrait photographySingle-hit identity generationMiaoya Camera (operations scaled back, a negative example)
E-commerce outfit change / commercial shootsSKU consistency + cost replacement of studio shootsMeitu Design Studio, WeShop, PicCopilot
Face swap / entertainment & socialEntertainment-oriented identity transferOpen-source ecosystem (ReActor / InstantID line) + reference-image capabilities of various platforms
Design platform / agentized designBriefing in, full asset suite outLOVART / Xingliu
Open-source self-hostingAutonomy over weights and orchestrationQwen-Image (1.0/2.0), FLUX klein, open-source face-swap/outfit ecosystem, LiblibAI (cloud aggregation)

Three structural trends run across all tracks:

  1. Reference images have become the main battleground: from Midjourney's --oref to Nano Banana Pro's 14 references + 5-person consistency, the ability to "put brand, people, and products into context" determines the ceiling of commercial value;
  2. The orchestration layer is moving up fast: Meitu Agent Teams, LOVART design agents, and ComfyUI workflows are three parallel forms that collectively turn the "multi-step production chain" into a product;
  3. The tension between open and closed source is becoming public: Qwen-Image 3.0 withdrawing its weights (2026-07) and FLUX.2 klein 4B sticking to Apache 2.0 form a contemporaneous contrast, making the sustainability of the open route a must-answer question in platform selection.

3. 16-Platform Cross-Comparison Matrix

平台开发商定位开放形态定价(时点价,详见各篇)L1L2L3L4L5L6适用边界
MidjourneyMidjourney Inc.(美)美学驱动的通用文生图Web + Discord$10/$30/$60/$120 月品牌创意、概念视觉
即梦 AI字节跳动(剪映团队)C 端一站式创作Web/App/小程序 + 火山 APIC 端 ;API 0.2~0.22 元/图中(曾被查处)中文创作者、短视频生态
可灵 AI 图像快手图像 + 视频统一创作Web/App + API国际站 $6.99~$127.99 月中强中强中高主体一致性创作、视频前置
Nano BananaGoogle DeepMind多模态指令式生成编辑Gemini 全家桶 + API/VertexNB2 1K $0.067;Pro $0.134**最强**企业生产、合规敏感场景
美图设计室美图公司电商设计工作流 + AgentWeb/桌面/App + 海外版¥30/¥88 起 **最强**电商卖家、批量商拍
妙鸭相机未序网络 / 阿里大文娱人像写真单点产品小程序/App¥9.9 / ¥29.9(2023 价)历史样本:单点爆款难留存
Leonardo.aiLeonardo Interactive(Canva 旗下)多模型聚合创作者平台Web + PAYG API$12/$30/$60;团队 $72/$144中强创作者规模化生产、API 集成
RunwayRunway AI(美)视频为核心的多模态平台Web + API + MCP$12/$28/$76 月;$0.01/credit中强影视级图像视频生产
FLUX.2Black Forest Labs(德)开源 + API 双轨图像基座API + HF 权重 + ComfyUIklein 4B $0.014 起(Apache 2.0);pro $0.03/MP**最强**排版文字、自建管线、可复现生产
Seedream / SeedEdit字节 Seed 团队统一生成编辑模型字节生态内嵌 + 火山 API0.2~0.22 元/张生态内生产、虚拟试穿多图参考
通义万相 / Qwen-Image阿里通义团队开源权重 + 云 API 双线HF 权重 + 百炼 API + Qwen Chat0.2 元/张(生成)、0.14 元/张(编辑)中→弱(3.0 断裂)中强图内有字场景、私有化部署
开源换脸换装生态学术界 + 社区换装(VTON)+ 换脸(Identity)方案集ComfyUI 节点 + 权重免费(许可分层)**最强****最强****最弱**自建管线、成本敏感批量场景
LiblibAI 2.0北京奇点星宇开源模型聚合 + 云端创作社区Web/App + 云 ComfyUI¥35/¥70/¥199 月 强(社区资产库)弱—中(曾遭央视点名)无 GPU 的开源生态用户
LOVART / 星流LiblibAI 体系设计垂类 AgentWeb**最强****最强**(Brand Kit)品牌全案设计、成套资产交付
WeShop蘑菇街团队孵化服饰换装商拍Web(国际站 + 国内站)中强中强强(SKU Gate)服饰电商换装、固定模特
PicCopilot阿里国际跨境电商营销素材Web框架强约束框架高风险跨境素材广度、多语言本地化

注:L1~L6 为六层 Harness 成熟度的定性评级(弱 / 中 / 中强 / 强 / 最强),依据各篇正文的逐层分析汇总;LiblibAI、LOVART、WeShop 三行依据组内对应文档的六层小结;PicCopilot 行证据最弱(其官方能力未获专题检索支持,仅框架性判断成立)。

4. Overall Six-Layer Harness Maturity Assessment

图 4-1|AI 图像平台的六层 Harness 能力模型

AI 图像平台的六层 Harness 能力模型 信息截止 2026-09-12 · 示意:基于本文分析绘制 L1 · 上下文工程层|决定模型“看到什么” 参考图上限、长提示词、结构化提示契约 L2 · 工具与执行层|决定模型“能做什么” 生成 / 编辑 / 换装 / 换脸工具集、API、节点 L3 · 编排与控制层|决定“按什么顺序做” 工作流、Agent Teams、多轮会话迭代 L4 · 记忆与状态层|决定“记住什么” 资产库、LoRA、品牌包、可复现工件 全行业短板 L5 · 评估与观测层|决定“做得好不好” Eval Set、一致性 Gate、成本观测 全行业短板 L6 · 治理与安全层|决定“不能做什么” 水印、标识、审核、许可、审计 结构解读:竞争已从“模型好不好看”转入“生成过程是否可预期”,六层成熟度构成统一分析框架。 全行业普遍短板在 L4(资产持久化)与 L5(评估标准)——正是平台间真正的分水岭。

数据来源:基于本文分析绘制的示意图。

4.1. Two Industry-Wide Weaknesses: L4 and L5

Putting all 16 samples side by side, the clearest industry conclusion is: the industry as a whole is weak in L4 (reproducibility and asset persistence) and L5 (lack of standards for effect evaluation).

  • The industry-wide L4 gap: most platforms do not treat "a user's generation history, parameters, reference images, and model versions" as portable first-class assets. Miaoya Camera is the extreme case — the digital avatar is locked inside a single product and cannot be migrated, which is considered one of the engineering root causes of its team disbanding in 2025-09. The three that do this best (Meitu's creative asset reuse, Runway Brand Kits, and ComfyUI workflows + LoRA all owned by the user) happen to be the samples with the healthiest commercialization or the most active ecosystems.
  • The industry-wide L5 gap: almost no platform publishes an official Eval Set or regression set. Industry evaluation mainly relies on three substitutes: vendor self-assessed leaderboards (Kling FlagEval, Qwen-Image topping 9 benchmarks — and notably Qwen's own self-assessed leaderboard shows an unfavorable fifth place, a rare exception), reverse-engineering from business metrics (Meitu's compute-point consumption and ARPPU multiples), and subjective human evaluation (Midjourney). The e-commerce vertical is the only scenario that turns evaluation into a hard gate (three-tier SKU consistency inspection), which precisely shows: without external hard constraints, vendors have no incentive to build an evaluation layer.
  • Reproducibility is closing the L4 gap: FLUX.2's fixed snapshot endpoints, ComfyUI's JSON workflow version control, and Midjourney's seed parameter system together turn "reproducing with the same parameters" from a wish into an engineering switch — this is the most valuable industry convergence this group observed.

4.2. Extreme Samples Across the Six Layers

LayerStrongest sampleWeakest sampleContrast implication
L1FLUX.2 (JSON structured prompts + hex brand colors + Grounding)Miaoya Camera (templates are everything)Degree of context structuring determines the controllability ceiling
L2ComfyUI open-source ecosystem / LiblibAI (nodes are tools)Midjourney / Miaoya (very short tool menus)Tool openness determines ecosystem density
L3Meitu Agent Teams / LOVART / ComfyUI workflowsMiaoya (atomic operations)Orchestration is the key to turning generation into production
L4LOVART Brand Kit / Runway Brand Kits / ComfyUI (assets owned by users)Miaoya (assets not migratable)Asset transferability determines retention
L5Meitu (closed-loop business metrics) / e-commerce route (SKU hard gate)Most platforms (no Eval Set)Without standards there is no way to regress
L6Nano Banana (SynthID + C2PA + public-figure restrictions)Open-source ecosystem (no mandatory labeling)Governance is a precondition for commercialization

4.3. Reference-Image Capability Comparison (L1 Core Metric)

PlatformReference-image limitCharacter consistencyMechanism name
MidjourneyStackableOmni Reference (--oref + --ow) / Style Reference
Jimeng / SeedreamOver a dozen (third-party API range 1–10)Multi-image reference + visual signals (Canny/Depth/Mask)
Kling O1 Image10 imagesAcross 100+ generations (channel claim)Identity Map (decoupling stable identity markers vs transient attributes)
Nano Banana Pro14 images5 peopleMulti-reference fusion + Search Grounding
Runway Gen-4 Image3 imagesLocks in with a single referenceReference Images (subject/scene/style division of labor)
FLUX.28 (API) / 10 (Playground); klein 48 consistent charactersMulti-Reference + Structured Prompting
Open source (ComfyUI)Determined by the workflowDetermined by LoRA / InstantID / PuLIDNode orchestration

4.4. Pricing Quick Reference

PlatformPayment modelKey unit priceConfidence
MidjourneySubscription$10 / $30 / $60 / $120 per monthMedium
Jimeng (Volcano Engine)Pay-per-use0.2 RMB/image (3.0 series), 0.22 RMB/image (4.0/4.6)Official, very high
KlingSubscription + Credits$6.99–$127.99/month; 660 Credits ≈ 3,300 imagesOfficial, high
Nano BananaSubscription + APINB2 1K $0.067; NB Pro $0.134Medium-high
Meitu Design StudioSubscription + compute points¥30 / ¥88 starting (third-party claim)Low
Miaoya CameraOne-time payment¥9.9 / ¥29.9 (2023 snapshot)Medium
Leonardo.aiSubscription + PAYG API$12 / $30 / $60; API from $5, 10 concurrentOfficial, high
RunwaySubscription + API$12 / $28 / $76; $0.01/creditOfficial, high
FLUX.2API + open-source weightsklein 4B from $0.014 (Apache 2.0); pro $0.03/MPOfficial, very high
Seedream (Volcano Engine)Pay-per-use0.2 / 0.22 RMBOfficial, very high
Tongyi Wanxiang / Qwen-ImagePay-per-use + open source0.2 RMB/image (generation); 0.14 RMB/image (editing)Official, very high
Open-source ecosystemFree (tiered licensing)Apache 2.0 / NCL / non-commercialPapers + HF, high
LiblibAISubscription + compute points¥35 / ¥70 / ¥199 per monthLow–medium
LOVART / XingliuSubscription
WeShopFree + subscription
PicCopilot

5. Platform Selection Decision Tree

A directly executable decision tree for each requirement scenario follows (for platforms at each branch endpoint, see the corresponding document):

你的核心场景是什么?
│
├── 电商换装 / 商拍
│   ├── 服饰类目、要固定模特与换装深度 → WeShop(15 篇)
│   ├── 跨境营销素材广度(多语言、模板、视频本地化)→ PicCopilot(16 篇)
│   ├── 批量设计工作流 + Agent 化 + 数据可核验 → 美图设计室(05 篇)
│   └── 成本敏感、有 GPU 运维能力、可自建合规 → 开源生态 CatVTON / Re-CatVTON(12 篇)
│
├── 品牌创意 / 营销视觉
│   ├── 美学优先、可接受弱编排 → Midjourney(01 篇)
│   ├── 需要联网事实与最强治理 → Nano Banana / Gemini(04 篇)
│   └── 中文语境、与短视频生态协同 → 即梦(02 篇)/ Seedream(10 篇)
│
├── 人像写真
│   ├── C 端单次写真 → 妙鸭相机(06 篇,注意其运营收缩状态)
│   ├── 需要主体/身份长期一致 → 可灵 O1 Identity Map(03 篇)/ Nano Banana 多参考(04 篇)
│   └── 有技术团队、要零样本身份注入 → 开源 InstantID / PuLID / EcomID(12 篇)
│
├── 换脸娱乐 / 影视后期
│   ├── 低分辨率快速迭代、自带 NSFW 检测 → ReActor(12 篇)
│   └── 高分辨率身份保持 → PuLID / EcomID(12 篇);视频侧 → Runway(08 篇)
│
├── 开源自建 / 私有化
│   ├── 图内有字(中文排版、信息图)→ Qwen-Image 1.0/2.0 权重(11 篇)
│   ├── 排版文字 + 可复现生产 → FLUX.2 flex / 固定快照端点(09 篇)
│   ├── 换装批量管线 → CatVTON(2.26GB 显存)(12 篇)
│   └── 无 GPU、要社区模型与工作流 → LiblibAI 云端(13 篇)
│
└── 设计中台 / 成套资产交付
    ├── 品牌全案(Logo 到短视频一揽子)→ LOVART(14 篇;注意星流国内版 2026-10-10 停服)
    └── 生成 API 深度集成进自家产品 → Leonardo PAYG(07 篇)/ 火山引擎 / 百炼 API

Three universal selection checks (regardless of which branch): ① whether the reference-image limit and character-consistency mechanism meet the scale of your business assets; ② whether a reproducibility means is provided (seed / fixed endpoint / workflow file); ③ whether L5 (consistency evaluation) and L6 (labeling compliance) are shouldered by the platform or by you — this is the real price difference between commercial platforms and the open-source ecosystem.

6. Compliance Chapter

AI image compliance constraints in China have completed the closed loop of "legislation → regulatory enforcement → judicial adjudication," and all documents in this group share this chapter as their baseline.

6.1. Regulatory Framework: the Labeling Measures and Deep Synthesis Rules

  • The Measures for Labeling AI-Generated Synthetic Content (Guoxinban Tongzi Issue No. 2 of 2025, jointly issued by the Cyberspace Administration of China, the Ministry of Industry and Information Technology, the Ministry of Public Security, and the National Radio and Television Administration, effective 2025-09-01):
  • Article 4: when providing download, copy, and export functions for generated synthetic content, explicit labeling that meets the requirements must be ensured in the file (for images, adding a prominent prompt label at an appropriate position);
  • Article 5: implicit labeling must be added to the file's metadata (generated-synthetic attributes, the service provider's name or code, content number, etc.), and adding digital watermarks is encouraged;
  • Article 6: dissemination platforms must verify the implicit labels in the metadata and handle them in three tiers (labeled / user-declared / traces detected), and append dissemination-attribute information;
  • Article 10 (red line): no organization or individual may maliciously delete, tamper with, forge, or conceal labels, nor provide tools or services for others to do so.
  • The supporting mandatory national standard, Cybersecurity Technology: Methods for Labeling AI-Generated Synthetic Content, was released at the same time.
  • The Provisions on the Administration of Deep Synthesis Services in Internet Information Services: Article 16 (implicit labeling) and Article 17 (deep synthesis services such as face replacement and face generation that could cause public confusion or misidentification must carry prominent labeling).
  • The Interim Measures for the Administration of Generative AI Services, Article 12: requires labeling of generated content.
  • Current compliance implementation in this group: apart from Nano Banana (SynthID invisible watermark + C2PA + tiered visible watermarks), no other platform has publicly disclosed labeling implementation; the open-source ecosystem (Doc 12) entirely requires users to build their own labeling components. This is a must-check item in platform selection.

6.2. Civil Boundaries: Portrait Rights Clauses of the Civil Code

  • Article 1018: a portrait is "the recognizable external image of a specific natural person reflected on a certain medium" — full identity is not required; "recognizable" alone qualifies for protection;
  • Article 1019: no organization or individual may infringe another's portrait right through defacement, smearing, or forgery using information-technology means; no portrait may be produced, used, or published without consent — AI face swapping directly falls under "forgery using information-technology means";
  • Article 1020: the fair-use circumstances (personal study, news reporting, official duty, public environment, etc.) do not include commercial face swapping;
  • Article 1023: a natural person's voice is protected by reference to the portrait-right rules (voice cloning / speech synthesis are likewise constrained).

6.3. Judicial Implementation: Beijing Internet Court's March 2026 Ruling

In a case where a well-known actor sued a short-drama producer and broadcaster over AI face-swap portrait-right infringement (effective 2026-03), four adjudication points have become the industry compliance baseline:

  1. "Recognizability" is the core adjudication standard: an AI face-swap image need not be completely identical to the original portrait — as long as the facial contours and features are highly similar and the general public can recognize them, using a specific natural person's portrait is established;
  2. Shifting of the burden of proof: a defendant claiming "coincidental AI resemblance" must reproduce the creation process; failure to reproduce incurs the adverse consequence of failing to prove — especially unfavorable to enterprises using open-source workflows (many parameters, hard-to-reproduce process), which should proactively retain generation logs and parameter snapshots;
  3. Copyright authorization cannot absorb portrait rights: obtaining authorization to the right of communication through information networks is not equivalent to authorization to use a portrait; a broadcaster that fails to fulfill reasonable review obligations likewise commits infringement;
  4. "Technological neutrality" is not a ground for exemption, and an extremely short duration of the disputed clip does not constitute a defense either.

6.4. Regulatory Enforcement: Jimeng AI Investigated and CCTV Naming

  • Jimeng AI (2026-04-28): was investigated and penalized by the cyberspace authorities for failing to effectively implement the requirements for labeling AI-generated synthetic content (per the Baidu Baike entry; re-verify against the official announcement). This is the only head-platform enforcement case within this group's research window, proving that enforcement of the Labeling Measures has reached first-tier big-vendor products.
  • LiblibAI (2026-04): was named by CCTV for serious regulatory loopholes in several AI software products; "Libu Libu AI" was said to be able to generate improper content using veiled prompts while remaining normally downloadable from app stores (re-verify against the regulatory announcement).
  • The two cases jointly show: the L6 governance layer is no longer a work item that can be retrofitted later, but a precondition for a platform's survival.

6.5. Criminal Risk Samples

  • Mr. Wang of Zhangjiagang, Jiangsu: AI face-swap defamation case — convicted of defamation, one year and three months in prison;
  • Mr. Sun of Ningyang, Shandong: AI face-swap combined with voice synthesis in an online-dating fraud case — convicted of fraud, three years' imprisonment with a four-year suspended sentence.

6.6. Compliance Point Comparison Across This Group's 16 Platforms

PlatformRelevant image-compliance risk surfacePublic labeling / watermark implementationCompliance conclusion in this group's documents
MidjourneyFace-swap / real-person material governance based mainly on community guidelinesNo official statement retrievedExplicit + implicit labeling implementation
Jimeng AIGenerated-synthetic content labeling (previously penalized); real-person material restrictionsNo official statement retrievedInvestigated by cyberspace authorities on 2026-04-28 (re-verify against official announcement)
Kling AI ImageFree-tier commercial-use restriction; real-person identity consistencyNo official statement retrievedPortrait risk constrained through membership agreement
Nano BananaPublic-figure portrait restrictionsSynthID mandatory invisible watermark + C2PA + tiered visible watermarksThe only public, complete labeling system in the whole group
Meitu Design StudioOutfit change involving real-person imageryEXIF metadata AI declaration (third-party claim; requires hands-on re-verification)Among the more compliant domestic tools (to be re-verified)
Miaoya CameraBiometric (face) information collection and user agreementNo official statement retrievedInitial user-agreement controversy → apology and revision (important compliance example)
Leonardo.aiFree-tier IP ownership; training-data usageNo official statement retrievedNeither input nor output used for training; clear rights transfer
RunwayTraining-data copyright litigationWatermark policy by tierGovernance mainly policy-based; litigation pending
FLUX.2Base-model licensing tiersApache 2.0 / NCL / closed-source API license boundaries clearest
SeedreamGenerated-synthetic content labeling (carrier Jimeng previously penalized)No official statement retrievedNo public labeling mechanism described at the model level
Tongyi Wanxiang / Qwen-ImageLabeling obligation of open-source deployment paths falls on the deployerNo labeling components providedUsing watermark-removal capability requires self-proving its legitimacy
Open-source face-swap/outfit ecosystemPortrait rights, deep-synthesis labeling, criminal-abuse riskNo mandatory labeling (ReActor has NSFW detection only)Safeguards must be built by users (the largest governance gap in the group)
LiblibAICommunity content-moderation loopholes (previously named by CCTV)Blockchain creation provenance (official claim)Platform governance is its weak spot
LOVART / XingliuGenerated-content labeling and brand-asset copyrightXingliu shutdown raises asset-migration compliance issues
WeShopFalse advertising (image-text mismatch), portrait rightsPlatform-side implementation [To be filled]Merchants must build their own "structure–attribute–compliance" three-tier inspection
PicCopilotFalse advertising (dual cross-border regulations)Both consistency and labeling require merchant confirmation item by item

6.7. Compliance Self-Check Checklist

Any team that uses AI image capabilities to deliver content externally should verify each item before launch:

  1. Has the output been given explicit labeling (prominent in-image notice)?
  2. Does the file metadata contain implicit labeling (attributes, service-provider code, content number)? If the platform does not provide it, have you built and injected it yourself?
  3. If real-person imagery is involved: has portrait authorization been obtained? Does the generated likeness pass the "recognizability" screening (err on the side of caution)?
  4. For face-swap / face-generation uses, is prominent labeling applied per Article 17 of the Deep Synthesis Rules?
  5. Have you retained a complete record of the reproducible creation process (to address the burden-of-proof shift)?
  6. For e-commerce scenarios: do generated images and the physical SKU pass the "structure–attribute–compliance" three-tier consistency inspection (to prevent false advertising)?
  7. Is there any misuse of capabilities such as label removal or watermark removal that touches the Article 10 red line of the Labeling Measures?

7. Document Navigation Index

This group contains 17 documents in total (16 platform studies + 1 group overview):

No.FileTopicOne-line positioning
README.mdGroup overview and cross-comparisonThis document: 16-platform comparison matrix, maturity assessment, decision tree, compliance chapter
0101-midjourney.mdMidjourneyThe aesthetic-driven general text-to-image benchmark, a "strong model + weak Harness" form factor
0202-jimeng.mdJimeng AI (ByteDance)Chinese-language all-in-one creation platform; a compliance negative example that was penalized
0303-kling-image.mdKling AI Image (Kuaishou)Kolors + O1 dual lines, Identity Map subject consistency
0404-nano-banana.mdGemini / Nano Banana (Google)Multimodal instruction-driven generation and editing; the strongest L6 governance in the group
0505-meitu-design-kit.mdMeitu Design Studio (Meitu)E-commerce design workflow + Agent Teams; the official case with the most solid data
0606-miaoya.mdMiaoya Camera (Alibaba)Single-hit portrait product; a full-lifecycle negative example with its team disbanding in 2025-09
0707-leonardo.mdLeonardo.aiA creator platform with multi-model aggregation + PAYG API
0808-runway.mdRunwayVideo-centric multimodal platform; the only MCP Server sample
0909-flux.mdFLUX.2 (Black Forest Labs)Open-source + API dual-track base; structured prompts and fixed snapshot endpoints
1010-seedream.mdSeedream / SeedEdit (ByteDance)Unified generation-and-editing architecture; multi-image reference virtual try-on
1111-qwen-image.mdTongyi Wanxiang / Qwen-Image (Alibaba)Open-source weights + domestic big-vendor route; the L4 risk case of the 3.0 closed-source break
1212-faceswap-oss.mdOpen-source face-swap/outfit ecosystemComparison of the VTON and Identity technical routes; L6 safeguards must be self-built
1313-liblibai.mdLiblibAI 2.0Open-source model aggregation + cloud creation community; a community asset-library form factor
1414-lovart.mdLOVART / XingliuDesign-vertical agent, with Brand Kit as a first-class citizen (Xingliu to be shut down)
1515-weshop.mdWeShop (Weixiang)Apparel outfit-change commercial shoots; a vertical sample with SKU consistency as a hard gate
1616-piccopilot.mdPicCopilot (Alibaba International)Cross-border marketing-asset breadth route (weakest evidence; requires official re-verification)
阅读建议

For a first entry into this group, read in the order 01 (benchmark) → 04 (strongest governance) → 05 (strongest orchestration) → 12 (open-source comparison) → this README, the shortest path to building an analytical intuition for the "six-layer Harness"; for selection scenarios, you can jump directly to the decision tree in Section 5.

7.1. Methodological Boundaries and Usage Notes

To help those citing this group's conclusions use them correctly, the following methodological boundaries are explained:

  1. Ratings are qualitative judgments, not quantitative scores: the L1–L6 ratings (Weak / Medium / Medium-strong / Strong / Strongest) in the Section 3 matrix are aggregated from each document's layered analysis, with the basis given in the corresponding document's 5.x sections; a subjective elasticity of ±1 level is allowed when comparing across platforms, and it is not recommended to convert the matrix directly into numeric scoring.
  2. Evidence strength is asymmetric: among the 16 samples, Meitu (financial reports), Google (official pricing and cases), and Volcano Engine / Alibaba Cloud / Runway / Leonardo / FLUX (official documentation) have relatively strong evidence; some key data of PicCopilot, LOVART, and WeShop is missing (listed in each document's information-gap statement). Official channels must be re-verified before citing the PicCopilot row's conclusions.
  3. Point-in-time sensitivity: version numbers and pricing have extremely short half-lives. Within the retrieval window alone, this group observed three major changes (Midjourney V8.2, Qwen-Image 3.0 going closed-source, and the Kling image line entering O1); citations should note the "2026-09 snapshot."
  4. Trade-offs from the Harness perspective: this group deliberately downgraded "model quality" (aesthetics, image fidelity) to a background variable, focusing on the engineering carrier layer. If a reader's main decision variable is output aesthetics, they should additionally consult third-party blind-test leaderboards beyond these documents; this group does not provide conclusions in that dimension.
  5. Case scope: apart from Meitu (financial-report data) and Google (THG Ingenuity / Max Fashion / L'agence), no official customer cases with quantified effectiveness data were retrieved for the other platforms; each document's main text marks this truthfully, and citations must not substitute vague wording such as "widely used in the industry."