Gemini / Nano Banana(Google)
1. 介绍
1.1. 平台概况
Nano Banana 是 Google DeepMind 的 Gemini 图像模型产品名系列,覆盖从入门到专业级的三档模型。在 AI Harness 六层能力模型中,Nano Banana 是本组 L6 治理与安全层最强、且唯一有公开可验证水印机制的闭源平台:所有 Google AI 生成图强制嵌入 SynthID 不可见水印,Nano Banana Pro 支持 C2PA 元数据标准,可见水印按订阅档位分级,并对明确受版权保护的特定公众人物肖像设限。
同时它是本组唯二有带量化数据商业案例的平台(另一为美图,数据来自财报),这使其成为研究"Harness 能力如何转化为可验证商业结果"的关键样本。
| 项 | 内容 | 置信度 |
|---|---|---|
| 开发商 | Google DeepMind | 极高 |
| 产品名映射 | Nano Banana = Gemini 2.5 Flash Image(入门);Nano Banana Pro = Gemini 3 Pro Image(2025-11-21 发布);Nano Banana 2 = Gemini 3.1 Flash Image | 中高 |
| 参考能力 | 最多 14 张参考图融合;复杂构图中保持最多 5 个人物的面部与服装细节 | 高(官方口径) |
| 输入长度 | 支持 64k 输入 token 的极长提示词 | 高 |
| 水印 | SynthID 强制嵌入、官方口径不可移除;Nano Banana Pro 支持 C2PA | 高 |
1.2. 模型谱系
| 产品名 | 对应模型 | 定位 | 输出上限 |
|---|---|---|---|
| Nano Banana(初代) | Gemini 2.5 Flash Image | 入门级 | 最高 1024px |
| Nano Banana Pro | Gemini 3 Pro Image | 专业级,2025-11-21 发布 | 1K / 2K / 4K |
| Nano Banana 2 | Gemini 3.1 Flash Image | 新增档位 | 512px / 1K / 2K / 4K |
在 Gemini App 中切换至思考型模型即使用 Nano Banana Pro。
1.3. 开放形态与定价
开放形态(多渠道一致,高置信):Gemini App / Web、Google AI Studio、Gemini API、Vertex AI、Google Ads、Slides、Vids、NotebookLM、Flow、Google Search(AI Mode)。
订阅入口(第三方来源,中置信):Google AI Plus $9.99/月(美国)、Google AI Pro $19.99/月;Ultra 档更高配额。
API 定价:
官方定价页未能直接抓取(ai.google.dev 因登录跳转未直抓)→ 官方实时定价 [待填写],成稿前须复核 https://ai.google.dev/pricing。
下表为 2026-02-26 时点的官方定价页截图整理值,,仅供量级参考,不得作为合同依据。
| 模型 | 512px | 1K | 2K | 4K | Token 单价 |
|---|---|---|---|---|---|
| Nano Banana 2(Gemini 3.1 Flash Image) | $0.045 | $0.067 | $0.101 | $0.151 | $60/百万 token |
| Nano Banana Pro(Gemini 3 Pro Image) | — | $0.134 | $0.134 | $0.240 | $120/百万 token |
| Nano Banana(初代,Gemini 2.5 Flash Image) | — | $0.039 | $0.039 | — | $30/百万 token |
2. 名词解释
2.1. AI 图像通用术语
| 术语 | 英文 / 缩写 | 释义 | Nano Banana 的对应实现 |
|---|---|---|---|
| 文生图 | Text-to-Image(T2I) | 仅由文本提示词生成图像 | 支持,且提示词可长达 64k token |
| 图生图 | Image-to-Image(I2I) | 以一张或多张图像为条件生成新图像 | 支持,最多 14 张参考图融合 |
| 局部重绘 | Inpainting | 对指定区域重新生成,区域外保持不变 | 被指令式编辑取代:可选或描述区域,无需遮罩 |
| 外扩 | Outpainting | 在画布外扩区域继续生成 | 通过多宽高比(含 21:9)与构图控制间接实现,[待填写] |
| 可控生成 | ControlNet | 以 Canny、Depth、Pose、Mask 等视觉信号控制生成结构 | 不使用 ControlNet;结构由"相机角度 / 焦点 / 景深 / 构图比例"等自然语言参数控制 |
| 参考图 | Reference Image | 作为身份、风格、结构约束输入的图像 | 硬编码上限 14 张,人物一致性上限 5 人 |
| 随机种子 | Seed | 固定后可在同参数下复现同一张图 | 未公开种子锁定与快照端点说明,[待填写] |
| 引导强度 | CFG | 提示词对生成结果的约束强度 | 不适用(自回归多模态模型,非扩散引导范式) |
| 低秩适配 | LoRA | 小参数量微调模块,用于固化人物、风格、服装资产 | 不支持用户侧 LoRA 训练与加载 |
| 图像提示适配 | IP-Adapter | 用图像编码器特征注入注意力,实现"以图为提示词" | 不适用(原生多模态直接接收图像输入) |
| 零样本身份注入 | InstantID | 单张参考图、无需微调即可迁移身份 | 不适用;身份保持由多参考融合(最多 5 人)原生实现 |
2.2. Nano Banana 特有术语
| 术语 | 英文 / 缩写 | 释义 |
|---|---|---|
| 不可见水印 | SynthID | Google 的像素级不可见数字水印,所有 Google AI 生成图均强制嵌入;可在 Gemini App 中上传图片追问"是否由 Google AI 生成"来检测;官方口径:不可移除 |
| 可见水印 | Gemini Sparkle(Gemini 星光水印) | 免费层与 Google AI Pro 用户保留;Google AI Ultra 订阅者,或经 Google AI Studio / Vertex AI 生成的图像可移除 |
| 内容来源与真实性联盟 | C2PA(Coalition for Content Provenance and Authenticity) | 内容来源与真实性元数据标准;Nano Banana Pro 支持(百度百科口径,中置信) |
| 搜索接地 | Grounding with Search | 接 Google 搜索知识库,基于实时信息(天气、赛事、食谱)或专业知识库生成图表与信息图 |
| 思考型模型 | Thinking Model | Gemini App 中切换至思考型模型即使用 Nano Banana Pro |
| 多参考一致性 | Multi-Reference Consistency | 最多 14 张参考图融合,在复杂构图中保持最多 5 个人物的面部与服装细节 |
| 生产简报 | Production Brief | 官方推荐的结构化提示词写法:subject / purpose / 文本 / 语言 / 构图 / 风格 / 宽高比 / 分辨率 / 必须保持不变的元素 |
| 多语言文字渲染 | Multilingual Text Rendering | 支持含繁体中文与简体中文在内的多语言文字渲染 |
| 手写草图转写实 | Sketch-to-Realistic | 把手写草图转换为写实图像 |
| 企业侧生成 | Vertex AI Generation | 经 Vertex AI 生成的图像可移除可见水印,并支持企业安全机制与弹性计费 |
2.3. 换装与换脸方向通用术语
说明:Nano Banana 不提供独立的虚拟试穿或换脸产品能力,但其多参考一致性(14 参考图 / 5 人物)与搜索接地使其成为第三方试穿方案(如 THG Ingenuity AI Stylist、Max Fashion 数字试衣间、ComfyUI 官方 VTON 工作流)最常用的生成底座之一。
| 术语 | 英文 / 缩写 | 释义 | Nano Banana 相关性 |
|---|---|---|---|
| 虚拟试穿 | VTON(Virtual Try-On) | 将目标服装"穿"到指定人物图像上并生成视觉可信结果 | 作为底座被第三方试穿方案调用(ComfyUI 官方 VTON 模板即使用 Nano Banana Pro) |
| 服装形变 | Garment Warping | 用 TPS(薄板样条)等几何变换把平铺服装对齐到人体姿态 | 不适用(不暴露形变控制) |
| 服装掩码 | Cloth Mask | 人体解析得到的服装区域二值图 | 不适用(指令式编辑无需掩码) |
| 试穿扩散 | Try-on Diffusion | 以扩散模型端到端完成服装与人体融合 | 不适用(非扩散范式) |
| 换脸 | Face Swap | 把 A 的脸替换到 B 的面部位置 | 不提供,且明确限制生成受版权保护的特定公众人物肖像 |
| 人脸重演 | Face Reenactment | 保留身份、迁移表情、口型与头部姿态 | 不提供 |
| 身份保持 | Identity Preservation | 生成结果在多大程度上仍"是那个人" | 通过多参考融合实现(最多 5 人面部细节保持) |
3. 功能说明
3.1. 生成与编辑能力
| 能力 | 支持 | 说明 |
|---|---|---|
| 文生图 / 图生图 | 支持 | 原生多模态输入 |
| 指令式编辑 | 支持 | 自然语言描述,无需遮罩 |
| 多图融合 | 支持 | 最多 14 张参考图 |
| 风格迁移 | 支持 | — |
| 局部编辑 | 支持 | 可选或描述区域;可调整相机角度、焦点、色彩分级、景深、昼夜光照 |
| 2K / 4K 输出 | 支持 | 按分辨率分档计价 |
| 多宽高比 | 支持 | 含 21:9 |
| 多语言文字渲染 | 支持 | 含繁体中文与简体中文 |
| 信息图与数据可视化 | 支持 | 结合 Search Grounding |
| 手写草图转写实 | 支持 | — |
| 搜索接地生成 | 支持 | Grounding with Search |
3.2. 多参考一致性与物理光学控制
- 多参考上限:14 张参考图融合。
- 人物一致性上限:复杂构图中保持最多 5 个人物的面部与服装细节。这两个数字是硬编码的上下文槽位,是本组中最明确公开的上下文容量契约(对比:可灵 O1 为 10 张、Runway Gen-4 为 3 张、FLUX.2 为 8~10 张、Midjourney 为可堆叠但无公开上限)。
- 物理与光学控制:可指定光线氛围、镜头类型、景深、视角、构图比例——把摄影学参数变成提示词的一等字段。
3.3. 结构化生产简报
官方推荐把提示词写成结构化生产简报,字段包括:
| 字段 | 说明 |
|---|---|
| subject | 主体 |
| purpose | 用途 |
| 文本 | 图中需渲染的文字 |
| 语言 | 文本语言 |
| 构图 | 构图方式 |
| 风格 | 视觉风格 |
| 宽高比 | 输出比例 |
| 分辨率 | 输出尺寸 |
| 必须保持不变的元素 | 一致性约束清单 |
这与 FLUX.2 的官方 JSON 字段(subject / background / lighting / style / camera_angle / composition)在思路上一致:把上下文从自由文本升级为可校验的字段契约。差别在于 FLUX.2 是机器可读的 JSON,Nano Banana 是官方推荐的自然语言写法规范。
3.4. 已知限制
- 无法生成涉及暴力、仇恨煽动的内容。
- 无法生成明确受版权保护的特定公众人物肖像——这是本组中少见的、被官方明确写出的身份限制。
- 可见水印按档位保留:免费层与 Google AI Pro 保留 Gemini 星光水印;Ultra 订阅者或经 AI Studio / Vertex AI 生成可移除。
- 定价需实时复核:官方定价页因登录跳转未能直接抓取,[待填写]。
4. 平台架构
图 4-1|Nano Banana 平台架构:入口—推理—底座三层
数据来源:基于本文分析绘制的示意图。
4.1. 原生多模态底座
- 底座为 Gemini 3 Pro / 3.1 Flash 多模态大模型,原生多模态(图文混合输入与输出),而非外挂扩散管线。
- 这一差异带来连锁影响:
- 无 CFG / Steps / Sampler 等扩散范式参数;
- 无 LoRA / ControlNet / IP-Adapter 生态位(这些是扩散模型的配套机制);
- 一致性靠模型自身的上下文理解,而非外部条件注入。
4.2. 推理范式
默认包含 reasoning 步骤:先理解意图与画面逻辑,必要时联网检索(Grounding),再生成。这与即梦 Seedream 5.0 的"联网检索 + CoT 思维链推理"是同一范式,也是本组可见的第三代生成架构共识——先想再画。
4.3. 分发矩阵
| 侧 | 入口 |
|---|---|
| 消费端 | Gemini App / Web、Google Search(AI Mode)、Workspace(Slides / Vids)、NotebookLM、Google Ads、Flow |
| 开发端 | Google AI Studio、Gemini API、Vertex AI、Google Antigravity |
企业侧关键差异:经 Vertex AI 生成的图像可移除可见水印,并支持企业安全机制与弹性计费。这意味着同一模型在"消费入口"与"企业入口"下有不同的合规与治理属性——这是选型时容易被忽略的关键点。
5. Harness 设计
5.1. 六层能力总览
| 层 | 名称 | Nano Banana 的实现 | 成熟度 | 证据强度 |
|---|---|---|---|---|
| L1 | 上下文工程 | 14 参考图 / 5 人物硬编码槽位;64k 输入 token;Search Grounding 把外部检索纳入上下文;结构化生产简报 | 强 | 高 |
| L2 | 工具与执行 | 生成 + 编辑 + 检索(Grounding)+ 多语言渲染;经 Gemini API 与 Vertex AI 暴露;Vertex AI 支持企业安全与弹性计费 | 中 | 高 |
| L3 | 编排与控制 | 多轮会话式迭代(生成 → 文本批评 → 定向修改);官方明确建议"一次只请求一个受控变更" | 中 | 高 |
| L4 | 记忆与状态 | 依赖 Google 生态持久化(Drive / NotebookLM / Workspace);未见面向"品牌资产包"的一等公民抽象 | 中 | 中 |
| L5 | 评估与观测 | 官方强调发布前逐字、逐数、逐脸、逐产品细节核验的人工验证步骤;未公开官方 Eval Set | 中 | 中 |
| L6 | 治理与安全 | 本组最强:SynthID 强制不可见水印 + C2PA + 可见水印分级 + 公众人物肖像限制 + Gemini App 内可验证入口 | 最强 | 高 |
5.2. L1 上下文工程层
Nano Banana 的 L1 有三个可度量特征,这在本组中极为罕见:
- 明确的容量契约:14 张参考图、5 个人物。这不是"支持多图"这类模糊表述,而是可写进技术方案与验收标准的硬上限。
- 超长提示词:支持 64k 输入 token,可把整份品牌规范、产品目录、历史素材说明一次性塞入上下文。
- 外部知识接地:Search Grounding 把 Google 搜索结果纳入生成上下文,使"按实时事实出图"成为可能(天气、赛事、食谱、专业知识库)。
官方推荐的结构化生产简报(9 个字段)进一步把上下文组织从"写作技巧"变成"填写规范"。这是本组除 FLUX.2 官方 JSON 契约外,唯一由平台官方给出的上下文组织规范。
5.3. L2 工具与执行层
- 工具集:生成、编辑、检索(Grounding)、多语言渲染。
- 暴露方式:Gemini API 与 Vertex AI。Vertex AI 侧额外提供企业安全机制与弹性计费。
- 无开放工具注册:不支持把外部工具注册进生成流程(对比 Leonardo.ai 支持把自定义精调模型注册为可调用 model ID)。
- 生态集成:被第三方平台与工作流聚合(如 ComfyUI 官方 VTON 工作流使用 Nano Banana Pro 作为生成节点),说明其 API 契约具备良好的可组合性。
5.4. L3 编排与控制层
- 多轮会话式迭代:生成 → 用文本批评 → 定向修改。
- 官方明确的操作纪律:"一次只请求一个受控变更"。这是一条罕见的官方编排指引——它把"变更粒度"上升为使用规范,实质上是在管理上下文漂移风险:一次改多处,未受控的部分也可能被模型顺手改掉。
- 仍缺:无 DAG、无子智能体派发、无工作流工件导出与版本控制、无中断恢复。
与可灵的"生成即对话"相比,两者范式相同,差别在于可灵强调"无需遮罩",Nano Banana 强调"一次一改"。
5.5. L4 记忆与状态层
- 依赖 Google 生态持久化:Drive、NotebookLM、Workspace。
- 未见面向"品牌资产包"的一等公民抽象——这是与 Runway 的 Brand Kits(最多 3 个)、可灵的主体创建(限量计价)、美图设计室的创作资产复用(官方明示)对比时的明显短板。
- 无 Checkpoint、无工件版本管理、无模型版本锁定端点的公开说明。
工程含义:如果企业需要"锁定的品牌角色 / 锁定 logo / 锁定色板"这类资产,Nano Banana 侧没有原生抽象,必须由业务系统自行实现并在每次请求时重新注入上下文——这会增加 64k token 上下文的占用与运维成本。
5.6. L5 评估与观测层
- 官方强调人工验证步骤:发布前逐字、逐数、逐脸、逐产品细节核验。这实质上是一份官方发布的验收清单,虽然不是自动化 Eval Set,但具备可操作性与可执行性,优于多数平台"无评估机制"的现状。
- 未公开官方 Eval Set、Golden Dataset 或回归集。
- 观测口径:API 侧按分辨率与 token 计价,成本可测算;但官方定价页未能直抓,[待填写]。
5.7. L6 治理与安全层
这是 Nano Banana 在本组中决定性领先的一层,也是本组唯一能对应中国《标识办法》显式 + 隐式双标识要求的闭源平台:
| 治理机制 | 实现 | 对应《标识办法》条款 |
|---|---|---|
| 隐式标识 | SynthID 强制嵌入,官方口径不可移除 | 第五条:应当添加隐式标识(元数据 / 数字水印) |
| 元数据标准 | Nano Banana Pro 支持 C2PA | 第五条:鼓励添加内容来源信息 |
| 显式标识 | 可见水印(Gemini 星光水印),按档位分级:免费层与 Pro 保留;Ultra / AI Studio / Vertex AI 可移除 | 第四条:提供下载导出时应当确保显式标识 |
| 可验证入口 | 在 Gemini App 中上传图片追问"是否由 Google AI 生成"即可检测 | 第六条:传播平台核验隐式标识 |
| 身份限制 | 不生成明确受版权保护的特定公众人物肖像 | 《民法典》第一千零一十九条 |
必须如实声明:以上机制对应《标识办法》的映射关系,是本报告依据公开资料作出的对照分析,并非 Google 官方对中国法规的合规声明。Google 是否按中国《标识办法》要求实现了"服务提供者名称或者编码、内容编号"等具体元数据字段,未见公开说明,[待填写]。
同时,可见水印在 Ultra / AI Studio / Vertex AI 入口可移除,这在企业侧是合理设计(企业自有素材),但也意味着仅依赖平台默认行为不足以满足显式标识义务,业务侧仍需自建注入环节。
5.8. 成熟度判断
Nano Banana 是"强模型 + 中上 Harness + 最强治理"形态。它的 L1 有明确容量契约与接地能力,L3 有官方操作纪律,L5 有可执行的官方验收清单,L6 是全组唯一具备"强制不可见水印 + 元数据标准 + 可验证入口"三重机制的闭源平台。短板在 L4——缺少品牌资产的一等公民抽象,这使它在"长期、多批次、需锁定资产"的生产场景中需要额外的外部承载层。
6. 实际案例
6.1. THG Ingenuity × Google Cloud "AI Stylist"
- 方案:基于 Gemini Enterprise Agent Platform 的虚拟试衣方案,已上架 Google Cloud Marketplace。
- 量化数据(Myprotein 初期部署):使用过至少一次 AI Stylist 的用户,对比站内运动服饰购物均值——
- 英国用户转化可能性接近 6 倍;
- 站内停留时长 6.3 倍;
- 平均客单价高 2.5%。
- 来源:Prolific North,相关度高(含量化数据)。
- 性质说明:数据来自方案方初期部署的公开报道,属厂商侧披露数据,引用时应注明来源与口径。
6.2. Max Fashion × Google Cloud
- 方案:中东零售巨头 Max Fashion 基于 Gemini Enterprise 平台推出企业级数字试衣间,生成高拟真实时预览,呈现服装在多样真实体型上的垂坠与动态。
- 引述(Max Fashion 全渠道高级副总裁 Bala Subramaniam):"虚拟试衣从根本上缩短了线上与线下体验的距离。手机浏览的顾客第一次能拥有与站在实体试衣间内相同的选购信心。"
- 来源:fabrix.hk 引 Retail Management Middle East,相关度中高。
- 量化数据:未披露具体转化或退货率数字。
6.3. L'agence × Google AI(2026-02 纽约时装周)
- 内容:嘉宾即时可视化自己穿走秀造型,支持即时预订。
- 来源:Business of Fashion,相关度中高。
- 量化数据:未披露。
7. 总结
7.1. 优势
- L6 治理全组最强:SynthID 强制不可见水印(官方口径不可移除)+ C2PA + 可见水印分级 + 公众人物肖像限制 + App 内可验证入口,形成完整的"生成—标识—可验证"闭环。
- 上下文容量契约明确:14 参考图 / 5 人物 / 64k token,可写进技术方案与验收标准。
- Search Grounding:把实时外部事实纳入生成上下文,支持信息图与数据可视化。
- 官方验收清单可用:逐字、逐数、逐脸、逐产品细节核验,可直接转化为企业 DoD。
- 是唯二有量化商业案例的平台:THG Ingenuity 转化接近 6 倍、停留时长 6.3 倍。
- 分发矩阵最广:从消费端 Search / Workspace 到开发端 API / Vertex AI,且企业入口(Vertex AI)在合规与计费上独立于消费入口。
7.2. 局限与适用边界
| 局限 | 影响 |
|---|---|
| 无品牌资产一等抽象 | 需外部系统维护并在每次请求重新注入,占用上下文 |
| 无工作流工件与版本管理 | 不可回归、不可 diff |
| 无模型版本锁定端点 | 模型升级后输出漂移 |
| 官方定价页未能直抓 | 成本测算不可靠,[待填写] |
| 可见水印在部分入口可移除 | 显式标识义务仍需业务侧自建 |
| 不支持 LoRA / ControlNet | 扩散生态资产无法复用 |
| 中国法规合规声明缺失 | 与《标识办法》的映射为报告分析,非官方声明 |
适用边界:适合需要强合规背书、需要实时事实接地、需要消费级触达(Search / Workspace)的场景;不适合需要长期锁定品牌资产、需要完全离线或私有化、需要精确结构化控制(如 FLUX.2 的 hex 品牌色)的生产场景。
7.3. 选型建议
- 企业侧一律走 Vertex AI:可移除可见水印、支持企业安全机制与弹性计费,且消费入口与合规属性不同。
- 自建品牌资产层:把品牌规范、角色参考、色板固化为外部资产库,按官方 9 字段生产简报格式在每次请求时注入。
- 遵循"一次一改"纪律:多轮迭代时每次只请求一个受控变更,避免未受控区域被顺带修改。
- 落实验收清单:把官方"逐字、逐数、逐脸、逐产品细节核验"写进 DoD,尤其涉及文字渲染与人脸时。
- 定价必须实时复核:成稿与采购前须访问 ai.google.dev 定价页确认当前单价。
- 不要依赖平台默认标识:即便 SynthID 强制嵌入,中国《标识办法》要求的显式标识与特定元数据字段仍需业务侧在导出链路中自行注入。
7.4. 合规提示
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号):第四条显式标识义务、第五条隐式标识义务(含服务提供者名称或编码、内容编号)、第六条传播平台核验义务、第十条不得恶意删除篡改伪造隐匿标识。
- 《中华人民共和国民法典》第一千零一十九条:不得以利用信息技术手段伪造等方式侵害他人肖像权。Nano Banana 对公众人物肖像的限制是平台侧护栏,不能替代权利人对权利人的授权。
- 北京互联网法院 2026-03 生效判决:可识别性 + 举证责任转移。即使平台有公众人物限制,企业仍应保存生成参数、参考图与提示词,以备举证。
信息缺口声明
- 官方定价页:ai.google.dev 因登录跳转未能直接抓取,官方实时定价 [待填写];表内数值为 2026-02-26 时点截图整理值,。
- C2PA 支持:来自百度百科口径,非 Google 官方文档原文。
- 各模型在中国的可用性:Gemini App / API 在中国大陆的可访问性与合规路径未检索到官方说明。[待填写]
- 与《标识办法》的合规映射:报告依据公开资料作出的对照分析,非 Google 官方合规声明;是否实现"服务提供者名称或编码、内容编号"等具体元数据字段 [待填写]。
- 种子锁定、模型版本锁定端点、快照机制:未检索到官方说明。[待填写]
- Google AI Plus / Pro / Ultra 的配额细节:仅第三方来源给出价格,配额未详。[待填写]
- Nano Banana 2(Gemini 3.1 Flash Image)的发布时间:检索结果仅标注"新增",未给出日期。[待填写]
8. 参考资料
- Google AI · Gemini API 定价页(官方,成稿前须复核) — Google。https://ai.google.dev/pricing
- The Rundown AI · Nano Banana Pro 能力、定价与入口(第三方,中置信) — https://www.therundown.ai/tools/nano-banana-pro
- iKala · Nano Banana Pro 功能与方案介绍(含模型对比表,第三方) — https://ikala.ai/zh-tw/blog/ikala-ai-insight/nano-banana-pro-functions-how-to-use/
- Brand Icon Image · Google Introduces Nano Banana Pro(官方发布会转述,第三方) — http://www.brandiconimage.com/2025/11/google-introduces-nano-banana-pro-with.html
- FoneArena · Google rolls out Nano Banana Pro powered by Gemini 3(第三方) — https://www.fonearena.com/blog?p=469327
- 百度百科 · Nano Banana Pro(含 SynthID / C2PA / 14 参考图 / 5 人物,二次来源) — https://baike.baidu.com/item/BananaPro/67405719
- Prolific North ·《THG Ingenuity predicts 600% conversion rate boost with Google AI Stylist rollout》(含量化数据) — https://www.prolificnorth.co.uk/news/thg-ingenuity-predicts-600-conversion-rate-boost-with-google-ai-stylist-rollout/
- Business of Fashion ·《How To Drive Luxury Conversion Through AI-Powered Virtual Try-On》(含 L'agence 案例) — https://www.businessoffashion.com/articles/technology/ai-virtual-try-on-luxury-dressx-report/
- FabriX ·《How AI Is Solving Fashion's Fitting Room Crisis & Driving E-Commerce Sales》(含 Max Fashion × Google Cloud) — https://fabrix.hk/blog/how-ai-is-solving-fashions-fitting-room-crisis
- ComfyUI 官方工作流「虚拟角色试穿 - 四合一」(使用 Nano Banana Pro 作为生成节点) — Comfy Org。https://comfy.org/zh/workflows/templates_rob_fashion_shoot_vton-4in1.app/
- 《人工智能生成合成内容标识办法》(国信办通字〔2025〕2 号) — 国家网信办等,2025-09-01 施行。https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- 《中华人民共和国民法典》第一千零一十九条 — 全国人民代表大会,2020。(正文引用条文,无官方链接)
- 经济参考报 ·《技术不是侵权"挡箭牌" 法院这样认定 AI"盗脸"》 — 新华社《经济参考报》,2026-04-17。http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm
Gemini / Nano Banana (Google)
1. Introduction
1.1. Platform Overview
Nano Banana is Google DeepMind's product-name series for Gemini image models, spanning three tiers from entry-level to professional. Within the AI Harness six-layer capability model, Nano Banana is the strongest closed-source platform in this group on the L6 Governance & Security layer, and the only one with a publicly verifiable watermarking mechanism: all Google AI-generated images forcibly embed the invisible SynthID watermark, Nano Banana Pro supports the C2PA metadata standard, visible watermarks are tiered by subscription plan, and it restricts portraits of specific public figures under clear copyright protection.
It is also one of only two platforms in this group with quantified commercial case studies (the other is Meitu, with data from financial reports), which makes it a key sample for studying how "Harness capabilities translate into verifiable business results."
| Item | Content | Confidence |
|---|---|---|
| Developer | Google DeepMind | Very High |
| Product name mapping | Nano Banana = Gemini 2.5 Flash Image (entry-level); Nano Banana Pro = Gemini 3 Pro Image (released 2025-11-21); Nano Banana 2 = Gemini 3.1 Flash Image | Medium-High |
| Reference capability | Up to 14 reference images fused; keeps facial and clothing detail of up to 5 people in complex compositions | High (official statement) |
| Input length | Supports very long prompts of 64k input tokens | High |
| Watermark | SynthID forcibly embedded, officially stated as non-removable; Nano Banana Pro supports C2PA | High |
1.2. Model Lineage
| Product name | Corresponding model | Positioning | Output cap |
|---|---|---|---|
| Nano Banana (first-gen) | Gemini 2.5 Flash Image | Entry-level | Up to 1024px |
| Nano Banana Pro | Gemini 3 Pro Image | Professional, released 2025-11-21 | 1K / 2K / 4K |
| Nano Banana 2 | Gemini 3.1 Flash Image | New tier | 512px / 1K / 2K / 4K |
Switching to the Thinking Model in the Gemini App uses Nano Banana Pro.
1.3. Open Access Forms and Pricing
Open access forms (consistent across channels, high confidence): Gemini App / Web, Google AI Studio, Gemini API, Vertex AI, Google Ads, Slides, Vids, NotebookLM, Flow, Google Search (AI Mode).
Subscription entry points (third-party sources, medium confidence): Google AI Plus $9.99/month (US), Google AI Pro $19.99/month; Ultra tier has higher quotas.
API pricing:
The official pricing page could not be scraped directly (ai.google.dev could not be fetched directly due to a login redirect) → official real-time pricing
[To be filled]; must be re-verified against https://ai.google.dev/pricing before finalizing.
The table below contains values compiled from screenshots of the official pricing page as of 2026-02-26, for order-of-magnitude reference only and not to be used as contractual basis.
| Model | 512px | 1K | 2K | 4K | Price per token |
|---|---|---|---|---|---|
| Nano Banana 2 (Gemini 3.1 Flash Image) | $0.045 | $0.067 | $0.101 | $0.151 | $60/million tokens |
| Nano Banana Pro (Gemini 3 Pro Image) | — | $0.134 | $0.134 | $0.240 | $120/million tokens |
| Nano Banana (first-gen, Gemini 2.5 Flash Image) | — | $0.039 | $0.039 | — | $30/million tokens |
2. Glossary
2.1. Common AI Image Terms
| Term | English / Abbreviation | Definition | Nano Banana's implementation |
|---|---|---|---|
| Text-to-image | Text-to-Image (T2I) | Generate images solely from text prompts | Supported, and prompts can be up to 64k tokens |
| Image-to-image | Image-to-Image (I2I) | Generate a new image conditioned on one or more images | Supported, fusing up to 14 reference images |
| Inpainting | Inpainting | Regenerate a specified region while leaving everything outside it unchanged | Replaced by instruction-based editing: select or describe a region, no mask required |
| Outpainting | Outpainting | Continue generating beyond the canvas | Achieved indirectly through multiple aspect ratios (incl. 21:9) and composition control, [To be filled] |
| Controllable generation | ControlNet | Control generation structure via visual signals such as Canny, Depth, Pose, Mask | Does not use ControlNet; structure is controlled via natural-language parameters such as "camera angle / focus / depth of field / composition ratio" |
| Reference image | Reference Image | An image input that constrains identity, style, structure | Hard-coded cap of 14 images, character consistency cap of 5 people |
| Random seed | Seed | Once fixed, reproduces the same image under identical parameters | No public documentation of seed locking or snapshot endpoints, [To be filled] |
| Guidance strength | CFG | How strongly the prompt constrains the generation result | Not applicable (autoregressive multimodal model, not the diffusion-guidance paradigm) |
| Low-rank adaptation | LoRA | Small-parameter fine-tuning module used to lock in character, style, clothing assets | User-side LoRA training and loading not supported |
| Image prompt adapter | IP-Adapter | Injects image-encoder features into attention to achieve "image as prompt" | Not applicable (natively multimodal, receives image input directly) |
| Zero-shot identity injection | InstantID | Transfers identity from a single reference image without fine-tuning | Not applicable; identity preservation is implemented natively via multi-reference fusion (up to 5 people) |
2.2. Nano Banana-Specific Terms
| Term | English / Abbreviation | Definition |
|---|---|---|
| Invisible watermark | SynthID | Google's pixel-level invisible digital watermark, forcibly embedded in all Google AI-generated images; can be detected by uploading an image in the Gemini App and asking whether it "was generated by Google AI"; official statement: non-removable |
| Visible watermark | Gemini Sparkle (Gemini starburst watermark) | Retained for free tier and Google AI Pro users; removable for Google AI Ultra subscribers or for images generated via Google AI Studio / Vertex AI |
| Coalition for Content Provenance and Authenticity | C2PA (Coalition for Content Provenance and Authenticity) | Content provenance and authenticity metadata standard; supported by Nano Banana Pro (per Baidu Baike, medium confidence) |
| Search grounding | Grounding with Search | Connects to the Google Search knowledge base to generate charts and infographics based on real-time information (weather, events, recipes) or specialized knowledge bases |
| Thinking model | Thinking Model | Switching to the Thinking Model in the Gemini App uses Nano Banana Pro |
| Multi-reference consistency | Multi-Reference Consistency | Fuses up to 14 reference images, maintaining facial and clothing detail of up to 5 people in complex compositions |
| Production brief | Production Brief | Officially recommended structured prompt format: subject / purpose / text / language / composition / style / aspect ratio / resolution / elements that must remain unchanged |
| Multilingual text rendering | Multilingual Text Rendering | Supports multilingual text rendering including Traditional Chinese and Simplified Chinese |
| Hand-drawn sketch to realistic | Sketch-to-Realistic | Converts a hand-drawn sketch into a realistic image |
| Enterprise-side generation | Vertex AI Generation | Images generated via Vertex AI can have the visible watermark removed, and support enterprise security mechanisms and elastic billing |
2.3. Common Terms for Clothing Change and Face Swap
Note: Nano Banana does not provide a standalone virtual try-on or face-swap product capability, but its multi-reference consistency (14 reference images / 5 people) and search grounding make it one of the most commonly used generation foundations for third-party try-on solutions (such as THG Ingenuity AI Stylist, Max Fashion's digital fitting room, and ComfyUI's official VTON workflow).
| Term | English / Abbreviation | Definition | Nano Banana relevance |
|---|---|---|---|
| Virtual try-on | VTON (Virtual Try-On) | "Wears" a target garment onto a specified person's image and produces visually credible results | Invoked as a foundation by third-party try-on solutions (ComfyUI's official VTON template uses Nano Banana Pro) |
| Garment warping | Garment Warping | Uses geometric transforms such as TPS (thin-plate spline) to align a flat garment to the body pose | Not applicable (does not expose warping control) |
| Cloth mask | Cloth Mask | Binary image of the garment region obtained by human parsing | Not applicable (instruction-based editing needs no mask) |
| Try-on diffusion | Try-on Diffusion | Fuses garment and body end-to-end with a diffusion model | Not applicable (not the diffusion paradigm) |
| Face swap | Face Swap | Replaces A's face onto B's facial position | Not provided, and explicitly restricts generating portraits of specific public figures under copyright protection |
| Face reenactment | Face Reenactment | Preserves identity while transferring expression, lip and head pose | Not provided |
| Identity preservation | Identity Preservation | To what extent the result is still "that person" | Achieved via multi-reference fusion (preserves facial detail of up to 5 people) |
3. Features
3.1. Generation and Editing Capabilities
| Capability | Supported | Notes |
|---|---|---|
| Text-to-image / image-to-image | Yes | Native multimodal input |
| Instruction-based editing | Yes | Natural-language description, no mask required |
| Multi-image fusion | Yes | Up to 14 reference images |
| Style transfer | Yes | — |
| Local editing | Yes | Select or describe a region; can adjust camera angle, focus, color grading, depth of field, day/night lighting |
| 2K / 4K output | Yes | Priced in tiers by resolution |
| Multiple aspect ratios | Yes | Including 21:9 |
| Multilingual text rendering | Yes | Including Traditional Chinese and Simplified Chinese |
| Infographics and data visualization | Yes | Combined with Search Grounding |
| Hand-drawn sketch to realistic | Yes | — |
| Search-grounded generation | Yes | Grounding with Search |
3.2. Multi-Reference Consistency and Physical-Optics Control
- Multi-reference cap: fuses 14 reference images.
- Character consistency cap: maintains facial and clothing detail of up to 5 people in complex compositions. These two numbers are hard-coded context slots and represent the most explicitly publicized context-capacity contract in this group (for comparison: Kling O1 is 10 images, Runway Gen-4 is 3, FLUX.2 is 8~10, Midjourney is stackable but has no public cap).
- Physical and optics control: can specify lighting atmosphere, lens type, depth of field, viewing angle, and composition ratio - turning photography parameters into first-class prompt fields.
3.3. Structured Production Brief
Officially, it is recommended to write prompts as a structured production brief, with fields including:
| Field | Description |
|---|---|
| subject | Subject |
| purpose | Purpose |
| text | Text to be rendered in the image |
| language | Text language |
| composition | Composition approach |
| style | Visual style |
| aspect ratio | Output ratio |
| resolution | Output dimensions |
| elements that must remain unchanged | List of consistency constraints |
This is conceptually consistent with FLUX.2's official JSON fields (subject / background / lighting / style / camera_angle / composition): upgrading context from free-form text into a verifiable field contract. The difference is that FLUX.2 is machine-readable JSON, while Nano Banana is the officially recommended natural-language authoring convention.
3.4. Known Limitations
- Cannot generate content involving violence or hate incitement.
- Cannot generate portraits of specific public figures under clear copyright protection - an identity restriction rarely seen in this group and explicitly stated by the vendor.
- Visible watermark retained by tier: the Gemini starburst watermark is retained for the free tier and Google AI Pro; removable for Ultra subscribers or images generated via AI Studio / Vertex AI.
- Pricing must be re-verified in real time: the official pricing page could not be scraped directly due to a login redirect,
[To be filled].
4. Platform Architecture
图 4-1|Nano Banana 平台架构:入口—推理—底座三层
数据来源:基于本文分析绘制的示意图。
4.1. Native Multimodal Foundation
- The foundation is the Gemini 3 Pro / 3.1 Flash multimodal LLM, natively multimodal (mixed image-text input and output), rather than an attached diffusion pipeline.
- This difference has cascading effects:
- No CFG / Steps / Sampler diffusion-paradigm parameters;
- No LoRA / ControlNet / IP-Adapter ecosystem (these are supporting mechanisms of diffusion models);
- Consistency relies on the model's own contextual understanding, not external condition injection.
4.2. Reasoning Paradigm
By default it includes reasoning steps: first understand intent and image logic, retrieve from the web when necessary (Grounding), then generate. This is the same paradigm as Jimeng Seedream 5.0's "web retrieval + CoT chain-of-thought reasoning", and is the third-generation generation-architecture consensus visible in this group - think first, then draw.
4.3. Distribution Matrix
| Side | Entry point |
|---|---|
| Consumer side | Gemini App / Web, Google Search (AI Mode), Workspace (Slides / Vids), NotebookLM, Google Ads, Flow |
| Developer side | Google AI Studio, Gemini API, Vertex AI, Google Antigravity |
Key enterprise-side difference: images generated via Vertex AI can have the visible watermark removed, and support enterprise security mechanisms and elastic billing. This means the same model has different compliance and governance attributes under the "consumer entry point" versus the "enterprise entry point" - a key point that is easy to overlook during selection.
5. Harness Design
5.1. Six-Layer Capability Overview
| Layer | Name | Nano Banana's implementation | Maturity | Evidence strength |
|---|---|---|---|---|
| L1 | Context Engineering | Hard-coded slots of 14 reference images / 5 people; 64k input tokens; Search Grounding brings external retrieval into context; structured production brief | Strong | High |
| L2 | Tools & Execution | Generation + editing + retrieval (Grounding) + multilingual rendering; exposed via Gemini API and Vertex AI; Vertex AI supports enterprise security and elastic billing | Medium | High |
| L3 | Orchestration & Control | Multi-turn conversational iteration (generate -> textual critique -> targeted revision); vendor explicitly recommends "request only one controlled change at a time" | Medium | High |
| L4 | Memory & State | Relies on Google ecosystem persistence (Drive / NotebookLM / Workspace); no first-class abstraction for "brand asset packs" | Medium | Medium |
| L5 | Evaluation & Observability | Vendor emphasizes manual verification steps - checking every word, number, face and product detail before release; no public official Eval Set | Medium | Medium |
| L6 | Governance & Security | Strongest in this group: mandatory invisible SynthID watermark + C2PA + tiered visible watermark + public-figure portrait restrictions + verifiable entry point in the Gemini App | Strongest | High |
5.2. L1 Context Engineering Layer
Nano Banana's L1 has three measurable characteristics, which are extremely rare in this group:
- An explicit capacity contract: 14 reference images, 5 people. This is not a vague statement like "supports multiple images", but a hard cap that can be written into technical specs and acceptance criteria.
- Ultra-long prompts: supports 64k input tokens, letting you stuff an entire brand guideline, product catalog, or historical-asset description into context at once.
- External knowledge grounding: Search Grounding brings Google search results into the generation context, making "generating images from real-time facts" possible (weather, events, recipes, specialized knowledge bases).
The officially recommended structured production brief (9 fields) further turns context organization from a "writing skill" into a "filling spec". Apart from FLUX.2's official JSON contract, this is the only context-organization standard given by a platform vendor in this group.
5.3. L2 Tools & Execution Layer
- Toolset: generation, editing, retrieval (Grounding), multilingual rendering.
- Exposure: Gemini API and Vertex AI. The Vertex AI side additionally provides enterprise security mechanisms and elastic billing.
- No open tool registration: external tools cannot be registered into the generation flow (by contrast, Leonardo.ai supports registering custom fine-tuned models as callable model IDs).
- Ecosystem integration: aggregated by third-party platforms and workflows (e.g., ComfyUI's official VTON workflow uses Nano Banana Pro as a generation node), showing that its API contract has good composability.
5.4. L3 Orchestration & Control Layer
- Multi-turn conversational iteration: generate -> critique with text -> targeted revision.
- Explicit vendor operational discipline: "request only one controlled change at a time". This is a rare official orchestration guideline - it raises "change granularity" to a usage norm, effectively managing context-drift risk: if you change several things at once, uncontrolled parts may also be changed by the model along the way.
- Still missing: no DAG, no sub-agent dispatch, no workflow-artifact export and version control, no interruption recovery.
Compared with Kling's "generation as conversation", the two share the same paradigm; the difference is that Kling emphasizes "no mask needed", while Nano Banana emphasizes "one change at a time".
5.5. L4 Memory & State Layer
- Relies on Google ecosystem persistence: Drive, NotebookLM, Workspace.
- No first-class abstraction for "brand asset packs" - an obvious shortcoming compared with Runway's Brand Kits (up to 3), Kling's subject creation (metered pricing), and Meitu Design Studio's creative-asset reuse (explicitly stated).
- No public documentation of Checkpoint, artifact version management, or a model-version locking endpoint.
Engineering implication: if a business needs assets such as a "locked brand character / locked logo / locked palette", Nano Banana has no native abstraction, so the business system must implement it itself and re-inject context on every request - which increases 64k-token context usage and operational cost.
5.6. L5 Evaluation & Observability Layer
- The vendor emphasizes manual verification steps: checking every word, number, face and product detail before release. This is effectively a vendor-issued acceptance checklist; although it is not an automated Eval Set, it is actionable and executable, better than the "no evaluation mechanism" status of most platforms.
- No public official Eval Set, Golden Dataset, or regression set.
- Observability basis: on the API side, pricing is per resolution and token, so cost is measurable; but the official pricing page could not be scraped directly,
[To be filled].
5.7. L6 Governance & Security Layer
This is the layer where Nano Banana is decisively ahead in this group, and it is the only closed-source platform here that can correspond to both the explicit and implicit dual-labeling requirements of China's Labeling Measures:
| Governance mechanism | Implementation | Corresponding provision of the Labeling Measures |
|---|---|---|
| Implicit labeling | SynthID forcibly embedded, officially stated as non-removable | Article 5: implicit labeling (metadata / digital watermark) shall be added |
| Metadata standard | Nano Banana Pro supports C2PA | Article 5: adding content-provenance information is encouraged |
| Explicit labeling | Visible watermark (Gemini starburst watermark), tiered by plan: retained for the free tier and Pro; removable on Ultra / AI Studio / Vertex AI | Article 4: explicit labeling shall be ensured when providing downloads and exports |
| Verifiable entry point | Detectable by uploading an image in the Gemini App and asking whether it "was generated by Google AI" | Article 6: propagation platforms verify implicit labeling |
| Identity restrictions | Does not generate portraits of specific public figures under clear copyright protection | Civil Code, Article 1019 |
Must be stated honestly: the mapping of the above mechanisms to the Labeling Measures is a comparative analysis made by this report based on public materials, not a compliance statement by Google regarding Chinese regulations. Whether Google has implemented specific metadata fields such as "service provider name or code, content ID" as required by China's Labeling Measures has no public statement, [To be filled].
At the same time, the visible watermark is removable at the Ultra / AI Studio / Vertex AI entry points, which is a reasonable design on the enterprise side (enterprise-owned assets), but it also means that relying only on platform default behavior is insufficient to satisfy the explicit-labeling obligation, and the business side still needs to build its own injection step.
5.8. Maturity Assessment
Nano Banana takes the form of "strong model + upper-middle Harness + strongest governance". Its L1 has a clear capacity contract and grounding capability, L3 has official operational discipline, L5 has an executable official acceptance checklist, and L6 is the only closed-source platform in this group with the triple mechanism of "mandatory invisible watermark + metadata standard + verifiable entry point". The shortcoming is L4 - the lack of a first-class brand-asset abstraction, which requires an additional external hosting layer in production scenarios that need "long-term, multi-batch, asset-locking".
6. Practical Cases
6.1. THG Ingenuity × Google Cloud "AI Stylist"
- Solution: a virtual try-on solution built on the Gemini Enterprise Agent Platform, already listed on the Google Cloud Marketplace.
- Quantified data (early Myprotein deployment): for users who used AI Stylist at least once, compared with the in-site average for sportswear purchases -
- UK users are close to 6× more likely to convert;
- On-site dwell time is 6.3×;
- Average order value is 2.5% higher.
- Source: Prolific North, high relevance (contains quantified data).
- Nature: the data comes from public reports on the solution provider's early deployment and is vendor-disclosed data; the source and basis should be noted when citing.
6.2. Max Fashion × Google Cloud
- Solution: Middle East retail giant Max Fashion launched an enterprise-grade digital fitting room built on the Gemini Enterprise platform, generating highly realistic real-time previews that show how garments drape and move across diverse, real body types.
- Quote (Bala Subramaniam, Senior Vice President of Omnichannel at Max Fashion): "Virtual try-on fundamentally closes the distance between online and offline experiences. Customers browsing on their phones can, for the first time, have the same shopping confidence as if standing in a physical fitting room."
- Source: fabrix.hk, citing Retail Management Middle East, medium-to-high relevance.
- Quantified data: no specific conversion or return-rate figures disclosed.
6.3. L'agence × Google AI (February 2026 New York Fashion Week)
- Content: guests can instantly visualize themselves wearing the runway looks, with instant booking supported.
- Source: Business of Fashion, medium-to-high relevance.
- Quantified data: not disclosed.
7. Summary
7.1. Strengths
- L6 governance is the strongest in the group: SynthID's mandatory invisible watermark (officially stated as non-removable) + C2PA + tiered visible watermark + public-figure portrait restrictions + an in-app verifiable entry point, forming a complete "generation—labeling—verifiable" closed loop.
- The context capacity contract is explicit: 14 reference images / 5 people / 64k tokens, which can be written into technical specifications and acceptance criteria.
- Search Grounding: brings real-time external facts into the generation context, supporting infographics and data visualization.
- The official acceptance checklist is usable: word-by-word, number-by-number, face-by-face, and product-detail-by-product-detail verification, directly convertible into an enterprise DoD.
- It is one of only two platforms with quantified commercial case studies: THG Ingenuity's conversion is close to 6× and dwell time 6.3×.
- The widest distribution matrix: from consumer-side Search / Workspace to developer-side API / Vertex AI, and the enterprise entry point (Vertex AI) is independent of the consumer entry point in compliance and billing.
7.2. Limitations and Applicability Boundaries
| Limitation | Impact |
|---|---|
| No first-class abstraction for brand assets | Must be maintained by an external system and re-injected on every request, occupying context |
| No workflow artifacts or version management | Not regression-testable, not diffable |
| No model-version locking endpoint | Output drift after model upgrades |
| Official pricing page could not be fetched directly | Cost estimation is unreliable, [To be filled] |
| Visible watermark removable at some entry points | The explicit-labeling obligation must still be built by the business side itself |
| LoRA / ControlNet not supported | Diffusion-ecosystem assets cannot be reused |
| No compliance statement for Chinese regulations | The mapping to the Labeling Measures is report analysis, not an official statement |
Applicability boundaries: suitable for scenarios that require strong compliance backing, real-time fact grounding, and consumer-level reach (Search / Workspace); not suitable for production scenarios that require long-term locking of brand assets, fully offline or privatized deployment, or precise structured control (such as FLUX.2's hex brand colors).
7.3. Selection Recommendations
- Always use Vertex AI on the enterprise side: the visible watermark is removable, enterprise security mechanisms and elastic billing are supported, and the consumer entry point differs in compliance attributes.
- Build your own brand-asset layer: consolidate brand guidelines, character references, and color palettes into an external asset library, and inject them on every request in the official 9-field production-brief format.
- Follow the "one change at a time" discipline: in multi-round iteration, request only one controlled change at a time, avoiding uncontrolled regions being modified along the way.
- Implement the acceptance checklist: write the official "word-by-word, number-by-number, face-by-face, and product-detail-by-product-detail verification" into the DoD, especially when text rendering and faces are involved.
- Pricing must be re-verified in real time: before finalizing the document and before procurement, visit the ai.google.dev pricing page to confirm current unit prices.
- Do not rely on the platform's default labeling: even with SynthID forcibly embedded, the explicit labeling and specific metadata fields required by China's Labeling Measures must still be injected by the business side itself in the export chain.
7.4. Compliance Notes
- The "Measures for Labeling AI-Generated Synthetic Content" (CAC General Office Document No. 2 of 2025): Article 4, the explicit-labeling obligation; Article 5, the implicit-labeling obligation (including the service provider's name or code and the content number); Article 6, the distribution platforms' verification obligation; Article 10, prohibition on maliciously deleting, altering, forging, or concealing labels.
- Article 1019 of the Civil Code of the People's Republic of China: no one may infringe another person's portrait right by forging or other means using information technology. Nano Banana's restriction on portraits of public figures is a platform-side guardrail and cannot substitute for the rights holder's authorization.
- The 2026-03 effective judgment of the Beijing Internet Court: identifiability + burden-of-proof shift. Even if the platform has public-figure restrictions, enterprises should still retain generation parameters, reference images, and prompts in preparation for evidentiary proceedings.
Information Gap Statement
- Official pricing page: ai.google.dev could not be fetched directly due to a login redirect; official real-time pricing
[To be filled]; the values in the table are compiled from a 2026-02-26 screenshot. - C2PA support: per the Baidu Baike statement, not the original text of Google's official documentation.
- Availability of each model in China: no official statement found on the accessibility and compliance path of the Gemini App / API in mainland China.
[To be filled] - Compliance mapping to the Labeling Measures: a comparative analysis made by the report based on public materials, not an official Google compliance statement; whether specific metadata fields such as "service provider name or code, content ID" are implemented
[To be filled]. - Seed locking, model-version locking endpoint, snapshot mechanism: no official statement found.
[To be filled] - Quota details for Google AI Plus / Pro / Ultra: only third-party sources give prices; quotas are unspecified.
[To be filled] - Release date of Nano Banana 2 (Gemini 3.1 Flash Image): search results only note "new tier" without a date.
[To be filled]
8. References
- Google AI · Gemini API pricing page (official, must be re-verified before finalizing) — Google. https://ai.google.dev/pricing
- The Rundown AI · Nano Banana Pro capabilities, pricing and entry points (third-party, medium confidence) — https://www.therundown.ai/tools/nano-banana-pro
- iKala · Nano Banana Pro features and solution introduction (incl. model comparison table, third-party) — https://ikala.ai/zh-tw/blog/ikala-ai-insight/nano-banana-pro-functions-how-to-use/
- Brand Icon Image · Google Introduces Nano Banana Pro (official launch-event retelling, third-party) — http://www.brandiconimage.com/2025/11/google-introduces-nano-banana-pro-with.html
- FoneArena · Google rolls out Nano Banana Pro powered by Gemini 3 (third-party) — https://www.fonearena.com/blog?p=469327
- Baidu Baike · Nano Banana Pro (incl. SynthID / C2PA / 14 reference images / 5 people, secondary source) — https://baike.baidu.com/item/BananaPro/67405719
- Prolific North · "THG Ingenuity predicts 600% conversion rate boost with Google AI Stylist rollout" (incl. quantified data) — https://www.prolificnorth.co.uk/news/thg-ingenuity-predicts-600-conversion-rate-boost-with-google-ai-stylist-rollout/
- Business of Fashion · "How To Drive Luxury Conversion Through AI-Powered Virtual Try-On" (incl. L'agence case) — https://www.businessoffashion.com/articles/technology/ai-virtual-try-on-luxury-dressx-report/
- FabriX · "How AI Is Solving Fashion's Fitting Room Crisis & Driving E-Commerce Sales" (incl. Max Fashion × Google Cloud) — https://fabrix.hk/blog/how-ai-is-solving-fashions-fitting-room-crisis
- ComfyUI official workflow 「虚拟角色试穿 - 四合一」 (uses Nano Banana Pro as the generation node) — Comfy Org. https://comfy.org/zh/workflows/templates_rob_fashion_shoot_vton-4in1.app/
- "Measures for Labeling AI-Generated Synthetic Content" (CAC General Office Document No. 2 of 2025) — Cyberspace Administration of China et al., effective 2025-09-01. https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm
- Article 1019 of the Civil Code of the People's Republic of China — National People's Congress, 2020. (article cited in body text, no official link)
- Economic Information Daily · "Technology is not an infringement 'shield'; the court finds AI 'face theft' this way" — Xinhua News Agency Economic Information Daily, 2026-04-17. http://dz.jjckb.cn/www/pages/webpage2009/html/2026-04/17/content_115180.htm