Manus 平台研究
1. 介绍
1.1 平台定位
Manus 由 Butterfly Effect Pte. Ltd.( Monica.im 团队)于 2025-03-06 发布,是一家注册于新加坡的公司(2024 年成立)推出的通用自主智能体产品。
按项目参数卡的统一口径,Manus 属于 Harness 的垂直产品化形态(通用场景)——与 Devin 的编码垂直相对,Manus 走的是「什么都能做一点」的通用路线:研究、写代码、做文档、生成可部署应用,全部在同一个产品内完成。
Manus 的核心设计主张是 「给 Agent 一台虚拟机,让它自己去干活」:它在云端启动隔离的 Linux 沙箱(29 个工具),Agent 在其中浏览网页、执行 Python、读写文件、调用外部 API,最终交付一个成品文件(报告 / 表格 / 幻灯片 / 可运行的 Web 应用)。
与 Devin 相比,Manus 的差异不在于技术深度,而在于场景宽度与交付物形态:Devin 交付 PR,Manus 交付「你能直接用的东西」。这决定了 Manus 的用户群更偏知识工作者(分析师、产品经理、研究者),而非仅工程师。
1.2 基本信息
| 项 | 值 |
|---|---|
| 开发商 | Butterfly Effect Pte. Ltd.(Monica.im 团队),新加坡,2024 年成立 |
| 发布时间 | 2025-03-06 |
| 形态 | 闭源商业产品 |
| 里程碑 | 上线 8 个月处理 147 万亿 token、驱动 8,000 万虚拟计算机、ARR 突破 $100M |
| 基准成绩 | GAIA:Level 1 86.5%、Level 2 70.1%、Level 3 57.7%,各档均优于 OpenAI 最佳 Agent |
| 版本 | Manus 1.5(Web App Builder);Manus 1.6(2026-05):Max 模式(双盲测试满意度 +19.2%)、Mobile Development、Design View |
| 底层模型 | 多智能体架构,基于 Anthropic Claude(早期为 Claude 3.7 Sonnet) |
| 合规 | SOC 2 Type I 与 Type II;ISO 27001 / 27701 |
| 数据路径 | 标准计划经新加坡;无正式 GDPR 认证(虽有 SOC 2 / ISO 27001) |
定价(2026-09 检索时点):
| 档位 | 价格 | 内容 |
|---|---|---|
| Free | ¥0 | 300 每日刷新 credits + 1,000 一次性;1 并发;2 定时任务 |
| Starter | $20/月 | 4,000 credits;300 每日刷新;20 并发 |
| Pro | $40/月 | 8,000 credits |
| 高档 | 约 $200/月 | 40,000 credits |
| Team | 约 $20/席位(2 人起) | — |
| 年付 | -17% | — |
1.3 资本与地缘政治事件
Manus 的发展史中有一段罕见的、直接被地缘政治干预的资本事件,这是评估其长期可用性时必须纳入的风险因子:
| 时间 | 事件 |
|---|---|
| 2025-12 | Meta 同意以约 $2B 收购 Butterfly Effect |
| 2026-04 | 中国国家发改委以国家安全为由要求撤销该交易 |
| 2026-06 | 正式分离 |
| 2026-08 | 恢复独立运营(含 Meta 时代账号数据的定时删除) |
影响评估:这一事件对 Manus 的用户产生了两个实质影响:
- 存续不确定性:产品经历过一次「险些被并入大厂又被剥离」的过山车,其长期路线图与数据政策存在变数。
- 数据处置:2026-08 恢复独立运营后,涉及 Meta 时代账号数据的定时删除计划——这对把 Manus 用于长期知识积累的用户是一个需要注意的数据连续性风险。
这段历史也构成了一个行业级信号:Agent 产业已不再是纯粹的商业竞争领域,头部产品可能直接受到国家安全审查的影响。
2. 名词解释
| 术语 | 英文 | 释义 |
|---|---|---|
| Planner Agent | Planner Agent | 把用户请求拆解为子任务 |
| Executor Agent | Executor Agent | 并行执行子任务 |
| Knowledge Agent / Verifier | Knowledge Agent | 研究、综合,并验证步骤是否真正完成(而非假设完成) |
| Linux 沙箱 | Linux Sandbox | 隔离的云端 Linux 执行环境,内置 29 个工具:browser、Python interpreter、file system、external APIs |
| 沙箱虚拟机 | Sandbox VM | Manus 为每个任务启动的云端隔离虚拟机;任务在其中独立执行,关闭浏览器后仍可继续运行 |
| Wide Research | Wide Research | 并行多智能体编排,同时处理数百个数据点/来源,每个条目获得独立一轮处理——解决「批量研究尾部质量衰减」问题 |
| Cloud Browser / Browser Operator | Browser Operator | 真实沙箱浏览器环境(非 API 模拟):导航页面、填表、用视觉模型确认每步动作;可在用户已登录会话中工作;可回放 |
| Manus's Computer | Manus's Computer | 实时执行透明界面:日志、中间工件、Agent 动作可实时查看并随时介入 |
| Mail Manus | Mail Manus | 转发邮件触发自主动作(简历筛选、会议准备、自定义工作流),可配置提示词规则 |
| Scheduled Tasks 2.0 | Scheduled Tasks | 定时监控与报告 |
| Design View | Design View | 图像 / 视频 / 3D 资产生成的交互式画布 |
| Knowledge 模块 | Knowledge Module | 持久记忆,学习用户偏好与内部文档 |
| Credit | Credit | 计费单位:复杂研究任务单次消耗 500~900 credits;无内置支出上限、无任务前成本预估;月度 credits 不结转;Agent 自身纠错重跑也消耗 |
| Max 模式 | Max Mode | Manus 1.6 引入,双盲测试满意度 +19.2% |
| GAIA | GAIA Benchmark | 通用 AI 助手基准,分 Level 1/2/3 三档难度 |
3. 功能说明
3.1 自主端到端任务执行
给定一个自然语言目标,Manus 独立规划、研究、撰写并交付成品文件。典型示例:15 分钟内产出一份 8 页带引用的市场分析。
交付物形态覆盖:Word / Excel / PPT / PDF / 代码文件 / 可部署 Web 应用 / 原生 iOS 与 Android 应用。
3.2 异步云执行
关闭浏览器后任务继续在云端运行,完成时通知。这是 Manus 相对多数对话式 Agent 的关键体验差异:它把「等待」从同步阻塞变成异步后台,符合人类对「委派」的心理模型。
3.3 Wide Research 并行研究
Wide Research 针对的是批量研究中的经典问题——尾部质量衰减:当要求 Agent 一次性处理 100 个条目时,往往前 20 个质量尚可,后 80 个越来越敷衍。
Manus 的解法是并行多智能体编排:每个条目获得独立一轮处理,而非在一个长上下文里顺序处理 100 个条目。这是用「并行 + 上下文隔离」换「质量一致性」,是 L1(上下文工程)与 L3(编排)协同解决质量问题的典型案例。
3.4 全栈应用与演示构建
单条提示即可生成可部署 Web 应用(含基础后端逻辑与登录)、原生 iOS/Android 应用。
重要风险提示:生成的全栈应用数据库设计往往幼稚,安全实践常缺失——应视为可点击原型,而非生产代码。这是所有「一句话生成应用」类产品的共同边界,Manus 也不例外。
3.5 Mail Manus 与定时任务的邮件自动化
- Mail Manus:转发邮件即可触发自主动作(简历筛选、会议准备、自定义工作流),可配置提示词规则。
- Scheduled Tasks 2.0:定时监控与报告,把一次性任务变成持续运行的自动化。
3.6 集成生态
Slack、WhatsApp、Telegram、Gmail、Google Calendar、Notion、Google Drive、GitHub、Stripe、Zapier。
Manus 1.6 新增 Mobile Development 与 Design View(图像 / 视频 / 3D 资产生成的交互式画布)。
4. 平台架构
4.1 多智能体架构
图 4-1|Manus 平台架构:多智能体编排 × 沙箱执行 × 实时透明
数据来源:基于本文分析绘制的示意图。
中央编排器(拆解用户请求)
├─ Planner Agent —— 拆子任务
├─ Executor Agent —— 并行执行子任务
└─ Knowledge / Verifier Agent —— 研究、综合并验证是否真正完成 底层模型为 Anthropic Claude(早期为 Claude 3.7 Sonnet)。
架构中最值得注意的设计是 Verifier 角色的独立存在:它不负责做事,只负责验证「步骤是否真正完成,而非假设完成」。这直击自主智能体最常见的失败模式——Agent 宣称完成了但实际没有验证。
4.2 执行环境:沙箱虚拟机
- 云端 VM + 隔离 Linux 沙箱,内置 29 个工具(browser、Python interpreter、file system、external APIs)。
- Cloud Browser:真实浏览器环境(非 API 模拟),导航页面、填表、用视觉模型确认每步动作,且可在用户已登录会话中工作,过程可回放。
- 桌面应用额外提供本地 "My Computer" 访问。
沙箱虚拟机是 Manus 产品能力的基础设施:它让 Agent 拥有「一台完整的计算机」而非一组受限 API,这是其工具表达力的来源,同时也是其安全风险与成本风险的来源。
4.3 透明层
Manus's Computer 提供实时执行透明界面:日志、中间工件、Agent 动作可实时查看,用户可随时介入。这在闭源 Agent 产品中相当罕见——多数闭源产品只给结果不给过程。
4.4 商业模式
freemium + credit 消耗制。Credit 是理解 Manus 经济性的核心:复杂研究任务单次消耗 500~900 credits,而月度 credits 不结转、Agent 自身纠错重跑也消耗、无任务前成本预估、无内置支出上限。
5. Harness 设计
5.1 六层能力总览
| 层 | 名称 | 评级 | 一句话判断 |
|---|---|---|---|
| L1 | 上下文工程层 | 中(不透明) | Knowledge 模块持久记忆 + 多智能体上下文拆分;压缩/缓存/优先级机制未公开 |
| L2 | 工具与执行层 | 强 | 云端 VM + 隔离 Linux 沙箱(29 工具)+ Cloud Browser 真实浏览器 + 视觉确认 |
| L3 | 编排与控制层 | 强 | Planner / Executor / Verifier 三层 + Wide Research 并行子智能体 + 异步长时运行 |
| L4 | 记忆与状态层 | 中 | Knowledge 模块 + 异步会话;跨任务长期记忆与工件管理细节未公开 |
| L5 | 评估与观测层 | 中 | Manus's Computer 实时透明 + 事后回放;无公开评估集 / 回归评测能力 |
| L6 | 治理与安全层 | 中 | 审计日志(Team+)、SSO(Team)、SOC 2 / ISO;但无 GDPR 认证,无支出上限 |
5.2 L1 上下文工程层
评级:中(不透明)。
可确认的机制:
- Knowledge 模块:持久记忆,学习用户偏好与内部文档。
- 多智能体上下文拆分:这是 Manus 解决上下文问题的核心工程手段——与其压缩长上下文,不如让每个子 Agent 只处理一小块。
不可确认的部分:压缩、缓存与优先级排序的具体机制未公开。Wide Research 的「每个条目独立一轮」设计暗示了 Manus 的工程取向——用上下文隔离替代上下文压缩,但这属于行为推测,官方未确认。
5.3 L2 工具与执行层
评级:强。
- 云端 VM + 隔离 Linux 沙箱,29 个工具,覆盖 browser、Python interpreter、文件系统、外部 API。
- Cloud Browser:真实浏览器自动化(非 API 模拟),并用视觉模型确认每一步动作——这比纯 DOM 操作的可靠性更高。
- 可生成并部署应用(Web / iOS / Android)。
这一层是 Manus 与 Devin 并列为「强」的地方:两者都给了 Agent 一台完整的计算机。差别在于 Manus 的计算机更偏「办公与浏览」,Devin 的更偏「工程与代码」。
5.4 L3 编排与控制层
评级:强。
- Planner / Executor / Verifier 三层架构:Verifier 的独立存在是设计亮点。
- Wide Research 并行子智能体:用并行 + 上下文隔离解决批量研究的尾部质量衰减。
- 异步长时运行:关浏览器继续跑,完成时通知。
- 可随时介入:Manus's Computer 允许用户中途打断与纠偏。
张力一(灵活性 ↔ 可预测性)在 Manus 上同样是「倒向可预测」:编排逻辑完全封装在产品内部,用户无法定义自己的工作流图,只能靠自然语言描述目标。作为面向知识工作者的产品,这是合理取舍;但这意味着同一目标的重复执行,路径与结果都可能不同,无法提供流程级的可重现保证。
5.5 L4 记忆与状态层
评级:中。
- Knowledge 模块:持久记忆,学习用户偏好与内部文档。
- 异步会话:关浏览器后任务继续,会话状态在云端保持。
跨任务长期记忆与工件管理细节未公开。用户可以观察到「它似乎记得我的偏好」,但无法查看、编辑或导出这部分记忆——对于企业用户,这构成一个数据主权问题。
5.6 L5 评估与观测层
评级:中。
- Manus's Computer 实时透明:过程可观察、可介入、可回放。这在闭源产品中已是难得。
- GAIA 基准成绩:L1 86.5% / L2 70.1% / L3 57.7%,是 Manus 唯一可引用的横向可比数据。
无公开评估集 / 回归评测能力。用户无法建立自己的测试集,也无法量化「Manus 在我的任务类型上是否进步了」。与阿里云百炼的 OpenJudge(50+ judge)相比,这是明显差距。
5.7 L6 治理与安全层
评级:中。
| 机制 | 说明 | 缺口 |
|---|---|---|
| 审计日志 | Team 及以上提供 | 免费档无 |
| SSO | Team 档提供 | 免费档无 |
| SOC 2 Type I / II | 有 | — |
| ISO 27001 / 27701 | 有 | — |
| GDPR 认证 | 无正式认证 | 数据经新加坡,EU 个人数据场景受限 |
| 支出上限 | 无内置支出上限 | credit 消耗不可封顶 |
社区层面的安全建议(值得所有用户采纳):
- 不要直接共享生产账号——应创建专用受限凭据给 Manus 使用。因为 Cloud Browser 可在用户已登录会话中工作,共享真实账号等于把自己的完整权限交给 Agent。
- credit 消耗需设月度预算告警——由于产品侧无上限,只能在使用侧自建监控。
5.8 三条内在张力在 Manus 上的投影
| 张力 | 在 Manus 上的具体表现 | 平台给出的答案 | 剩余风险 |
|---|---|---|---|
| 灵活性 ↔ 可预测性 | 完全产品化,用户无法定义编排 | 用「自然语言给目标、产品给结果」换易用性 | 同一目标重复执行路径与结果可能不同 |
| 开放性 ↔ 治理 | Cloud Browser 可在已登录会话中工作,能力最强也风险最大 | 审计日志 + SSO + SOC 2 / ISO | 无 GDPR 认证;社区明确建议用专用受限凭据而非生产账号 |
| 成本 ↔ 深度 | 深度研究任务单次 500~900 credits,纠错重跑也计费 | 分档 credits 包 | 无支出上限、无任务前预估、月度不结转——成本完全不可预测 |
成本张力是 Manus 用户最主要的抱怨来源,也是本平台最需要在选型阶段就设计好对策的部分。可行的工程对策包括:把大任务拆成小任务分批执行(获得多次成本反馈点)、为团队设置共享 credit 池的告警、避免在高峰期执行长任务(高峰期服务不稳定 + 重试消耗)。
6. 实际案例
未检索到带客户名与量化 ROI 的企业落地案例。 检索到的材料均为产品能力描述、基准成绩与个人/团队使用评测。
可用的定性论据:
| 论据 | 内容 | 来源性质 |
|---|---|---|
| 用户时间节省 | 用户报告每周节省约 8 小时 | 第三方聚合评分口径, |
| 跨平台评分 | 4.0 / 5(488,125 条评价) | 聚合口径, |
| 平台规模 | 上线 8 个月处理 147 万亿 token、驱动 8,000 万虚拟计算机、ARR 突破 $100M | 厂商/媒体口径 |
| 基准成绩 | GAIA L1 86.5% / L2 70.1% / L3 57.7% | 厂商口径(但为标准化基准,可比性相对高) |
诚实标注:上表中无任何一项构成「企业落地案例」。任何引用 Manus 业务价值的论述都必须标注「未检索到公开量化数据」。
7. 总结
7.1 优点
- GAIA 基准领先:L1 86.5% / L2 70.1% / L3 57.7%,各档均优于 OpenAI 最佳 Agent。
- 通用自主执行覆盖面最广:研究 / 代码 / 文档 / 应用生成一体化,一个产品覆盖多种知识工作。
- Wide Research 并行研究质量高:精准解决批量研究的尾部质量衰减问题。
- 实时透明可审计:Manus's Computer 在闭源产品中罕见。
- Verifier 角色独立:验证「真正完成」而非「假设完成」,直击自主 Agent 核心失败模式。
- 集成生态丰富:Slack / Notion / Gmail / Calendar / GitHub / Stripe / Zapier。
- 异步执行:关浏览器继续跑,符合委派心理模型。
7.2 缺点
- credit 消耗高度不可预测(最大痛点):500~900/复杂任务;无预估、无上限、不结转;纠错重跑也计费。
- 无正式 GDPR 认证:数据经新加坡,EU / 日本个人数据合规场景受限。
- 高峰期服务不稳定,且重试同样消耗 credits。
- 生成的全栈应用非生产级:数据库设计幼稚,安全实践缺失。
- L1 / L4 / L5 不透明:上下文机制、长期记忆、评估能力均不可验证或干预。
- 资本与地缘政治事件带来存续不确定性:Meta 收购被撤销、恢复独立运营、Meta 时代账号数据定时删除。
- 未检索到带客户名与 ROI 的企业落地案例。
7.3 适用边界
适合:
- 自由研究者 / 市场分析师产出客户级报告(带引用)。
- 产品团队快速原型验证(明确视为可点击原型)。
- 小机构跑内容或数据管道。
- 批量研究类任务(Wide Research 的最佳适配区)。
- 个人知识工作者的日常委派(邮件、日程、文档)。
不适合:
- 需要精确可靠输出、时间敏感的专业工作流。
- EU / 日本个人数据合规严格的场景(无 GDPR 认证)。
- 需要严格预算封顶的团队。
- 需要自建评估集与回归验证的组织。
- 把生成物直接投入生产的工程场景。
7.4 选型建议
| 如果你的首要约束是 | Manus 是否合适 | 理由 |
|---|---|---|
| 批量研究 / 报告生成 | 强合适 | Wide Research 是差异化能力 |
| 快速原型验证 | 强合适 | 覆盖面广,交付物可直接演示 |
| 个人知识工作委派 | 合适 | 异步执行 + 集成生态 |
| 成本严格封顶 | 不合适 | 无支出上限,credit 不可预测 |
| EU / 日本个人数据 | 不合适 | 无 GDPR 认证 |
| 生产级输出 | 不合适 | 生成物需视为原型 |
| 长期数据连续性 | 需谨慎 | 资本事件导致数据政策变动 |
一句话结论:Manus 是通用自主智能体中覆盖面最广、基准成绩最好的一个,Wide Research 与 Verifier 设计都体现了真实的工程洞察;但它的成本模型(无上限、无预估、不结转)与合规缺口(无 GDPR)使其难以进入受监管的企业核心流程。它是极强的个人生产力工具,而非企业级平台。
信息缺口声明
- 企业落地案例:3 轮检索后未检索到带客户名与量化 ROI 的企业落地案例。
- L1 / L4 / L5 层机制:Manus 为闭源产品,上下文压缩/缓存/优先级机制、跨任务长期记忆与工件管理细节、评估集与回归评测能力均未公开。报告中标注为「机制未公开」,不得推测填补。
- 29 个工具的完整清单:未检索到官方公布的工具明细,仅知覆盖 browser、Python interpreter、file system、external APIs 四类。
- 恢复独立运营后的产品与定价变化:2026-08 恢复独立运营后是否有定价或功能调整,未检索到确认信息,标注 。
- 用户时间节省数据:「每周节省约 8 小时」与「4.0/5(488,125 条评价)」均为第三方聚合评分口径,非独立研究,标注 。
- 版本号时点敏感:Manus 1.6(2026-05)为 2026-09 检索时点版本。
- Transparency / AI 生成内容标注:未检索到 Manus 官方关于透明度说明或 AI 生成内容标注政策的公开资料。
- Meta 时代账号数据删除的具体范围与时间表:仅知「恢复独立运营(含 Meta 时代账号数据的定时删除)」,具体范围未公开,标注 。
8. 参考资料
- Manus AI Complete Guide 2026 — AIpedia。https://en.ai-pedias.com/blog/manus-ai-complete-guide-2026
- Manus Review 2026 — AI:PRODUCTIVITY。https://aiproductivity.ai/tools/manus
- Manus AI — AI Cloud Base。https://aicloudbase.com/tool/manus-ai
- Manus Ai Review 2026: Features, Pricing, and Verdict — AI Journal Now。https://aijournalnow.com/manus-ai-review/
- Is Manus Free? Plans, Limits & Pricing (2026) — HokAI。https://hokai.io/hub/agents/manus
- 2026 企业智能体开发平台全景评测:八大主流平台横向对比 — 稀土掘金。https://juejin.cn/post/7654244323158016038
- 2026年AI智能体平台全维度横评:从"养龙虾"到企业级部署 — CSDN。https://blog.csdn.net/weixin_56622231/article/details/159515126
- AgentScope 2.0: From Transparent Development to System Engineering(沙箱与权限设计对照参考)。https://java.agentscope.io/v2/en/blogs/agentscope-v2-release.html
- Claude Agent SDK — Agent Patterns Catalog(执行透明与权限设计对照参考)。https://www.agentpatternscatalog.org/compositions/claude-agent-sdk
Manus Platform Research
1. Introduction
1.1 Platform Positioning
Manus was released on 2025-03-06 by Butterfly Effect Pte. Ltd. (the Monica.im team), a company incorporated in Singapore (founded in 2024), as a general-purpose autonomous agent product.
Following the unified framing of the project parameter card, Manus is a vertically productized form of Harness (general scenario) — in contrast to Devin's coding vertical, Manus follows the general-purpose route of "doing a bit of everything": research, writing code, making documents, and generating deployable applications, all completed within a single product.
Manus's core design proposition is "give the Agent a virtual machine and let it do the work itself": it launches an isolated Linux sandbox in the cloud (29 tools), within which the Agent browses web pages, executes Python, reads and writes files, and calls external APIs, ultimately delivering a finished artifact (a report / spreadsheet / slide deck / runnable Web application).
Compared with Devin, Manus's difference lies not in technical depth but in scenario breadth and deliverable form: Devin delivers PRs, while Manus delivers "things you can use directly." This determines that Manus's user base skews toward knowledge workers (analysts, product managers, researchers) rather than engineers only.
1.2 Basic Information
| Item | Value |
|---|---|
| Developer | Butterfly Effect Pte. Ltd. (Monica.im team), Singapore, founded 2024 |
| Release date | 2025-03-06 |
| Form | Closed-source commercial product |
| Milestones | Processed 147 trillion tokens, drove 80 million virtual computers, and surpassed $100M ARR in 8 months since launch |
| Benchmark results | GAIA: Level 1 86.5%, Level 2 70.1%, Level 3 57.7%, all tiers better than OpenAI's best Agent |
| Version | Manus 1.5 (Web App Builder); Manus 1.6 (2026-05): Max mode (double-blind test satisfaction +19.2%), Mobile Development, Design View |
| Underlying model | Multi-agent architecture based on Anthropic Claude (earlier Claude 3.7 Sonnet) |
| Compliance | SOC 2 Type I and Type II; ISO 27001 / 27701 |
| Data path | Standard plans via Singapore; no formal GDPR certification (despite SOC 2 / ISO 27001) |
Pricing (as of the 2026-09 retrieval point):
| Tier | Price | Includes |
|---|---|---|
| Free | ¥0 | 300 daily-refreshed credits + 1,000 one-time; 1 concurrency; 2 scheduled tasks |
| Starter | $20/month | 4,000 credits; 300 daily refresh; 20 concurrency |
| Pro | $40/month | 8,000 credits |
| High-end | ~$200/month | 40,000 credits |
| Team | ~$20/seat (2 people min.) | — |
| Annual | -17% | — |
1.3 Capital and Geopolitical Events
Manus's development history contains a rare capital event directly intervened in by geopolitics — a risk factor that must be included when assessing its long-term viability:
| Date | Event |
|---|---|
| 2025-12 | Meta agreed to acquire Butterfly Effect for ~$2B |
| 2026-04 | China's National Development and Reform Commission demanded the transaction be rescinded on national security grounds |
| 2026-06 | Formal separation |
| 2026-08 | Independent operation resumed (including the scheduled deletion of account data from the Meta era) |
Impact assessment: this event had two substantive effects on Manus's users:
- Continuity uncertainty: the product has gone through a rollercoaster of "nearly being folded into a big company and then spun back out," leaving its long-term roadmap and data policy in flux.
- Data disposition: after independent operation resumed in 2026-08, there is a scheduled deletion plan for account data from the Meta era — a data-continuity risk that users relying on Manus for long-term knowledge accumulation should note.
This history also constitutes an industry-level signal: the Agent industry is no longer a purely commercial competitive arena — leading products may be directly subject to national security review.
2. Glossary
| Term | English | Definition |
|---|---|---|
| Planner Agent | Planner Agent | Decomposes the user's request into subtasks |
| Executor Agent | Executor Agent | Executes subtasks in parallel |
| Knowledge Agent / Verifier | Knowledge Agent | Researches, synthesizes, and verifies that steps were actually completed (rather than assumed to be) |
| Linux sandbox | Linux Sandbox | Isolated cloud Linux execution environment with 29 built-in tools: browser, Python interpreter, file system, external APIs |
| Sandbox VM | Sandbox VM | The isolated cloud VM that Manus launches for each task; the task runs independently inside it and can continue running even after the browser is closed |
| Wide Research | Wide Research | Parallel multi-agent orchestration that processes hundreds of data points/sources simultaneously, giving each entry its own independent round of processing — solving the "tail-quality degradation" problem of batch research |
| Cloud Browser / Browser Operator | Browser Operator | A real sandbox browser environment (not API simulation): navigates pages, fills forms, and confirms each step with a vision model; can work within the user's logged-in session; replayable |
| Manus's Computer | Manus's Computer | Real-time execution transparency interface: logs, intermediate artifacts, and Agent actions can be viewed in real time and intervened in at any time |
| Mail Manus | Mail Manus | Forwarding an email triggers autonomous actions (resume screening, meeting preparation, custom workflows), with configurable prompt rules |
| Scheduled Tasks 2.0 | Scheduled Tasks | Scheduled monitoring and reporting |
| Design View | Design View | Interactive canvas for image / video / 3D asset generation |
| Knowledge Module | Knowledge Module | Persistent memory that learns user preferences and internal documents |
| Credit | Credit | Billing unit: a complex research task consumes 500~900 credits; no built-in spending cap, no pre-task cost estimate; monthly credits do not roll over; the Agent's own error-correction reruns also consume credits |
| Max Mode | Max Mode | Introduced in Manus 1.6, double-blind test satisfaction +19.2% |
| GAIA | GAIA Benchmark | General-purpose AI assistant benchmark divided into three difficulty tiers: Level 1/2/3 |
3. Feature Description
3.1 Autonomous End-to-End Task Execution
Given a natural-language goal, Manus independently plans, researches, writes, and delivers a finished artifact. Typical example: producing an 8-page market analysis with citations within 15 minutes.
Deliverable forms cover: Word / Excel / PPT / PDF / code files / deployable Web applications / native iOS and Android applications.
3.2 Asynchronous Cloud Execution
Tasks keep running in the cloud after the browser is closed, with notification upon completion. This is the key experiential difference between Manus and most conversational Agents: it turns "waiting" from a synchronous block into an asynchronous background process, matching the human mental model of "delegation."
3.3 Wide Research Parallel Research
Wide Research targets a classic problem in batch research — tail-quality degradation: when an Agent is asked to process 100 entries at once, the first 20 are often acceptable in quality while the last 80 become increasingly perfunctory.
Manus's solution is parallel multi-agent orchestration: each entry gets its own independent round of processing rather than processing 100 entries sequentially in one long context. This trades "parallelism + context isolation" for "quality consistency," and is a typical case of L1 (context engineering) and L3 (orchestration) jointly solving a quality problem.
3.4 Full-Stack Application & Demo Building
A single prompt can generate a deployable Web application (including basic backend logic and login) and native iOS/Android applications.
Important risk notice: the generated full-stack applications often have naive database designs and frequently missing security practices — they should be treated as clickable prototypes, not production code. This is the common boundary of all "one-sentence app generation" products, and Manus is no exception.
3.5 Mail Manus and Email Automation via Scheduled Tasks
- Mail Manus: forwarding an email triggers autonomous actions (resume screening, meeting preparation, custom workflows), with configurable prompt rules.
- Scheduled Tasks 2.0: scheduled monitoring and reporting, turning one-off tasks into continuously running automation.
3.6 Integration Ecosystem
Slack, WhatsApp, Telegram, Gmail, Google Calendar, Notion, Google Drive, GitHub, Stripe, Zapier.
Manus 1.6 adds Mobile Development and Design View (an interactive canvas for image / video / 3D asset generation).
4. Platform Architecture
4.1 Multi-Agent Architecture
图 4-1|Manus 平台架构:多智能体编排 × 沙箱执行 × 实时透明
数据来源:基于本文分析绘制的示意图。
中央编排器(拆解用户请求)
├─ Planner Agent —— 拆子任务
├─ Executor Agent —— 并行执行子任务
└─ Knowledge / Verifier Agent —— 研究、综合并验证是否真正完成 The underlying model is Anthropic Claude (earlier Claude 3.7 Sonnet).
The most noteworthy design in the architecture is the independent existence of the Verifier role: it does not do the work but only verifies that "steps were actually completed, not assumed to be." This directly targets the most common failure mode of autonomous agents — the Agent claims it finished but never actually verified.
4.2 Execution Environment: Sandbox VM
- Cloud VM + isolated Linux sandbox with 29 built-in tools (browser, Python interpreter, file system, external APIs).
- Cloud Browser: a real browser environment (not API simulation) that navigates pages, fills forms, and confirms each step with a vision model, and can work within the user's logged-in session; the process is replayable.
- The desktop app additionally provides local "My Computer" access.
The sandbox VM is the infrastructure of Manus's product capabilities: it gives the Agent "a complete computer" instead of a limited set of APIs — the source of its tool expressiveness, and at the same time the source of its security and cost risks.
4.3 Transparency Layer
Manus's Computer provides a real-time execution transparency interface: logs, intermediate artifacts, and Agent actions can be viewed in real time, and the user can intervene at any time. This is fairly rare among closed-source Agent products — most closed products give only results, not process.
4.4 Business Model
Freemium + a credit-consumption model. Credit is the core to understanding Manus's economics: a complex research task consumes 500~900 credits per run, monthly credits do not roll over, the Agent's own error-correction reruns also consume credits, there is no pre-task cost estimate, and there is no built-in spending cap.
5. Harness Design
5.1 Six-Layer Capability Overview
| Layer | Name | Rating | One-line assessment |
|---|---|---|---|
| L1 | Context Engineering Layer | Medium (opaque) | Knowledge Module persistent memory + multi-agent context splitting; compression/caching/prioritization mechanisms not disclosed |
| L2 | Tools and Execution Layer | Strong | Cloud VM + isolated Linux sandbox (29 tools) + Cloud Browser real browser + vision confirmation |
| L3 | Orchestration and Control Layer | Strong | Planner / Executor / Verifier three tiers + Wide Research parallel sub-agents + asynchronous long-running execution |
| L4 | Memory and State Layer | Medium | Knowledge Module + asynchronous sessions; cross-task long-term memory and artifact management details not disclosed |
| L5 | Evaluation and Observability Layer | Medium | Manus's Computer real-time transparency + post-hoc replay; no public evaluation sets / regression evaluation capability |
| L6 | Governance and Security Layer | Medium | Audit logs (Team+), SSO (Team), SOC 2 / ISO; but no GDPR certification, no spending cap |
5.2 L1 Context Engineering Layer
Rating: Medium (opaque).
Mechanisms that can be confirmed:
- Knowledge Module: persistent memory that learns user preferences and internal documents.
- Multi-agent context splitting: this is the core engineering means by which Manus addresses the context problem — rather than compressing long contexts, each sub-Agent only handles a small piece.
What cannot be confirmed: the specific mechanisms of compression, caching, and prioritization are not disclosed. Wide Research's "each entry in its own round" design hints at Manus's engineering orientation — replacing context compression with context isolation — but this is a behavioral inference, not officially confirmed.
5.3 L2 Tools and Execution Layer
Rating: Strong.
- Cloud VM + isolated Linux sandbox with 29 tools covering browser, Python interpreter, file system, and external APIs.
- Cloud Browser: real browser automation (not API simulation) that confirms each step with a vision model — more reliable than pure DOM operations.
- Can generate and deploy applications (Web / iOS / Android).
This layer is where Manus ties with Devin as "Strong": both give the Agent a complete computer. The difference is that Manus's computer skews more toward "office work and browsing," while Devin's skews more toward "engineering and code."
5.4 L3 Orchestration and Control Layer
Rating: Strong.
- Planner / Executor / Verifier three-tier architecture: the independent existence of the Verifier is a design highlight.
- Wide Research parallel sub-agents: use parallelism + context isolation to solve the tail-quality degradation of batch research.
- Asynchronous long-running execution: keeps running after the browser is closed, notifying on completion.
- Intervention at any time: Manus's Computer allows users to interrupt and correct course mid-task.
Tension one (flexibility ↔ predictability) also "leans toward predictability" on Manus: orchestration logic is fully encapsulated inside the product, and users cannot define their own workflow graphs — they can only describe goals in natural language. As a product for knowledge workers, this is a reasonable trade-off; but it also means repeated execution of the same goal may produce different paths and results, offering no process-level reproducibility guarantee.
5.5 L4 Memory and State Layer
Rating: Medium.
- Knowledge Module: persistent memory that learns user preferences and internal documents.
- Asynchronous sessions: after the browser is closed, tasks continue and session state is maintained in the cloud.
Cross-task long-term memory and artifact management details are not disclosed. Users can observe that "it seems to remember my preferences," but cannot view, edit, or export this part of the memory — which constitutes a data sovereignty issue for enterprise users.
5.6 L5 Evaluation and Observability Layer
Rating: Medium.
- Manus's Computer real-time transparency: the process is observable, interruptible, and replayable. This is already rare among closed-source products.
- GAIA benchmark results: L1 86.5% / L2 70.1% / L3 57.7%, the only horizontally comparable data Manus can cite.
No public evaluation set / regression evaluation capability. Users cannot build their own test sets, nor quantify whether "Manus has improved on my task types." Compared with Alibaba Cloud Bailian's OpenJudge (50+ judges), this is a clear gap.
5.7 L6 Governance and Security Layer
Rating: Medium.
| Mechanism | Description | Gap |
|---|---|---|
| Audit logs | Provided for Team and above | Not available on free tier |
| SSO | Provided on Team tier | Not available on free tier |
| SOC 2 Type I / II | Yes | — |
| ISO 27001 / 27701 | Yes | — |
| GDPR certification | No formal certification | Data routed via Singapore; EU personal-data scenarios limited |
| Spending cap | No built-in spending cap | Credit consumption cannot be capped |
Community-level security recommendations (worth adopting for all users):
- Do not share production accounts directly — create dedicated, restricted credentials for Manus. Because Cloud Browser can work within the user's logged-in session, sharing a real account is equivalent to handing your full permissions to the Agent.
- Set monthly budget alerts for credit consumption — since the product has no cap, monitoring must be self-built on the usage side.
5.8 The Projection of Three Inherent Tensions on Manus
| Tension | How it manifests on Manus | The platform's answer | Remaining risk |
|---|---|---|---|
| Flexibility ↔ Predictability | Fully productized; users cannot define orchestration | Trades usability for "describe the goal in natural language, the product delivers the result" | Repeated execution of the same goal may produce different paths and results |
| Openness ↔ Governance | Cloud Browser can work within a logged-in session — the most capable and the riskiest | Audit logs + SSO + SOC 2 / ISO | No GDPR certification; the community explicitly recommends dedicated restricted credentials over production accounts |
| Cost ↔ Depth | Deep research tasks cost 500~900 credits per run; error-correction reruns are also billed | Tiered credit packages | No spending cap, no pre-task estimate, no monthly rollover — costs are completely unpredictable |
The cost tension is the primary source of complaint among Manus users, and the part this platform most needs to design countermeasures for at the selection stage. Feasible engineering countermeasures include: splitting large tasks into smaller batched tasks (obtaining multiple cost-feedback points), setting alerts on a shared credit pool for the team, and avoiding long tasks during peak hours (peak-hour service instability + retry consumption).
6. Case Studies
No enterprise deployment case with a named customer and quantified ROI was found. The retrieved materials are all product-capability descriptions, benchmark results, and individual/team usage reviews.
Available qualitative evidence:
| Evidence | Content | Source nature |
|---|---|---|
| User time saved | Users report saving about 8 hours per week | Third-party aggregate rating source |
| Cross-platform rating | 4.0 / 5 (488,125 reviews) | Aggregate source |
| Platform scale | Processed 147 trillion tokens, drove 80 million virtual computers, and surpassed $100M ARR in 8 months since launch | Vendor/media source |
| Benchmark results | GAIA L1 86.5% / L2 70.1% / L3 57.7% | Vendor source (but a standardized benchmark, relatively comparable) |
Honest annotation: none of the items in the table above constitutes an "enterprise deployment case." Any claim citing Manus's business value must be annotated with "no public quantified data was found."
7. Summary
7.1 Strengths
- Leading GAIA benchmarks: L1 86.5% / L2 70.1% / L3 57.7%, all tiers better than OpenAI's best Agent.
- Broadest coverage of general autonomous execution: research / code / documents / application generation integrated into one product covering multiple kinds of knowledge work.
- High-quality Wide Research parallel research: precisely solves the tail-quality degradation problem of batch research.
- Real-time transparency and auditability: Manus's Computer is rare among closed-source products.
- Independent Verifier role: verifies "actually completed" rather than "assumed completed," directly targeting the core failure mode of autonomous agents.
- Rich integration ecosystem: Slack / Notion / Gmail / Calendar / GitHub / Stripe / Zapier.
- Asynchronous execution: keeps running with the browser closed, matching the delegation mental model.
7.2 Weaknesses
- Highly unpredictable credit consumption (biggest pain point): 500~900 per complex task; no estimate, no cap, no rollover; error-correction reruns are also billed.
- No formal GDPR certification: data routed via Singapore, limiting compliance scenarios for EU / Japanese personal data.
- Service instability during peak hours, and retries consume credits too.
- Generated full-stack applications are not production-grade: naive database design and missing security practices.
- L1 / L4 / L5 are opaque: context mechanisms, long-term memory, and evaluation capabilities cannot be verified or intervened in.
- Capital and geopolitical events create continuity uncertainty: the Meta acquisition was rescinded, independent operation resumed, and Meta-era account data is scheduled for deletion.
- No enterprise deployment case with a named customer and ROI was found.
7.3 Applicability Boundaries
Well suited to:
- Independent researchers / market analysts producing client-grade reports (with citations).
- Product teams rapidly validating prototypes (explicitly treating them as clickable prototypes).
- Small organizations running content or data pipelines.
- Batch-research tasks (Wide Research's best-fit zone).
- Daily delegation for individual knowledge workers (email, scheduling, documents).
Not suited to:
- Time-sensitive professional workflows requiring precise, reliable output.
- Scenarios with strict EU / Japanese personal-data compliance (no GDPR certification).
- Teams that need a strict budget cap.
- Organizations that need self-built evaluation sets and regression validation.
- Engineering scenarios that put generated output directly into production.
7.4 Selection Recommendations
| If your primary constraint is | Is Manus suitable? | Reason |
|---|---|---|
| Batch research / report generation | Strongly suitable | Wide Research is a differentiating capability |
| Rapid prototype validation | Strongly suitable | Broad coverage; deliverables can be demoed directly |
| Individual knowledge-work delegation | Suitable | Asynchronous execution + integration ecosystem |
| Strict cost cap | Not suitable | No spending cap; credit usage unpredictable |
| EU / Japanese personal data | Not suitable | No GDPR certification |
| Production-grade output | Not suitable | Generated output must be treated as a prototype |
| Long-term data continuity | Needs caution | Capital events cause data-policy changes |
One-line conclusion: Manus is the general-purpose autonomous agent with the broadest coverage and the best benchmark results; the Wide Research and Verifier designs both reflect genuine engineering insight. But its cost model (no cap, no estimate, no rollover) and compliance gaps (no GDPR) make it difficult to enter regulated enterprise core processes. It is a very strong personal productivity tool, not an enterprise-grade platform.
Information-Gap Statement
- Enterprise deployment cases: after 3 rounds of retrieval, no enterprise deployment case with a named customer and quantified ROI was found.
- L1 / L4 / L5 layer mechanisms: Manus is a closed-source product — the context compression/caching/prioritization mechanisms, cross-task long-term memory and artifact-management details, and evaluation-set and regression-evaluation capabilities are all not disclosed. The report marks these as "mechanisms not disclosed" and must not be filled in by speculation.
- The complete list of the 29 tools: no officially published tool breakdown was found; only that they cover four categories: browser, Python interpreter, file system, external APIs.
- Product and pricing changes after resuming independent operation: whether there were pricing or feature adjustments after independent operation resumed in 2026-08 has not been confirmed; marked as
[To be verified]. - User time-saving data: "about 8 hours saved per week" and "4.0/5 (488,125 reviews)" both come from third-party aggregate rating sources, not independent research; marked as
[To be verified]. - Version number is time-sensitive: Manus 1.6 (2026-05) is the current version at the 2026-09 retrieval point.
- Transparency / AI-generated-content labeling: no public documents from Manus on transparency statements or AI-generated-content labeling policy were found.
- The specific scope and schedule of Meta-era account-data deletion: only "independent operation resumed (including the scheduled deletion of Meta-era account data)" is known; the specific scope is not disclosed; marked as
[To be verified].
8. References
- Manus AI Complete Guide 2026 — AIpedia. https://en.ai-pedias.com/blog/manus-ai-complete-guide-2026
- Manus Review 2026 — AI:PRODUCTIVITY. https://aiproductivity.ai/tools/manus
- Manus AI — AI Cloud Base. https://aicloudbase.com/tool/manus-ai
- Manus Ai Review 2026: Features, Pricing, and Verdict — AI Journal Now. https://aijournalnow.com/manus-ai-review/
- Is Manus Free? Plans, Limits & Pricing (2026) — HokAI. https://hokai.io/hub/agents/manus
- 2026 Enterprise Agent Development Platform Panoramic Review: A Horizontal Comparison of Eight Major Platforms — Juejin. https://juejin.cn/post/7654244323158016038
- 2026 Full-Dimension Review of AI Agent Platforms: From "Farming Lobsters" to Enterprise-Grade Deployment — CSDN. https://blog.csdn.net/weixin_56622231/article/details/159515126
- AgentScope 2.0: From Transparent Development to System Engineering (reference for sandbox and permission design). https://java.agentscope.io/v2/en/blogs/agentscope-v2-release.html
- Claude Agent SDK — Agent Patterns Catalog (reference for execution transparency and permission design). https://www.agentpatternscatalog.org/compositions/claude-agent-sdk