摩尔线程:MTT S5000 与夸娥万卡集群


1. 介绍

1.1. 厂商定位

摩尔线程(688795.SH)是「国产 GPU 第一股」:2025-12-05 上交所科创板挂牌。其技术路线在本组独树一帜——GPU 全功能路线:不做单一 AI 推理加速器(DSA),而是自研全功能 GPU(图形渲染 + AI 计算 + 视频处理),以 MUSA 统一系统架构承载,对标国际主流 GPU 生态。这一路线决定了它的 AI 智算业务(2025 上半年占营收 94.73%)建立在通用 GPU 之上,而非专用加速器。

本篇的叙事主线:GPU 全功能路线 + 万卡集群商业化追赶者;所有性能/利用率指标均为公司口径,缺第三方基准(见 6.3 节与信息缺口声明)。

1.2. 基本信息卡

项目内容
公司摩尔线程智能科技(北京)股份有限公司(688795.SH)
定位全功能 GPU 设计厂商(AI 智算 + 图形 + 专用计算)
上市2025-12-05 科创板挂牌,「国产 GPU 第一股」
旗舰产品MTT S5000(「平湖」架构,2025 年大规模商业化)
集群产品夸娥(KUAE)智算集群,万卡级,浮点算力 10 ExaFLOPS(公司口径)
软件架构MUSA 统一系统架构(全栈自主:芯片/指令集/软件栈)
信息截止2026-09-12

1.3. 财务与上市表现

指标数值口径说明
上市首日2025-12-05 收盘 600.5 元/股(较发行价 +425%),当日成交 ¥153.07 亿元中国证券网 / 快科技(A/B)
2025 上半年营收¥7.02 亿元,AI 智算占 94.73%,毛利率 69.14%公司披露(A)
2025 全年研发费用¥13.05 亿元,占营收 86.68%年报(A)
2026 Q1扭亏为盈中国经济网报道(A/B)
授权专利514 项(至 2025-06)公司材料(B)
已量产芯片5 颗 GPU 芯片、迭代四代架构公司材料(B)

1.4. 在 AI Harness 体系中的位置

摩尔线程是本组唯一一家同时公开了训练效率指标与集群管理平台分层的国产厂商:KUAE Platform(集群管理)+ KUAE ModelStudio(大模型服务)的分层结构,使 L3/L5 有可评估的产品化形态。其 Harness 角色:为「信创 + 智算中心」场景提供从芯片到集群调度平台的完整竖井。


2. 名词解释

术语英文/缩写释义
MTTMoore Threads摩尔线程产品前缀(MTT S 系列 GPU)
全功能 GPUFull-Feature GPU同时支持图形渲染、AI 计算、视频编解码的通用 GPU 路线
MUSAMoore Threads Unified System Architecture摩尔线程统一系统架构:自研指令集与全栈软件,兼容国际主流 GPU 生态
平湖MTT S5000 采用的芯片架构代号
夸娥KUAE摩尔线程智算集群品牌,万卡级,公司口径浮点算力 10 ExaFLOPS
KUAE Platform夸娥集群管理平台(集群调度/运维层)
KUAE ModelStudio夸娥大模型服务平台(模型服务层)
MFUModel FLOPs Utilization模型浮点算力利用率;公司口径 Dense 60% / MoE 40%
有效训练时间占比Effective Training Time Ratio排除故障与空闲后的有效训练时间占比;公司口径 90% 以上
线性扩展效率Linear Scaling Efficiency集群规模扩展时的算力线性度;公司口径 95%
PD 分离Prefill / Decode Disaggregation推理阶段分离部署;公司列为技术布局方向
MTT C256公司发布的超节点架构规划
花港十万卡级集群架构技术攻关项目代号
MTT AIBOOK基于「长江」SoC 的边缘终端产品

3. 功能说明与产品线

3.1. MTT S5000 旗舰芯片

公司口径(业绩说明会 2026-04/05,A 级):

  1. 「平湖」架构,2025 年大规模商业化;
  2. 单卡稠密 AI 算力最高 1000 TFLOPS,FP8—FP64 全精度覆盖;
  3. Day-0 适配 DeepSeek、GLM、MiniMax、Kimi、Qwen 等 SOTA 大模型;
  4. 推理口径(首届 MUSA 开发者大会 2025-12,公司口径):DeepSeek R1 671B 全量模型,MTT S5000 单卡 Prefill 吞吐破 4000 tokens/s、Decode 破 1000 tokens/s。

需要强调:上述推理数据与算力指标均为厂商自报,无 MLPerf 或第三方同场景对比(与 H100/H200/MI300X 的对比数据缺失,见信息缺口声明)。

3.2. 夸娥(KUAE)万卡集群

公司口径(年报 + 业绩说明会,A 级·公司口径):

指标数值说明
部署状态已部署上线
浮点算力10 ExaFLOPS万卡级
Dense 大模型训练 MFU60%公司口径,多项指标称达国际主流水平
MoE 训练 MFU40%公司口径
有效训练时间占比90% 以上公司口径
线性扩展效率95%公司口径

3.3. MUSA 统一系统架构

  1. 全栈自主:芯片、指令集、软件栈均为自研,可独立演进;
  2. 兼容国际主流 GPU 生态:降低既有 CUDA 生态用户的迁移成本;
  3. 生态活动:首届 MUSA 开发者大会(2025-12)披露推理性能与生态进展。

3.4. 路线图

  1. MTT C256:超节点架构规划(公司年报披露,A 级);
  2. 「花港」:十万卡级集群架构技术攻关中;
  3. MTT AIBOOK:基于「长江」SoC 的边缘终端产品。

4. 平台架构

4.1. 从芯片到集群的技术栈

图 5-1|摩尔线程:从 MTT S5000 到夸娥万卡集群的技术栈

摩尔线程技术栈:芯片 → MUSA → 夸娥集群 信息截止 2026-09-12 · 指标均为公司口径 · 示意:基于本文分析绘制 芯片层 MTT S5000「平湖」架构 · 单卡稠密 AI 算力最高 1000 TFLOPS · FP8—FP64 全精度 · Day-0 适配 DeepSeek/GLM/Qwen/Kimi/MiniMax MUSA 统一系统架构 自研指令集 + 全栈软件 · 兼容国际主流 GPU 生态 · 四代架构迭代 / 5 颗芯片量产 夸娥(KUAE)万卡集群(本图重点 · 公司口径) 10 ExaFLOPS 万卡浮点算力 Dense MFU 60% MoE MFU 40% 有效训练时间 90%+ 故障检测与弹性恢复 线性扩展 95% 万卡规模线性度 KUAE Platform(集群管理) 集群调度 / 运维 / 容错 KUAE ModelStudio(模型服务) KV Cache 管理 · PD 分离 · 批处理调度(布局方向) 路线图:MTT C256 超节点架构规划 · 「花港」十万卡级集群攻关 · MTT AIBOOK 边缘终端 平均交付周期 90 天(公司口径)

数据来源:摩尔线程 2025 年年度报告、业绩说明会、首届 MUSA 开发者大会披露;示意图基于本文分析。

4.2. KUAE 平台分层

夸娥体系的产品化分层是本篇架构分析的核心:KUAE Platform 管集群(调度/运维/容错),KUAE ModelStudio 管模型服务(推理服务化)。这一分层与 NVIDIA「Mission Control + Dynamo/NIM」的结构对应关系明显——国内厂商在集群管理平台的形态上已趋同于国际头部实践。


5. Harness 设计

5.1. 六层能力总览

支撑产品/机制成熟度
L1 上下文工程推理平台 KV Cache 管理(公司列为技术布局方向)
L2 工具与执行MUSA 软件栈 + Day-0 大模型适配
L3 编排与控制KUAE Platform 集群调度 + ModelStudio 模型服务分层
L4 记忆与状态推理解耦(PD 分离)、批处理调度(布局方向);万卡级故障检测与弹性恢复
L5 评估与观测有效训练时间占比 90%+、线性扩展效率 95%(公司口径指标体系)中强(指标公开但无第三方验证)
L6 治理与安全全栈自主可控(信创场景适配)强(合规维度)

5.2. L1 / L4 上下文与状态层

公司年报明确将「KV Cache 管理、推理解耦(PD 分离)、批处理调度、低时延服务编排」列为技术布局方向(A 级·公司口径)——注意这是布局方向而非已交付能力,公开材料中未见对应产品组件的详细规格。L4 的万卡级故障检测与弹性恢复有集群稳定性指标佐证(有效训练时间 90%+)。

5.3. L3 编排与控制层

KUAE Platform(集群管理)与 KUAE ModelStudio(大模型服务)的双平台分层,提供集群调度与模型服务的产品化入口。与 09 篇开源 serving 生态相比,摩尔线程选择自建平台层而非直接复用 vLLM/K8s 生态——这在信创交付语境下是务实选择,但生态开放度待观察。

5.4. L5 评估与观测层

摩尔线程的可贵之处在于公开了训练效率指标体系本身:Dense MFU 60%、MoE MFU 40%、有效训练时间占比 90%+、线性扩展效率 95%——这套指标与 10 篇训练框架生态的 MFU 经验区间(70B Dense:Megatron 45—55%+)可比对。但必须强调:全部为公司口径,无 MLPerf 或独立第三方复核(弱于同组 NVIDIA 的 A 级验证)。

5.5. L6 治理与安全层

全栈自主可控(芯片/指令集/软件栈)是其在信创与关键行业场景的核心治理卖点;省级智算中心客户(政府背景)的连续中标构成侧证。


6. 实际案例

6.1. 夸娥万卡集群:公司口径指标

万卡集群已部署上线(公司口径 10 ExaFLOPS),配套指标见 3.2 节。案例写法锚点:夸娥集群 MFU 指标MTT S5000 单卡 R1 671B 推理吞吐两个公司口径数据——两者均缺第三方基准,是评估摩尔线程真实水平时必须携带的免责前提。

6.2. 智算中心与商业订单

订单/部署内容口径
省级智算中心KUAE 中标湖北、山东、四川等 6 省级智算中心,合计约 ¥20 亿元百度百科引公司披露(C,引公司披露)
万卡训练集群2026-03 签订 ¥6.6 亿元万卡训练集群大单公司年报(A)
千卡集群北京、南京等地部署 3 个千卡集群公司披露(C)
运营商合作与中国移动青海公司战略合作公司披露(C)
交付效率平均交付周期 90 天公司口径

6.3. 数据可信度说明

本篇几乎所有性能与利用率数据均为公司口径(A 级·公司披露 ≠ 第三方验证),执行以下三条纪律:

  1. 公司口径逐处标注:MFU、ExaFLOPS、tokens/s 等指标出现处均已标注「公司口径」;
  2. 不与竞品做未经第三方验证的对比:MTT S5000 与 H100/H200/MI300X 无同场景第三方对比数据,本文不做任何性能对标表述;
  3. 商业订单以公司披露/年报为准:¥6.6 亿元大单出自年报(A 级),¥20 亿元智算中心合计出自百科转引(C 级),可信度分级已逐条标注。

7. 总结

优势

  1. GPU 全功能路线的稀缺性:国内唯一以全功能 GPU 上市的公司,图形 + AI 双能力在智算中心与信创场景有组合价值;
  2. 指标体系公开:主动披露 MFU/扩展效率/有效训练时间等训练侧核心指标,透明度高于多数国产同行;
  3. 商业化验证:上市首年扭亏(2026 Q1)、¥6.6 亿元万卡大单、6 省智算中心覆盖;
  4. 平台产品化:KUAE 双平台分层提供可交付的集群管理形态。

劣势

  1. 无第三方基准:全部性能指标为公司口径,MLPerf 提交缺失,跨厂商比较不可行;
  2. MUSA 生态规模待检验:算子覆盖率、开源仓库规模无官方数字;
  3. 研发强度高企:研发费用占营收 86.68%,盈利能力对营收增长敏感。

适用边界:省级/区域智算中心建设(信创与自主可控要求)、需要全功能 GPU(AI + 图形)混合负载的场景、万卡级训练集群的国产化尝试。

选型建议:把「公司口径」指标作为初始筛选而非验收标准,验收时要求在自有模型/负载上实测 MFU 与扩展效率;万卡级部署前验证 MUSA 在目标训练框架(Megatron 风格 API,见 10 篇)上的算子覆盖率;关注 MTT C256 超节点与「花港」十万卡架构的落地节奏。

信息缺口声明

  1. 第三方基准(MLPerf 等)验证缺失:夸娥集群 MFU/扩展效率均为公司口径,无独立复核;
  2. MTT S5000 与国际同代产品(H100/H200/MI300X)的同场景第三方对比数据缺失;
  3. MUSA 生态开源仓库规模与算子覆盖率无官方数字;
  4. 省级智算中心 ¥20 亿元合计口径出自百科转引公司披露(C 级),建议以公司公告二次核验。

8. 参考资料

  1. KUAE(夸娥)智算中心 — 百度百科(交叉公司披露),2025—2026。https://baike.baidu.com/item/KUAE(夸娥)智算中心/63862920
  2. 摩尔线程 2025 年度暨 2026 年第一季度业绩说明会 — 中国证券网路演中心,2026。https://roadshow.cnstock.com/fbh/mexc2025
  3. 「国产 GPU 第一股」诞生!摩尔线程正式上市 — 快科技,2025-12。https://news.mydrivers.com/1/1090/1090711.htm
  4. 国产 GPU 商业化提速 摩尔线程一季度扭亏为盈 — 中国经济网,2026-04。https://www.ce.cn/xwzx/gnsz/gdxw/202604/t20260427_2932385.shtml
  5. 摩尔线程 2025 年年度报告 — 东方财富数据中心,2026。https://data.eastmoney.com/notices/detail/688795/AN202604261821587780.html
  6. 摩尔线程官网(MUSA 产品线)— 摩尔线程,2025。https://www.mthreads.com/

Moore Threads: MTT S5000 and the KUAE 10K-Card Cluster

1. Introduction

1.1. Vendor Positioning

Moore Threads (688795.SH) is the "first stock of domestic GPUs": it listed on the Shanghai Stock Exchange's STAR Market on 2025-12-05. Its technical route stands apart within this group — the full-feature GPU route: rather than making a single-purpose AI inference accelerator (DSA), it self-develops full-feature GPUs (graphics rendering + AI compute + video processing), carried on the MUSA unified system architecture, benchmarking against the international mainstream GPU ecosystem. This route determines that its AI intelligent-computing business (94.73% of revenue in H1 2025) is built on general-purpose GPUs, not dedicated accelerators.

The narrative thread of this article: full-feature GPU route + commercial catch-up player in 10K-card clusters; all performance/utilization metrics are company figures, lacking third-party benchmarks (see Section 6.3 and the information-gap declaration).

1.2. Basic Information Card

ItemDetail
CompanyMoore Threads Intelligent Technology (Beijing) Co., Ltd. (688795.SH)
PositioningFull-feature GPU design vendor (AI intelligent computing + graphics + specialized compute)
ListingListed on STAR Market 2025-12-05, "first stock of domestic GPUs"
Flagship productMTT S5000 ("Pinghu" architecture, large-scale commercialization in 2025)
Cluster productKUAE intelligent-computing cluster, 10K-card scale, 10 ExaFLOPS floating-point compute (company figures)
Software architectureMUSA unified system architecture (full-stack self-developed: chip / instruction set / software stack)
Information cutoff2026-09-12

1.3. Financial and Listing Performance

MetricValueBasis Notes
Listing first dayClosed at 600.5 CNY/share on 2025-12-05 (+425% vs. issue price), turnover ¥153.07 亿 on the dayChina Securities Network / Kuaikeji (A/B)
H1 2025 revenue¥7.02 亿, AI intelligent computing 94.73%, gross margin 69.14%Company disclosure (A)
2025 full-year R&D expense¥13.05 亿, 86.68% of revenueAnnual report (A)
2026 Q1Turned profitableChina Economic Net report (A/B)
Granted patents514 (as of 2025-06)Company materials (B)
Mass-produced chips5 GPU chips, four generations of architecture iteratedCompany materials (B)

1.4. Position in the AI Harness System

Moore Threads is the only domestic vendor in this group that publicly discloses both training-efficiency metrics and a layered cluster-management platform: the layered structure of KUAE Platform (cluster management) + KUAE ModelStudio (LLM service) gives L3/L5 a productized form that can be assessed. Its Harness role: providing a complete silo from chip to cluster-scheduling platform for the "Xinchuang + intelligent-computing center" scenarios.


2. Glossary

TermEnglish / AbbreviationDefinition
MTTMoore ThreadsMoore Threads product prefix (MTT S-series GPUs)
全功能 GPUFull-Feature GPUA general-purpose GPU route supporting graphics rendering, AI compute, and video encode/decode simultaneously
MUSAMoore Threads Unified System ArchitectureMoore Threads unified system architecture: self-developed instruction set and full-stack software, compatible with the international mainstream GPU ecosystem
平湖Codename of the chip architecture used by the MTT S5000
夸娥KUAEMoore Threads' intelligent-computing cluster brand, 10K-card scale, company-figure floating-point compute of 10 ExaFLOPS
KUAE PlatformKUAE cluster-management platform (cluster scheduling / operations layer)
KUAE ModelStudioKUAE LLM-service platform (model-service layer)
MFUModel FLOPs UtilizationModel FLOPs utilization; company figures: Dense 60% / MoE 40%
有效训练时间占比Effective Training Time RatioProportion of effective training time after excluding failures and idle time; company figure: above 90%
线性扩展效率Linear Scaling EfficiencyComputing linearity when the cluster scales; company figure: 95%
PD 分离Prefill / Decode DisaggregationSeparated deployment of inference stages; listed by the company as a technology layout direction
MTT C256The super-node architecture plan published by the company
花港Codename of the technical-research project for 100K-card cluster architecture
MTT AIBOOKEdge-terminal product based on the "Changjiang" SoC

3. Feature Description and Product Line

3.1. MTT S5000 Flagship Chip

Company figures (earnings-call sessions 2026-04/05, grade A):

  1. "Pinghu" architecture, large-scale commercialization in 2025;
  2. Up to 1000 TFLOPS of dense AI compute per card, full precision coverage from FP8 to FP64;
  3. Day-0 adaptation to SOTA LLMs such as DeepSeek, GLM, MiniMax, Kimi, and Qwen;
  4. Inference figures (first MUSA Developer Conference 2025-12, company figures): DeepSeek R1 671B full model on a single MTT S5000, Prefill throughput exceeding 4000 tokens/s and Decode exceeding 1000 tokens/s.

It must be emphasized: the above inference data and compute metrics are all self-reported by the vendor, with no MLPerf or third-party same-scenario comparison (comparative data vs. H100/H200/MI300X is missing; see the information-gap declaration).

3.2. KUAE 10K-Card Cluster

Company figures (annual report + earnings-call session, grade A · company figures):

MetricValueNotes
Deployment statusDeployed and in operation
Floating-point compute10 ExaFLOPS10K-card scale
Dense LLM training MFU60%Company figure; multiple metrics claimed to reach international mainstream levels
MoE training MFU40%Company figure
Effective training time ratioAbove 90%Company figure
Linear scaling efficiency95%Company figure

3.3. MUSA Unified System Architecture

  1. Full-stack self-developed: chip, instruction set, and software stack are all self-developed and can evolve independently;
  2. Compatible with the international mainstream GPU ecosystem: lowers the migration cost for existing CUDA-ecosystem users;
  3. Ecosystem activity: the first MUSA Developer Conference (2025-12) disclosed inference performance and ecosystem progress.

3.4. Roadmap

  1. MTT C256: super-node architecture plan (disclosed in the company annual report, grade A);
  2. "Huagang": technical research underway for 100K-card cluster architecture;
  3. MTT AIBOOK: an edge-terminal product based on the "Changjiang" SoC.

4. Platform Architecture

4.1. Technology Stack, from Chip to Cluster

Figure 5-1 | Moore Threads: the technology stack from MTT S5000 to the KUAE 10K-card cluster

摩尔线程技术栈:芯片 → MUSA → 夸娥集群 信息截止 2026-09-12 · 指标均为公司口径 · 示意:基于本文分析绘制 芯片层 MTT S5000「平湖」架构 · 单卡稠密 AI 算力最高 1000 TFLOPS · FP8—FP64 全精度 · Day-0 适配 DeepSeek/GLM/Qwen/Kimi/MiniMax MUSA 统一系统架构 自研指令集 + 全栈软件 · 兼容国际主流 GPU 生态 · 四代架构迭代 / 5 颗芯片量产 夸娥(KUAE)万卡集群(本图重点 · 公司口径) 10 ExaFLOPS 万卡浮点算力 Dense MFU 60% MoE MFU 40% 有效训练时间 90%+ 故障检测与弹性恢复 线性扩展 95% 万卡规模线性度 KUAE Platform(集群管理) 集群调度 / 运维 / 容错 KUAE ModelStudio(模型服务) KV Cache 管理 · PD 分离 · 批处理调度(布局方向) 路线图:MTT C256 超节点架构规划 · 「花港」十万卡级集群攻关 · MTT AIBOOK 边缘终端 平均交付周期 90 天(公司口径)

Data sources: Moore Threads 2025 annual report, earnings-call sessions, and disclosures at the first MUSA Developer Conference; the diagram is drawn based on this article's analysis.

4.2. KUAE Platform Layering

The productized layering of the KUAE system is the core of this article's architecture analysis: KUAE Platform manages the cluster (scheduling/operations/fault tolerance), while KUAE ModelStudio manages model services (inference-as-a-service). This layering maps clearly onto the structure of NVIDIA's "Mission Control + Dynamo/NIM" — domestic vendors have converged on the practices of leading international players in the form of cluster-management platforms.


5. Harness Design

5.1. Six-Layer Capability Overview

LayerSupporting Product / MechanismMaturity
L1 Context EngineeringInference-platform KV Cache management (listed by the company as a technology layout direction)Medium
L2 Tooling & ExecutionMUSA software stack + Day-0 LLM adaptationMedium
L3 Orchestration & ControlKUAE Platform cluster scheduling + ModelStudio model-service layeringMedium
L4 Memory & StateInference decoupling (PD separation), batch scheduling (layout direction); 10K-card fault detection and elastic recoveryMedium
L5 Evaluation & ObservabilityEffective training time ratio 90%+, linear scaling efficiency 95% (company-figure indicator system)Medium-strong (metrics public but no third-party verification)
L6 Governance & SecurityFull-stack self-controlled (Xinchuang scenario adaptation)Strong (compliance dimension)

5.2. L1 / L4 Context and State Layers

The company annual report explicitly lists "KV Cache management, inference decoupling (PD separation), batch scheduling, and low-latency service orchestration" as technology layout directions (grade A · company figures) — note these are layout directions, not delivered capabilities, and no detailed specifications of corresponding product components appear in public materials. L4's 10K-card fault detection and elastic recovery are supported by cluster-stability metrics (effective training time above 90%).

5.3. L3 Orchestration and Control Layer

The dual-platform layering of KUAE Platform (cluster management) and KUAE ModelStudio (LLM service) provides productized entry points for cluster scheduling and model services. Compared with the open-source serving ecosystem of Article 09, Moore Threads chose to build its own platform layer rather than directly reuse the vLLM/K8s ecosystem — this is a pragmatic choice in the Xinchuang-delivery context, but the openness of the ecosystem remains to be seen.

5.4. L5 Evaluation and Observability Layer

What is valuable about Moore Threads is that it publicly discloses the training-efficiency indicator system itself: Dense MFU 60%, MoE MFU 40%, effective training time ratio 90%+, linear scaling efficiency 95% — this set of indicators is comparable to the MFU experience ranges of the training-framework ecosystem in Article 10 (70B Dense: Megatron 45—55%+). But it must be emphasized: all are company figures, with no MLPerf or independent third-party verification (weaker than the grade-A verification of NVIDIA in this group).

5.5. L6 Governance and Security Layer

Full-stack self-control (chip / instruction set / software stack) is its core governance selling point in Xinchuang and key-industry scenarios; successive wins of provincial intelligent-computing-center customers (with government backgrounds) constitute supporting evidence.


6. Practical Cases

6.1. KUAE 10K-Card Cluster: Company-Figure Metrics

The 10K-card cluster has been deployed and is in operation (company figure: 10 ExaFLOPS), with supporting metrics in Section 3.2. Case-writing anchors: two company-figure data points — the KUAE cluster MFU metrics and the MTT S5000 single-card R1 671B inference throughput — both lack third-party benchmarks, and are the disclaimer premise that must be carried when assessing Moore Threads' true level.

6.2. Intelligent-Computing Centers and Commercial Orders

Order / DeploymentDetailBasis
Provincial intelligent-computing centersKUAE won 6 provincial intelligent-computing centers including Hubei, Shandong, and Sichuan, totaling about ¥20 亿Baidu Baike citing company disclosure (C, citing company disclosure)
10K-card training clusterSigned a ¥6.6 亿 10K-card training-cluster order in 2026-03Company annual report (A)
Thousand-card clustersDeployed 3 thousand-card clusters in Beijing, Nanjing, and elsewhereCompany disclosure (C)
Operator cooperationStrategic cooperation with China Mobile QinghaiCompany disclosure (C)
Delivery efficiencyAverage delivery cycle of 90 daysCompany figure

6.3. Data Credibility Notes

Almost all performance and utilization data in this article are company figures (grade A · company disclosure ≠ third-party verification), following these three disciplines:

  1. Company figures annotated at each occurrence: wherever metrics such as MFU, ExaFLOPS, and tokens/s appear, they are marked "company figures";
  2. No comparison against competitors without third-party verification: there is no same-scenario third-party comparative data for MTT S5000 vs. H100/H200/MI300X, and this article makes no performance benchmarking claims;
  3. Commercial orders follow company disclosure / annual report: the ¥6.6 亿 order comes from the annual report (grade A), while the ¥20 亿 intelligent-computing-centers total comes from a Baidu Baike citation (grade C); the credibility grading is annotated item by item.

7. Summary

Strengths:

  1. Scarcity of the full-feature GPU route: the only company in China to go public with a full-feature GPU, and the combination of graphics + AI capabilities has associative value in intelligent-computing-center and Xinchuang scenarios;
  2. Public indicator system: proactively discloses core training-side metrics such as MFU / scaling efficiency / effective training time, with transparency higher than most domestic peers;
  3. Commercial validation: turned profitable in its first year after listing (2026 Q1), a ¥6.6 亿 10K-card order, and coverage of intelligent-computing centers in 6 provinces;
  4. Platform productization: the dual KUAE platform layering provides a deliverable cluster-management form.

Weaknesses:

  1. No third-party benchmarks: all performance metrics are company figures, MLPerf submissions are missing, and cross-vendor comparison is not feasible;
  2. MUSA ecosystem scale yet to be validated: no official figures for operator coverage or open-source repository scale;
  3. High R&D intensity: R&D expense is 86.68% of revenue, and profitability is sensitive to revenue growth.

Applicable boundaries: provincial/regional intelligent-computing-center construction (Xinchuang and self-control requirements), scenarios needing full-feature GPU (AI + graphics) mixed loads, and domestic efforts at 10K-card training clusters.

Selection recommendations: treat "company-figure" metrics as initial screening rather than acceptance criteria, and require measurement of MFU and scaling efficiency on your own models/loads at acceptance time; validate MUSA's operator coverage on the target training framework (Megatron-style API, see Article 10) before 10K-card deployment; and watch the landing pace of the MTT C256 super-node and the "Huagang" 100K-card architecture.

Information-Gap Declaration

  1. Missing verification by third-party benchmarks (MLPerf, etc.): the KUAE cluster's MFU / scaling efficiency are all company figures with no independent review;
  2. Missing same-scenario third-party comparative data between the MTT S5000 and international same-generation products (H100/H200/MI300X);
  3. No official figures for MUSA ecosystem open-source repository scale and operator coverage;
  4. The ¥20 亿 provincial intelligent-computing-centers total is based on a Baidu Baike citation of company disclosure (grade C); it is recommended to re-verify against company announcements.

8. References

  1. KUAE Intelligent-Computing Center — Baidu Baike (cross-checked with company disclosure), 2025—2026. https://baike.baidu.com/item/KUAE(夸娥)智算中心/63862920
  2. Moore Threads 2025 Annual and 2026 Q1 Earnings-Call Session — China Securities Network Roadshow Center, 2026. https://roadshow.cnstock.com/fbh/mexc2025
  3. "First Stock of Domestic GPUs Is Born! Moore Threads Officially Listed" — Kuaikeji, 2025-12. https://news.mydrivers.com/1/1090/1090711.htm
  4. "Domestic GPU Commercialization Speeds Up; Moore Threads Returns to Profit in Q1" — China Economic Net, 2026-04. https://www.ce.cn/xwzx/gnsz/gdxw/202604/t20260427_2932385.shtml
  5. Moore Threads 2025 Annual Report — Eastmoney Data Center, 2026. https://data.eastmoney.com/notices/detail/688795/AN202604261821587780.html
  6. Moore Threads Official Website (MUSA product line) — Moore Threads, 2025. https://www.mthreads.com/