Skip to main content

15 posts tagged with "Domestic Chips"

Chinese domestic AI chip development

View all tags

DeepSeek Confirms Betting on Huawei Chips for LLM Training: From the 160,000-Unit 950DT Rumor to a "Must Succeed" Commitment

· 5 min read
Industry Research Team

On September 22, 2026, multiple media outlets reported that DeepSeek's CEO explicitly stated the company is betting heavily on Huawei chips, expects to receive a new batch of Huawei chips for large-model training, and stressed that "this choice must succeed." Following earlier reports that "DeepSeek planned to procure 160,000 Ascend 950DT units for inference," this is the most significant alignment signal yet in the domestic compute ecosystem — an upgrade from inference procurement to a training bet.


1. The Signal Chain: Four Steps to a "Training Bet"​

Piecing together the public information from the past six months, the DeepSeek–Ascend collaboration shows a clear escalation path:

TimeSignalNature
Mid-2026DeepSeek V4 completes ecosystem migration from CUDA to CANNSoftware adaptation
August 2026Report: DeepSeek plans to procure 160,000 Ascend 950DT units for inference deploymentInference procurement intent
2026-09-17HC2026: 960DT ready three quarters ahead of schedule, 950DT ramping in Q4Supply delivery
2026-09-22DeepSeek CEO: betting on Huawei chips for LLM training, "must succeed"Training bet

The key change is the workload tier: inference deployment means "cutting costs with off-the-shelf compute," while a training bet means "staking the existence of next-generation models on domestic chips" — training clusters demand an order of magnitude more in stability, interconnect efficiency, and software-stack maturity.

2. Huawei's Ability to Deliver​

DeepSeek's willingness to commit rests on a series of verifiable progress points on Huawei's supply side (all official figures):

  • One generation per year, delivered: the 950PR is in mass production, the 950DT ramps in 2026 Q4, and the 960DT moved three quarters earlier than its original 2027 Q4 plan to ready in 2027 Q1 / launch in Q2 — the roadmap's credibility validated twice in a row;
  • Deployment scale: Ascend supernodes have been commercially deployed at scale in over 1,000 sets, covering internet, finance, healthcare, and manufacturing;
  • Ecosystem maturity: CANN has entered routine open-source operation, with external developers exceeding 61% for the first time and 5,200 monthly active developers; there are over 40 Ascend-native training models, making it the only domestic technology route supporting pretraining;
  • Training evidence chain: China Telecom's Xing 4.0-29B-A4B agentic MoE model, open-sourced in September, was announced as trained end-to-end on Ascend (company claim) — "Ascend can train large models" is no longer just Huawei's self-attestation.

See the Ascend 950DT and Ascend 960 spec pages for details; for supernode analysis, see Ascend 960 Official Launch.

3. Why DeepSeek?​

DeepSeek's choice has strong structural drivers:

  1. Supply certainty: under export-control constraints, NVIDIA's flagship supply to China keeps tightening; domestic compute is the only plannable large-scale training supply;
  2. Cost structure: DeepSeek has always been known for extreme engineering efficiency (the V3 training cost set the industry benchmark), and domestic compute plus supernode system efficiency fits its approach;
  3. Betting on ecosystem dividends: CANN open-sourcing plus deep binding with leading model vendors means model vendors can participate in shaping the toolchain's direction — something impossible within CUDA's closed system;
  4. Self-fulfilling demonstration effect: a leading lab's public commitment pulls back on Huawei's production scheduling and upstream HBM and system investment, making "must succeed" a rational commitment rather than a slogan.

4. Risks and Open Questions​

Viewed coolly, this route still has three items that need time to verify:

  • Actual training scale: the quantity, model type (950DT or 960DT), and delivery schedule of the new batch of chips are all undisclosed;
  • MFU methodology: the MFU / latency gains Huawei cites all come from Markov-lab simulations, with no independent third-party measurements yet; the real effective compute of a training cluster depends on long-term data from large-scale production environments;
  • Per-card generation gap: the 960DT's 4 PFLOPS (FP4) versus Rubin R200's 50 PFLOPS (FP4) — the per-card gap objectively exists, and training efficiency depends on whether supernode scale and software optimization can compensate; that is precisely the decisive battleground of "system-level competition."

5. Summary​

  • DeepSeek confirms betting on Huawei chips for training large models — the first training-grade commitment from a leading lab in domestic compute;
  • Signal chain: CUDA→CANN migration → 160,000-unit 950DT inference procurement rumor → 960 ready ahead of schedule → training bet;
  • Supporting factors: one-generation-per-year delivery, 1,000+ supernode sets, and 61% external developers in the CANN open-source ecosystem;
  • What to watch: actual arrival of the new chips and training-cluster scale, third-party MFU data, and training-compute disclosure in DeepSeek's next release.

Further Reading​

References​

  • Toutiao Tech Morning Report: DeepSeek confirms betting on Huawei chips for large-model training (2026-09-22)
  • HUAWEI CONNECT 2026 official announcements (2026-09-17)
  • Earlier report: DeepSeek plans to procure 160,000 Ascend 950DT units (2026-08)

This article is compiled from public reports and vendors' official statements. Details such as chip quantities and training-cluster scale are subject to subsequent disclosures from DeepSeek and Huawei.

Three Signals for the Domestic Compute Ecosystem in One Week: China Telecom Open-Sources Ascend-Trained Model, Inspur 128-Chip Inference Appliance, EVAS Raises RMB 2 Billion

· 5 min read
Industry Research Team

Beyond the official Ascend 960 unveiling and the Zhenwu V900 launch, late September also delivered three ecosystem signals that headline coverage easily drowned out, yet whose structural significance rivals flagship chips: a carrier open-sourcing a model trained on Ascend, an OEM delivering a hundred-chip-class domestic inference system, and a cloud-compute chip startup closing a new funding round. This article breaks down each one.


1. China Telecom Open-Sources Xing4.0: Carrier-Grade Endorsement for End-to-End Ascend Training​

On September 23 (as reported), China Telecom open-sourced Xing4.0-29B-A4B — an agentic mixture-of-experts (MoE) model that the company says was trained end-to-end on Huawei Ascend accelerators.

Why it matters:

  • Third-party proof that "it can train": until now, evidence that "Ascend can train large models" came mainly from Huawei itself (40+ natively trained models); a carrier running training with its own data, engineering teams, and clusters — and open-sourcing it — is the first scaled training case outside the Huawei ecosystem;
  • Agentic positioning: the 29B-A4B (29B total / 4B active) MoE form factor targets agent workloads directly — aligned with the open-source cadence of DeepSeek and Qwen, rather than a small experimental model;
  • Open-source spillover: publishing model weights plus the training recipe means Ascend training engineering know-how can be reused by other institutions — the standard playbook for ecosystem diffusion.

Combined with Huawei's disclosure that "CANN has entered routine open-source operations, with external developers at 61%", the Ascend ecosystem is shifting from "vendor-led" to "community-operated".

2. Inspur MetaBrain SD200 Ultra: Measured Throughput Claims for a 128-Chip Domestic Inference System​

Inspur launched the MetaBrain SD200 Ultra inference system around the same time:

MetricValue (company figures)
Domestic AI chips128 chips (vendor did not disclose specific models)
Model servedKimi K3
Token throughput2.8T tokens (system throughput figure under that methodology)
Latency5.85 ms

Three readings:

  • Hundred-chip-class system integration: 128 chips cooperating across a heterogeneous setup to run very large MoE inference is a test of interconnect topology, scheduling, and fault tolerance — OEMs are now demonstrably capable of assembling domestic chips into large systems;
  • Aimed at top open-source models: targeting Kimi K3 (already live on AWS Bedrock) as the benchmark workload shows the acceptance standard for domestic inference systems is "runs today's most popular models", not a self-referential demo;
  • Mind the methodology: the 2.8T token throughput and 5.85ms latency are both vendor claims; the tested model version, concurrency, and batch size were not disclosed, so procurement evaluations should re-test under real workloads.

3. EVAS Closes RMB 2 Billion B+ Round: The Cloud Compute Chip Race Still Attracts Capital​

EVAS (Yixing Intelligence) completed a B+ round of RMB 2 billion, at a post-money valuation near RMB 15 billion, with Oriza Capital following on; earlier in the first half of the year it had closed a RMB 1.5 billion Series B and brought in China Mobile as a strategic investor. The company focuses on next-generation cloud compute chips.

In 2026, when the flagship chip lineup (Ascend / Cambricon / Moore Threads / Hygon / Zhenwu) looks settled, a cloud compute chip newcomer still raising RMB 2 billion suggests:

  • Primary-market conviction in the "second tier of domestic compute": with top vendors' capacity booked into 2027, overflow demand gives newcomers a window;
  • The role of carrier strategic investment: China Mobile is both a buyer and an investor — the demand side of domestic compute is using capital to lock in the supply side;
  • Differentiated room remains for cloud inference DSA routes (see the ecosystem positioning of routes like Tsingmicro and Houmo).

4. Assembling the Week's Signals: The Loop Is Taking Shape​

Put the three signals together with this month's flagship launches, and every segment of the domestic compute loop now has players filling the gaps:

SegmentThis Month's Evidence
Flagship chipsAscend 960 early official unveiling (9-17), Zhenwu V900 launch (9-22)
Training validationChina Telecom's Xing4.0 end-to-end trained on Ascend and open-sourced, DeepSeek betting on Ascend training
Inference systemsInspur SD200 Ultra running Kimi K3 on 128 chips
Software ecosystemCANN routine open-sourcing, external developers at 61%
Capital supplyEVAS B+ round of RMB 2 billion (post-money near RMB 15 billion)
Demand sideDeepSeek's alignment, Kimi K3 live on overseas cloud platforms

Independent players are investing across all four segments — "chips, systems, models, capital" — the biggest difference from the "isolated breakthroughs" of 2024-2025.

5. Takeaways​

  • China Telecom Xing4.0-29B: the first open-source model end-to-end trained on Ascend outside the Huawei ecosystem;
  • Inspur SD200 Ultra: a 128-chip domestic inference system at 2.8T token throughput / 5.85ms (company figures, pending re-testing);
  • EVAS RMB 2 billion B+ round: the cloud compute chip second tier is still getting real money;
  • What to watch: Xing4.0's training cluster scale and MFU, third-party re-tests of the SD200 Ultra, and EVAS's tape-out cadence.

Further Reading​

References​

  • The GPU Daily (2026-09-24): China Telecom open-sources Xing4.0-29B; Inspur MetaBrain SD200 Ultra
  • Toutiao (2026-09-21): EVAS completes RMB 2 billion B+ round, post-money valuation near RMB 15 billion
  • CSDN AI Daily (2026-09-21): Huawei CANN enters routine open-source operations, OceanStor M900 launched

This article is compiled from public reports. Xing4.0 training details and SD200 Ultra performance figures are company statements; testing conditions are subject to subsequent disclosures.

Alibaba Zhenwu V900 Unveiled at Apsara Conference: 3x M890 Performance, 216GB Memory, 500,000-Card Cluster, Mass Production in 2027 Q1

· 5 min read
Industry Research Team

Just five days after Huawei officially announced the Ascend 960 at HUAWEI CONNECT on September 17, Alibaba unveiled its new unified training-and-inference AI chip, Zhenwu V900, at the Apsara Conference in Hangzhou on September 22 — which Alibaba calls "the most powerful Chinese self-developed AI chip in terms of compute performance to date." This article is compiled from Alibaba's official announcements and reports from Sina Tech, C114, Huanqiu.com, and other sources.


1. Single Chip: 3x M890, 216GB + 1200GB/s​

MetricZhenwu V900 (this launch)Zhenwu M890 (previous generation)
Performance3x Zhenwu M890 (official figure; absolute value undisclosed)FP16 600 TFLOPS (as catalogued on this site)
Memory216 GB144 GB (HBM3)
Chip-to-chip interconnect1200 GB/s—
PrecisionNative FP8 / FP4 (including high-precision training)—
PositioningUnified training and inference (trillion-parameter-scale training + low-precision inference)Unified training and inference
Mass production2027 Q1, scaled deployment in Alibaba Cloud data centersAlready deployed at scale

Three takeaways:

  • Memory crosses into the 200GB+ tier: 216GB puts it in the same capacity class as the Ascend 960DT (288GB self-developed HBM) and NVIDIA Rubin (288GB HBM4), leveling the capacity threshold for long-context and very large MoE models;
  • 1200GB/s chip-to-chip interconnect: laying the foundation for the supernode's "memory-semantics interconnect" — the chip-level prerequisite that pairs with ICN Switch;
  • Precision coverage across all scenarios: officially described as "high-precision training, low-precision and ultra-low-precision inference across all scenarios," with native FP8/FP4 support aligned with the common spec of 2026's new cards.

The absolute single-chip compute figure has not been officially disclosed; this site records it using the official relative figure of "3x M890." See the Zhenwu M890 spec page for M890 details.

2. System Level: ICN Switch + Panjiu Supernode, a 500,000-Card Single Cluster​

V900's real selling point is not the single chip but system-level collaboration:

  • ICN Switch self-developed interconnect chip: once connected, the supernode gains native memory semantics and unified memory addressing, with a thousand cards interconnected at full bandwidth — over a thousand V900s can "work together like a single super chip";
  • Panjiu supernode server: integrates V900 (compute) + ICN Switch (interconnect) + Panmai intelligent NIC (network) + Zhenyue SSD controller (storage), a fully self-developed compute-storage-network stack;
  • 500,000-card single cluster: combined with Alibaba Cloud's next-generation intelligent computing center network architecture, a single cluster can scale up to 500,000 cards.

This mirrors Huawei's "11 key chips" approach: the unit of competition has shifted from the single chip to whole-system delivery capability — Zhenwu handles compute, Yitian handles general-purpose computing, Panmai handles networking, Zhenyue handles storage, and ICN Switch handles interconnect.

3. Business and Roadmap​

  • The Zhenwu family has served over 650+ enterprise customers (as of June 2026), spanning autonomous driving, finance, large models, embodied AI, energy, and manufacturing;
  • Supernodes based on the M890 are already deployed at scale, running models with over 2 trillion parameters such as Qwen3.8 and Kimi K3; Alibaba Cloud will add new serving nodes in Q4 to expand supernode supply;
  • Roadmap: V900 enters mass production and sales in 2027 Q1; Zhenwu J900 is planned for release in 2027 Q3;
  • CPU synergy: Yitian 720 / 730 arrive in 2027 (the 730 is the first to adopt T-Head's fully self-developed CPU microarchitecture, with single-core SPECint2017/GHz up to 1.4x that of Yitian 710); the 2029 Yitian 750 will interconnect directly with Zhenwu AI chips via the ICN bus;
  • Alibaba Group CEO Eddie Wu said T-Head's AI chip annual shipment volume will increase substantially.

4. Competitive Coordinates: Two Swords in One Week​

Viewing the two mid-September launches side by side, the landscape of "system-level competition" among domestic AI chips is now clear:

DimensionHuawei Ascend 960 (9-17)Alibaba Zhenwu V900 (9-22)
Launch eventHC2026Apsara Conference 2026
Per-card memory288GB (self-developed HBM)216GB
Memory bandwidth9.6 TB/sUndisclosed
InterconnectLinJu UnifiedBus / NPO optical interconnectICN Switch memory-semantics interconnect
SupernodeAscend 960 supernode (4,096 cards, 8 EFLOPS FP8)Panjiu supernode (500,000-card single cluster)
Availability2027 Q22027 Q1

Huawei is taking the "supernode + open-source CANN ecosystem" route, while Alibaba is taking the "full cloud stack + open-source model ecosystem" route; both still trail in per-card specs, but what they deliver are procurable, operable ultra-large-scale clusters. The substitution logic against NVIDIA's CUDA ecosystem is shifting from "performance parity" to "supply certainty + system efficiency."

5. Summary​

  • V900: 3x M890, 216GB, 1200GB/s chip-to-chip interconnect, native FP8/FP4, mass production in 2027 Q1;
  • System: ICN Switch with thousand-card full bandwidth, Panjiu supernode with a fully self-developed compute-storage-network stack, 500,000-card single cluster;
  • Cadence: J900 in 2027 Q3, Yitian 730 in 2027, Yitian 750 in 2029 — T-Head's first three-year generational roadmap;
  • What to watch: actual mass production and Alibaba Cloud deployment scale in 2027 Q1, disclosure of V900's absolute compute, and the possibility of external supply of ICN Switch.

Further Reading​

References​

  • Huanqiu.com: T-Head launches new unified training-and-inference AI chip Zhenwu V900 (2026-09-22)
  • Sina Tech / TechWeb: Alibaba T-Head launches new AI chip Zhenwu V900, mass production in the first quarter of 2027
  • C114: Alibaba launches new AI chip Zhenwu V900, mass production in 2027 Q1

This article is compiled from Alibaba's official announcements and public reports. The absolute single-chip compute of V900 has not been officially disclosed; supernode and cluster scale figures are official claims, subject to actual future delivery.

昇腾 960 官宣发布:FP4 算力翻倍、全球首个 NPO 超节点,华为确认一年一代

· 6 min read
Industry Research Team

2026 年 9 月 17 日,华为全联接大会(HC2026),华为常务董事、ICT 基础设施业务总裁汪涛正式发布昇腾 960。与 8 月数字中国峰会上的"路线图预告"不同,这次是完整的规格官宣——而且提前三个季度就绪。本文基于华为官网新闻稿及多方报道整理,逐项拆解 960 的规格、超节点与路线图含义。


1. 昇腾 960:双版本,训练先行​

昇腾 960 分为两个版本,节奏错开一个季度:

指标昇腾 960DT(训练版)昇腾 960PR(推理版)
FP8 算力2 PFLOPS(较 950 翻倍)待披露
FP4 算力4 PFLOPS待披露
显存288GB(自研 HBM)待披露
显存带宽9.6 TB/s待披露
配套超节点Atlas 860(风冷)Atlas 960(液冷)
就绪 / 上市2027 Q1 / 2027 Q22027 Q3

三个读数:

  • 提前三个季度就绪:960DT 原计划 2027 年 Q4,现在 2027 Q1 就绪、Q2 上市——与 950DT"提前上线华为云"一脉相承,路线图可信度在持续兑现;
  • 精度口径与英伟达对齐:FP8 / FP4 主口径(辅以自研 HiF8 / HiFP8),FP4 已是 2026 新卡的通用口径;
  • 288GB 自研 HBM + 9.6TB/s:显存容量与带宽同步翻倍,对训练长上下文大模型是实打实的容量红利。

单卡规格详见昇腾 960 规格页;上一代昇腾 950DT 规格页。

2. 昇腾 960 超节点:全球首个 NPO 超节点​

这是本场发布真正的"重器"——昇腾 960 超节点是全球第一个采用 NPO(近封装光学)的超节点:

指标昇腾 960 超节点
单节点规模4096 卡
系统算力8 EFLOPS FP8 / 16 EFLOPS FP4
HBM 总容量1 PB
互联 RTT低至 2μs
可用度99.8%

NPO 意味着什么:传统光模块挂在交换机面板上,信号要先出封装再转电—光;NPO 把光引擎挪到交换芯片封装近旁,大幅缩短 SerDes 距离、降低功耗与时延。华为的具体数字:

  • 5500 个自研 Hi-ONE 光引擎(业界首个量产 NPO,单引擎 7.2T);
  • 替代 4.8 万颗 800G 光模块;
  • 降低功耗 550kW 以上,同时支撑 2μs 级互联时延。

对比 8 月数字中国峰会的口径(Atlas 960 单超节点 15488 卡),官方最新口径为单超节点 4096 卡(1 个超集群可由多超节点组成)——以华为官网最新新闻稿为准。

⚠️ 华为还给出了"MFU 提升 2.75 倍""推理时延降低 70%"等收益数据,均出自华为马尔科夫实验室仿真,尚无独立第三方实测,阅读时注意口径。

3. 一年一代:970(2028)→ 980(2029)​

华为首次把"一年一代"从愿景变成官宣承诺:

年份产品状态
2026昇腾 950 系列950PR 已量产,950DT Q4 放量
2027昇腾 960本次官宣,提前就绪
2028昇腾 970HC2026 确认
2029昇腾 980HC2026 确认

支撑这一节奏的是华为提出的"韬定律":算力规格每代翻倍,访存带宽、访存容量、互联带宽同步大幅提升。

生态侧的同场数据也值得记录:CANN 外部开发者占比首次超过 61%、月活开发者 5200 人;910C 超节点部署已超 1000 套。

4. 竞争坐标:单卡有代差,系统级对打​

把 960DT 放到 2027 年的棋盘上看:

  • 对 NVIDIA Rubin(R200:288GB HBM4 / 50 PFLOPS FP4):单卡 FP4 算力约为 R200 的 8%(4 vs 50 PFLOPS),单卡代差客观存在;但 4096 卡超节点 + NPO 互联是系统级竞争——用"可交付的超大集群"对打单卡性能;
  • 对采购方:960 的价值主张是"在国产量产约束内拿到最大可用集群"。DeepSeek 拟采购 16 万颗 950DT 跑推理的订单已经证明:当供给与生态到位,头部实验室愿意主动选国产;
  • 对国产链条:950→960 的显存翻倍直接拉动国产 HBM 迭代,这是整条链的胜负手。

5. 小结​

  • 960DT:FP8 2P / FP4 4P、288GB HBM、9.6TB/s,2027 Q2 上市,提前三个季度;
  • 960 超节点:全球首个 NPO 超节点,4096 卡 / 8 EFLOPS FP8 / 1PB HBM / RTT 2μs;
  • 路线图:一年一代官宣至 2029,950、960 均提前兑现,规划可信度显著上升;
  • 看什么:2027 Q2 实际交付节奏、国产 HBM 产能、CANN 生态的第三方模型覆盖度。

相关阅读​

参考资料​

  • 华为官网:发布全球首个采用 NPO 的超节点(昇腾 960 超节点)(2026-09-17)
  • 电子工程专辑:国产 AI 芯片新突破!华为昇腾 960 芯片将提前发布
  • 环球网:华为全联接大会 2026 主题演讲报道

本文基于华为官网新闻稿与公开报道整理。MFU、时延等收益数据为华为实验室仿真口径;2028 / 2029 产品仅为路线图官宣,规格以未来发布为准。

昇腾 950 超节点 Q4 上市,中兴、新华三、曙光全线跟进:国产超节点从概念走向交付

· 6 min read
Industry Research Team

芯片追不上,系统来补。这是国产算力过去一年最清晰的战略共识。2026 年下半年,国产超节点正式从"概念发布"进入"产品密集发布、联合适配和商业部署"阶段:华为 Atlas 950 SuperPoD 真机已亮相并计划 Q4 上市,中兴、新华三、中科曙光、壁仞、沐曦等一众厂商相继入局。


1. 华为 Atlas 950:业界最大 1024 卡超节点,Q4 上市​

继 7 月 WAIC 2026 真机首秀后,华为昇腾 950 超节点的商用脚步持续加快。核心规格:

指标Atlas 950 SuperPoD
互联规模最大 1024 × 昇腾 950DT(灵衢高速互联)
AI 算力1 EFLOPS FP8/mxFP8/HiF8、2 EFLOPS FP4
全局内存256TB 统一编址,片上内存最大 1024 × 96GB @ 4.0TB/s
互联带宽单柜最大 64 × 1.68 TB/s(双向),3μs 超低 RTT
形态满配 128 计算柜 + 32 互联柜,占地约 1000㎡,8192 颗 950DT
供电散热全液冷,100kW 供电,380V AC/336V DC/240V DC
上市时间2026 年第四季度

在 WAIC 展台之外,两个数字更能说明商业化进度:昇腾 384 超节点已商用落地 750 多套,规模应用于互联网、运营商、金融、教育、医疗等行业,并且是国内唯一训练出 SOTA 模型的超节点;更远期规划中,Atlas 950 SuperCluster(超 50 万颗昇腾芯片)同样定于 2026 Q4,Atlas 960 SuperPoD(15488 卡)与 Atlas 960 SuperCluster(超 100 万卡)将于 2027 Q4 接力。

华为的路线图背后是"Tau Scaling Law"——在 EUV 光刻机受限的前提下,通过系统级扩展突破单芯片物理极限,并计划到 2031 年将晶体管密度推升至 1.4nm 级工艺水平。

2. 超节点赛道全面开花:三条技术路线​

WAIC 之后,国产超节点形成了三类清晰的技术路线:

第一类:华为全栈自研。 同时掌握昇腾芯片、灵衢互联、Atlas 硬件、CANN 软件栈和 MindSpore 框架。CANN 已全面开源:社区上线 67 个项目、开源代码超 1244 万行、月活开发者突破 3500 人;昇腾开发者总数超 400 万。

第二类:中兴、新华三的兼容路线。 中兴发布 OEX 超节点新品,依托协议标准化与接口统一,全面兼容多元 GPU 生态;新华三 UniPoD S80000 系列可从 32 卡扩展至 1024 卡、最大支持 16384 卡互联,将 Scale-up/Scale-out 网络、液冷、供电、管理和故障恢复整合进同一套架构。

第三类:国产 GPU 联合创新。 中兴联合曦智、壁仞、沐曦、燧原、天数智芯等,基于 OEX+dOCS 架构打造国产高性能 Matrix 超节点;壁仞计划发布 BR20x 系列 GPU,基于自研 BLink 2.0 互联协议实现单超节点 1024 卡 Scale-up;沐曦推出"曦景"S 系列超节点。

在更大的 Scale-out 层面,中科曙光"曙光 8000(登峰)"全国产十万卡 AI 超集群已落成并接入国家超算互联网——它不是单体超节点,而是由大量计算节点、超节点和网络存储组成的超集群,检验的是国产算力的大规模调度、应用适配和工程交付能力。

3. 为什么"超节点"成了国产算力的主战场​

根本原因在于衡量标准变了:国产算力的评价体系,正从"比较单颗芯片的峰值性能"转为"一整套系统能够调动多少有效算力"。华为副董事长徐直军对此直言:单芯片上英伟达仍领先且短期难以追赶,但在超节点和集群层面,华为有信心。

这一转向的产业逻辑:

  • 出口管制下的现实选择:单卡制程受限 → 用高速互联把更多 NPU 组织成一台"大计算机";
  • 需求侧真实拉动:万亿参数 MoE 模型的训练与推理,天然需要大规模低时延互联,超节点恰是对症下药;
  • 生态护城河前移:CANN 开源 + MindSpore 兼容 PyTorch,把 CUDA 生态的竞争从"算子覆盖"层面拉回到"开发者规模"层面。

东兴证券指出,全球超节点赛道已聚集微软、Meta、亚马逊、三大运营商、阿里、字节、腾讯、百度、曙光、中兴、浪潮、新华三、海光、沐曦等数十家厂商——格局未定,但英伟达一家独大的局面正在被谷歌 TPU、AMD Helios、华为 Atlas 三股力量同时挑战。

4. 三道关卡:从"造出来"到"卖得动"​

业内人士普遍认为,国产超节点距离真正成熟还有三道坎:

  1. 单芯片性能与单柜算力密度:受制程限制,单卡性能差距仍存,现阶段靠多柜互联和规模扩展另辟蹊径;
  2. 规模扩展效率:超节点扩展到一定规模后性能提升边际递减,通信复杂度、能耗和容错成本持续上升;
  3. 可验证的商业价值:下一阶段的比拼是有效算力利用率、单位 token 成本、模型适配广度与客户复购意愿。

一句话总结:2026 年是"中国超节点元年",2027 年见真章——届时 Atlas 950 与 Vera Rubin NVL 系列将在全球两个平行市场各自接受规模化交付的检验。


相关链接​

参考资料​


本文基于 WAIC 2026 现场报道、昇腾社区官方规格与公开研报整理。超节点性能数据均为厂商公布口径,独立第三方评测结果尚待观察。

国产 AI 芯片半年报大检阅:寒武纪净赚 23 亿、壁仞营收暴涨 20 倍、燧原登板,"抢芯大战"白热化

· 6 min read
Industry Research Team

8 月底至 9 月初,国产 AI 芯片厂商半年报密集披露,加上燧原科技 9 月 2 日启动 IPO 申购,被称为"国产四小龙"的摩尔线程、沐曦、壁仞、燧原全部完成上市,加上早已在科创板的寒武纪——国产 AI 芯片的资本市场拼图就此补齐。更重要的是,财报数字第一次集体印证了一件事:国产算力正从"能用"跨向"好用",商业化拐点已现。


1. 半年报成绩单:增长是主旋律,盈利是分水岭​

厂商2026H1 营收同比盈利状态技术路线
寒武纪59.96 亿元+108.1%归母净利 23.11 亿自研 MLU 指令集(DSA)
摩尔线程17.36 亿元+147.4%净亏 1156 万(收窄)全功能 GPU(兼容 CUDA 路线)
壁仞科技12.36 亿元+1997.6%未盈利通用 GPU
沐曦13.24 亿元+44.7%净利 6.12 亿(首次扭亏)通用 GPU
燧原11.20 亿元+279.1%未盈利DSA 专用架构(TopsRider)

几个值得注意的细节:

  • 沐曦率先跨过盈利线:8 月 31 日披露的半年报显示净利润 6.12 亿元、同比扭亏(上年同期亏损 1.86 亿);不过扣非净利润仍为 -4900 万(亏损收窄 75.8%)——含金量仍在爬坡。
  • 壁仞低基数暴增:近 20 倍的同比增速来自上年同期极低的收入基数,但 12.36 亿的绝对体量已与燧原、沐曦同量级,第二梯队座次重新洗牌。
  • 研发强度惊人:沐曦研发费用占营收 39.7%、摩尔线程 44.3%、壁仞高达 65%——高研发投入是全员未完全盈利的根本原因,也是未来竞争力的来源。
  • 摩尔线程毛利率承压:从 78% 降至 57%,成本增速(+245%)远超营收增速(+147%)——大规模量产期的品控与爬坡成本开始显现。

2. 燧原登板:四小龙资本拼图补齐​

燧原科技 9 月 2 日启动公开申购,发行 4303 万股新股(约占上市后总股本 10%),募资目标 60 亿元,采用 DSA 专用架构 + 自研 TopsRider 软件平台(不兼容 CUDA)。至此:

  • 科创板:寒武纪(2020 年)、摩尔线程、沐曦、壁仞(2026 年)
  • 燧原:2026 年 9 月完成上市

国产 AI 芯片第一梯队全部进入公开市场,融资通道打开后,研发投入的"军备竞赛"将进一步升级。

3. 需求端:百万卡缺口,订单排到三年后​

财报爆发的另一面,是需求端的极度饥渴:

  • 产能缺口:行业调研显示,2026 年国产 AI 芯片需求规模约 400 万颗,实际交付约 300 万颗,存在百万级缺口;
  • 订单周期:内蒙古乌兰察布远景星河基地(规划支撑百万卡级并行算力)表示"在手订单已排到三年后";
  • 供给创新:算力紧缺催生"集装箱式算力中心"——20/40 英尺标准集装箱为载体,插电接水即可在 24 小时内完成部署,把传统"一年工期"压缩到"一天上线";
  • 市场空间:IDC 数据显示 2025 年中国 AI 加速卡出货约 400 万片,国产占约 41%(165 万片);CIC 预测中国 AI 加速器市场 2028 年将超万亿元,其中国产方案份额约 90%。

4. 两种路线的两种赌注​

市场研究机构预测 2026 年中国高端 AI 芯片市场中,国产方案份额将接近 90%,其中华为约 62%、寒武纪约 14%,剩余由摩尔线程、沐曦、壁仞等共同占据。而寒武纪与摩尔线程这对"双龙头",正在押注两种截然不同的未来:

寒武纪:用现金买确定性。 半年报资产负债表上,存货 82.48 亿 + 预付款 29.14 亿,两项合计 111.6 亿、占总资产 61%——在实体清单限制下"产能即订单",提前锁定未来 12-18 个月的晶圆产能,代价是经营现金流净额同比下滑 66%,且上半年已计提存货跌价损失 3.97 亿。底气来自订单确定性:57 亿元股权激励计划的考核目标是 2026 年营收不低于 135 亿、2026-2028 年累计不低于 1000 亿。

摩尔线程:用亏损买时间。 全功能 GPU 的"大而全"路线(AI 计算 + 图形渲染 + 物理仿真 + 视频编解码)前期投入巨大,需要在现金流耗尽之前跨过规模化的门槛。上半年亏损已收窄至千万级,IPO 募资到位后,时间窗口正在变宽。

5. 软件生态:被忽视的胜负手​

需求外溢(英伟达 H20 系列在中国市场遇冷)推动 DeepSeek、智谱 GLM、月之暗面 Kimi 等头部模型纷纷适配国产加速器(昇腾、沐曦、摩尔线程等),这给了国产芯片难得的"真实负载打磨机会"。但正如行业共识:芯片可以三年一代,生态需要十年之功。寒武纪自研指令集的封闭高效、摩尔线程 MUSA 对 CUDA 生态的兼容路线、燧原 DSA 的专用极致——哪条路能沉淀出真正的开发者生态,才是决定 2030 年格局的关键变量。


相关链接​

参考资料​


本文基于上市公司半年报、Pandaily 与央视财经等公开报道整理。财务数据为公司披露口径,市场份额为研究机构预测,不构成任何投资建议。

Domestic Big Three 2026 H2: Localization Rate Crosses 40% Toward 60%, Ascend 960 Roadmap, MLU690 and S5000 Ecosystems Ramp Up

· 6 min read
Industry Research Team

In 2026, China's AI chip market landscape has shifted from "NVIDIA unipolar dominance" to "overseas vendors leading, domestic multi-route catch-up." According to industry research, China's overall AI accelerator market was ~4M units in 2025, of which 1.65M were domestic, with share first breaking 40%; as products iterate and fabs follow up, the localization rate is expected to rise to 60%-70% by 2027. This article focuses on the latest H2 2026 progress of Huawei Ascend, Cambricon, and Moore Threads — the domestic "Big Three."


1. Huawei Ascend: 950 Capacity Fully Booked, 960 Roadmap Unveiled​

Ascend's core advantage is "architecture + full-stack ecosystem synergy," with ~800K units shipped in 2025, capturing 50% of the total domestic vendor share. The product iteration cadence is clear:

TimeProductNote
2025 Q1Ascend 910CMain transitional model
2026 Q1Ascend 950PRInference flagship
2026 Q4 (planned)Ascend 950DTTraining flagship, drives domestic HBM iteration
2027-2028Ascend 960 / 970Roadmap products

950 series capacity has entered a "fully booked" state: 950PR entered mass production in April 2026; June monthly capacity jumped to 500K-600K units (nearly 10x MoM), with a full-year target of 1.2M units at 100% certainty; ByteDance locked in 350K units for $5.6B, while Tencent / Alibaba / Baidu combined locked in 400K units.

Ascend 960 roadmap specs (per roadmap disclosure):

MetricAscend 960
ArchitectureAscend 6th gen (Da Vinci v6)
FP8 compute~4 PFLOPS
Memory288GB
Memory bandwidth9.6 TB/s
Super-nodeAtlas 960 SuperPoD, 15,488 cards, Lingqu optical-electrical converged bus
Debut2027 Q4 (roadmap)

The previous-gen Ascend 384 super-node has cumulatively shipped over 750 sets, deployed across 20+ industries including internet, operators, finance, education, and healthcare — Huawei calls it "the only domestic super-node that has trained a SOTA model."


2. Cambricon MLU690: H2 Mass Production, Entering ByteDance Bidding Window​

Cambricon is the core domestic compute leader in the absence of an Ascend IPO, with the technology gap continuously narrowing:

  • Siyuan 590 (7nm): Performance equivalent to 80% of A100, already supports DeepSeek, continuously adapting to mainstream large models like Qwen 3 and GLM
  • Siyuan 690 series: Will enter mass production in H2 2026, expected to achieve order scale-up during ByteDance's H2 bidding window
  • Revenue certainty: Equity incentive targets show >100% revenue growth for the next 3 years: 2026 revenue target 13.5B RMB, 2027 27B RMB, 2028 60B RMB

Cambricon fully benefits from the industry dividend of "domestic CSP capex + full adaptation of domestic large models and domestic chips," making it the most direct elasticity play on rising localization rate.


3. Moore Threads MTT S5000: Full-Function GPU + Ecosystem Breakthrough​

Moore Threads takes a differentiated "full-function GPU" route, with the flagship MTT S5000 based on the 4th-gen "Pinghu" MUSA architecture:

MetricMTT S5000
Dense AI compute1000 TFLOPS
Memory80GB
Memory bandwidth1.6 TB/s
Inter-card interconnect784 GB/s
PrecisionFP8 to FP64 full precision (training + inference)
SecurityFirst batch to pass national "Safe and Reliable Evaluation" (Level I)

Its engineering capability is verified: the Kuae (KUAE) intelligent computing cluster based on S5000 achieves 95% training linear scaling efficiency, with compute efficiency loss within 5% at ten-thousand-card scale; supports checkpoint-resume training with effective training time ratio >90%; and has trained a MoE-236B base model with >25 trillion tokens of corpus from scratch.

The ecosystem is Moore Threads' deepest moat: MUSA has achieved 100% core math library compatibility, 3000+ PyTorch operator compatibility, covers 55 categories of core AI operators, has official vLLM and SGLang support, Day-0 adaptation of mainstream models, and 800K+ developers. Its PD heterogeneous-disaggregation solution achieves equivalent replacement of international high-end GPUs at a 2:1 ratio with S5000, significantly reducing inference cost.

The 5th-gen "Huagang" architecture (released 2025-12) supports FP4 to FP64 full precision, with 50% higher compute density and 10x better energy efficiency than the previous gen, supporting 100K+ card clusters; cumulative R&D investment in the "Huashan" (train-infer integrated) and "Lushan" (graphics rendering) new chips based on this architecture exceeds 900M RMB.


4. Software Ecosystem Decides: Day-0 Adaptation Becomes Routine​

Beyond hardware, software ecosystem realization is the watershed for domestic compute in 2026:

  • Huawei's CANN heterogeneous computing architecture and MindSeries suite are fully open-sourced, with the community incubating 67 projects, 12.44M+ lines of code, and 3,500+ monthly active developers
  • The "release-and-adapt" closed loop between domestic large models and domestic chips has basically formed: Tencent Hunyuan T3 (295B), DeepSeek-V4, and GLM-5.2 all completed Day-0 adaptation
  • 2026 is regarded as the "first year of domestic super-nodes"; Huatai Securities estimates China's super-node architecture market will reach 341.4B RMB by 2028, with a 2026-2028 CAGR of 194%

5. Industry Judgment: From "Can It Be Built" to "Can It Be Used Well"​

The domestic Big Three are converging along three paths:

  1. Huawei: Locks government/enterprise and internet big customers with super-node system-level capability + full-stack software
  2. Cambricon: Impacts the revenue inflection point by narrowing the training-side gap + scaling up via big-customer bidding
  3. Moore Threads: Covers cloud-edge-end full scenarios with full-function GPU generality + mature CUDA-compatible ecosystem

The common shortcoming of all three remains advanced process and HBM supply — precisely the core link of overseas controls. But as domestic HBM iterates and fabs follow up, a realistic path to 60%-70% localization by 2027 exists.

References​


This article is compiled from public industry research, broker views, and corporate announcements as of August 2026. Some shipment and market-share figures are third-party estimates, not officially confirmed data.

WAIC 2026 Recap: Huawei Atlas 950 SuperPoD Live Hardware Wins SAIL Grand Award, Domestic Compute Enters the "System-Level" Showdown

· 5 min read
Industry Research Team

The 2026 World Artificial Intelligence Conference (WAIC) was held July 17-20, 2026 at the Shanghai World Expo Center, themed "Intelligent Partners, Creating the Future Together." Over 1,100 companies showcased 3,000+ exhibits, with 300+ products debuting globally. For the compute-card industry, this concentrated review of domestic compute sent a clear signal: the competitive main line is shifting from "single-chip peak compute" to "SuperNode system-level effective compute."

1. Huawei Atlas 950 SuperPoD: live debut, wins SAIL grand award​

Huawei's Atlas 950 SuperPoD live hardware made its first public appearance at WAIC 2026, on-site carrying 16 compute cabinets with 1,024 Ascend cards total. With three system-level innovations — "ultra-wide bandwidth, ultra-low latency, unified memory addressing" — it stood out from hundreds of domestic and international entries to win the conference's top honor, the SAIL (Super AI Leader) Award.

Core parameters (confirmed on-site at WAIC)​

MetricAtlas 950 SuperPoD
Exhibited scale16 compute cabinets / 1,024 Ascend cards
Max interconnect scale8,192 Ascend NPU cards fully interconnected (full config)
Interconnect protocolHuawei in-house "Lingqu" (UnifiedBus) 2.0
Total compute1 EFLOPS FP8 / 2 EFLOPS FP4 (1,024 cards); full 8,192-card ~8 EFLOPS FP8
Unified memory256 TB globally unified memory address space
Interconnect latency3 μs ultra-low RTT; TB-level NPU interconnect bandwidth
Full config128 compute cabinets + 32 interconnect cabinets = 160 cabinets, ~1000㎡, carrying 8,192 Ascend 950DT
LaunchFull config planned for Q4 2026
CoolingFully liquid-cooled blind-plug architecture

Huawei disclosed for the first time: the previous-gen Ascend 384 SuperNode has cumulatively shipped 750+ units commercially, deployed across 20+ industries including internet, operators, finance, education, healthcare, transportation, and manufacturing, calling it "the only domestic SuperNode that has trained SOTA models."

2. Software ecosystem: CANN fully open-sourced, developers at scale​

Beyond hardware, Huawei highlighted open-source software ecosystem progress:

  • CANN heterogeneous compute architecture and MindSeries base software suite were fully open-sourced end of 2025;
  • The CANN open-source community has incubated 67 projects, 12.44M+ lines of code, with 3,500+ monthly active developers;
  • Huawei has co-developed 7,000+ solutions with 3,000+ industry partners, serving 2,000+ core government/enterprise customers;
  • WAIC showcased 60+ real business scenarios, 20+ benchmark cases, covering the full chain from technology breakthrough to scaled commercial deployment.

3. Domestic chips' Day-0 adaptation becomes routine​

On July 6, 2026, Tencent released the MoE model Hunyuan T3 (295B parameters, 256K context); domestic chips rapidly completed Day-0 adaptation:

VendorChipAdaptation status
Moore ThreadsMTT S5000Completed rapid Hunyuan T3 adaptation (previously adapted DeepSeek-V4, GLM-5.2)
MetaXXiyun C seriesIn-house MXMACA stack first to full-chain Day-0 adaptation, zero-code deployment

Moore Threads also showcased the MTT C256 SuperNode (first-of-its-kind single-layer Scale-up 256-card full interconnect, sub-microsecond latency) and three AI-factory solutions — "model training factory / token production factory / agent production factory."

4. More domestic compute debut highlights​

Vendor / productHighlight
Orient AlphaChip DF1000World's first "software-defined + near-memory computing" 3D chip, interconnect pitch compressed to sub-micron
ZhongHao XinYing "Xuyu"Fully in-house next-gen TPU-architecture AI-specific chip, with Taize 2.0 server
Enflame × IluvatarDomestic high-performance Matrix SuperNode based on OEX+dOCS architecture, shortlisted for the conference "Excellent AI Leader Award"
Rongming MicroelectronicsAdvancing next-gen VPU, evolving from video processing to "visual-agent compute base"

The domestic AI chip lineup also included Moore Threads, MetaX, Enflame, Houmo, Cixiong, Suaneng, SemiDrive, Phytium, Aixin, Iluvatar, and others.

Industry interpretation: from "can it be built" to "is it used well"​

WAIC 2026 reflects a fundamental shift in the competitive stage of domestic AI chips:

  1. SuperNode becomes the main battlefield: beyond single-chip performance, system-level capabilities — "inter-chip interconnect + cluster scale + cooling" — become the breakthrough key. Huawei Lingqu and Enflame/Iluvatar OEX are both pushing here. Huatai Securities defines 2026 as the "first year of domestic SuperNodes," estimating China's SuperNode architecture market could reach ¥341.4B by 2028, with 2026-2028 CAGR of 194%.
  2. Software ecosystem delivers: Day-0 adaptation has gone from slogan to routine; the "launch-and-adapt" closed loop between domestic large models (DeepSeek-V4, GLM-5.2, Hunyuan T3) and domestic chips is essentially formed.
  3. Demand-side endorsement: China Mobile earlier released its 2026-2027 AI SuperNode centralized procurement announcement — about 6,208 cards, over ¥2B — accelerating domestic SuperNode scaled commercialization.

References​


This article is compiled from WAIC 2026 (July 17-20) on-site and official disclosures, and will continuously track the 950 SuperNode Q4 launch.

Domestic GPU IPO Wave: The "Four Little Dragons" Assemble on Capital Markets, Moore Threads MTT S5000 Benchmarks Against H100

· 5 min read
Industry Research Team

From December 2025 to July 2026 — just half a year — at least 6 AI chip companies have listed or are about to list on capital markets. Together with already-listed Cambricon, Hygon, and Iluvatar, the domestic GPU corps' total market cap is approaching ¥2 trillion. This marks the critical climb from domestic GPUs being "usable" to "useful."

1. The "Four Little Dragons" assemble on capital markets​

CompanyListing statusRaise / issue priceSponsor
Moore ThreadsListed (STAR Market sh688795, 2025-12-05)Issue price ¥114.28, raised ¥8BCITIC Securities
MetaXIPO accepted (2026-06-30)¥3.904B (total investment ¥5B)Huatai United
EnflamePassed review (2026-06-15)¥6B—
BirenHKEX / sprinting——

Already-listed camp: Cambricon (sh688256, STAR Market 2020-07-20), Hygon, Iluvatar (HKEX). Moore Threads turned a book profit of ¥29.35M in Q1; MetaX narrowed losses 57.7% and gave a 2026 breakeven timeline.

2. Moore Threads MTT S5000: benchmarking against H100​

Moore Threads announced its flagship AI train+inference GPU MTT S5000 successfully completed full-pipeline adaptation validation of Zhipu's new-generation large model GLM-5 — measured performance "breaks the domestic compute ceiling":

MetricMTT S5000
Architecture4th-gen "Pinghu" architecture
FP8 compute1 PFLOPS (1,000 TFLOPS)
Memory bandwidth1.6 TB/s
PositioningFull-function train+inference GPU, benchmarks against NVIDIA H100
ProductionMass-produced; clusters online supporting trillion-parameter training

Deployment validation: jointly completed full-pipeline training of embodied-brain model RoboBrain 2.5 with BAAI; partnered with SiliconFlow for high-performance DeepSeek-V3 inference, single-card speed near international top products. IPO funds go to three directions: next-gen AI train+inference chip, next-gen graphics chip, next-gen AI SoC chip.

WAIC 2026 new progress: Moore Threads showcased the MTT C256 SuperNode (first-of-its-kind single-layer Scale-up 256-card full interconnect, sub-microsecond latency) and three AI-factory solutions — "model training factory / token production factory / agent production factory"; the company pre-announced H1 2026 revenue of ¥1.65B-1.75B, up 135%-149% YoY.

3. Cambricon: dual flagships MLU590/690​

ChipProcessComputeMemoryCustomer / status
MLU590 (思元590)7nm ChipletINT8 512 TOPS / FP16 345 TFLOPS96 GB HBM2eByteDance inference mainstay, ~80% of A100 overall, mass shipments early 2026
MLU690 (思元690)5nm-class (SMIC N+2)FP16 700+ TFLOPS / INT8 2800+ TOPS196 GB HBM3 (3.35 TB/s)Dual-die packaging, MLU-Link 890 Gbps; ~70% of H100 (80-90% pure inference); ByteDance largest customer, mass production early 2026

Cambricon is the only domestic AI chip vendor with a "unified edge-cloud architecture" — one MLU instruction set spans 思元 220 (edge) → 370 (border) → 590/690 (cloud), with one NeuWare toolchain across compute tiers.

Capital and performance double explosion: Cambricon's total market cap exceeded ¥1 trillion on June 30, 2026, becoming the STAR Market's first "trillion-yuan stock," up 75%+ YTD. On performance, Q1 2026 revenue ¥2.885B (+160% YoY), deducted net profit ¥934M; full-year 2025 revenue ¥6.497B (+453% YoY), net profit attributable to parent ¥2.059B, ending long-term losses. ByteDance has cumulatively deployed over 100k 思元 590/690, its largest customer.

4. DeepSeek-V4 effect: changing the expectation coordinate system​

On April 24, 2026, DeepSeek released the trillion-parameter flagship DeepSeek-V4. Unlike a year earlier when V3's launch sparked debate over "can domestic chips even run large models," this time multiple domestic chips — Huawei Ascend, Cambricon, Hygon, MetaX, Moore Threads, Kunlun, T-Head, Iluvatar — completed adaptation on launch day.

The evaluation coordinate system is shifting: from "what percentage of NVIDIA's same-generation product performance" to "can it carry the real workloads of top-tier large models."

Industry interpretation​

  1. Capital ammunition in place: dense IPOs provide ample funding for domestic GPU R&D iteration and capacity expansion, moving from "technology breakthrough" to "commercial virtuous cycle."
  2. Train+inference becomes the mainstream route: Moore Threads takes the full-function GPU route (graphics+AI+general compute), differentiating from Huawei Ascend's "AI-focused."
  3. Software ecosystem is the decider: Day-0 adaptation and the maturity of unified software stacks (MUSA / NeuWare / MXMACA) are replacing raw peak compute as the core yardstick of domestic GPU "usability."

References​


This article continuously tracks the domestic GPU listing process and product iteration.

Huawei Ascend 950 Series Capacity & Orders Deep Dive: 950PR Monthly Capacity Jumps 10×, ByteDance Locks In 350k Units for $5.6B

· 4 min read
Industry Research Team

The Ascend 950 series (950PR inference / 950DT training) has become the core supply of domestic AI compute. Per multiple brokerages and industry research, 950 series capacity is 100% booked with scarce spot supply; the full-year 1.2M-unit target is "100% certain," with expectations of an upward revision to 1.5M. This article summarizes capacity and order data as of July 2026.

1. Capacity pace: ~10× MoM jump in June​

Time950PR monthly capacityNotes
May 202650k-60k unitsNear full production
June 2026500k-600k units~10× MoM; SMIC, Hua Hong tier-1 suppliers on overtime
Q3 2026 (est.)700k-800k unitsPer month
Full-year 2026 target1.2M unitsUpward revision to 1.5M expected

Supply chain delivery is tight: high-speed backplanes and liquid-cooling connectors' lead time stretched from 2 weeks to 6-8 weeks; orders are booked into 2027.

2. Order structure: top cloud providers + operators + overseas​

CustomerLocked volumeAmount / Notes
ByteDance350k 950PR$5.6B, concentrated delivery from Q3 2026
Tencent / Alibaba / Baidu~250k 950PR + 150k 950DTCombined ~400k units
Three major operators200k+ unitsCentralized procurement, for intelligent compute centers and AI private networks
OverseasSouth Korea 2,000 units, Malaysia 3,000 servers, Russia ten-thousand-card clusterFrom pilot to commercial

3. Shipment forecast: firmly #1 domestic​

Per CCA (Kezhi) Consulting estimates:

Metric20252026 (forecast)
Huawei Ascend total shipments812k cards1.026M cards
Of which 950PR—~800k units
Of which 950DT—~100k-200k units

Huawei has completed the product transition from the 910 series to the 950 series. The internet industry has become Ascend's largest application market; competitive advantage is extending from single-hardware performance to software ecosystem and system capabilities.

4. Going overseas: formal South Korea entry in Q4​

Per Korean media ETNews, Huawei plans Q4 2026 to formally enter the South Korean market with the Ascend series and Atlas 950 SuperPod:

  • Local distributor agreements signed; two channel partners including SK Shieldus selected
  • Main products: 950PR (mass-produced and delivered since April) and 950DT (launched Q4)
  • Official line: 950PR inference performance is 2.87× that of H20, priced at about 1/4 of it

5. WAIC 2026: 1024-card live debut confirmed​

At WAIC 2026 (July 17-20, Shanghai), Huawei's Atlas 950 SuperPoD live hardware made its first public appearance — a 16 compute-cabinet, 1,024 Ascend-card scale — and won the conference's top honor, the SAIL Award:

  • Core metrics: total compute 1 EFLOPS FP8 / 2 EFLOPS FP4, 256 TB globally unified memory addressing, Lingqu 2.0 interconnect, 3 μs ultra-low RTT latency
  • Full configuration: 128 compute cabinets + 32 interconnect cabinets = 160 cabinets, ~1000㎡ footprint, carrying 8,192 Ascend 950DT, planned Q4 2026 launch
  • Commercial foundation: previous-gen 384 SuperNode has cumulatively shipped 750+ units, deployed in 20+ industries
  • Software ecosystem: CANN fully open-sourced end of 2025; community incubated 67 projects, 12.44M+ lines of code, 3,500+ monthly active developers

WAIC's debut confirmed the 950 series' "SuperNode-first" product logic: beyond single-card compute, system-level effective compute (interconnect bandwidth + unified memory + low latency) is the key dimension for domestic compute to benchmark against international flagships.

Ascend roadmap recap​

ProductPositioningKey metrics (official roadmap)
950PRInference1 PFLOPS (FP8) / 2 PFLOPS (FP4), 2 TB/s interconnect
950DTTrainingSuperNode core, launched Q4
960Train/inference2 PFLOPS (FP8) / 4 PFLOPS
970Next-genIn planning

Industry interpretation​

  1. Domestic substitution moves from inference to training: 950PR (inference) ramps first, 950DT (training) follows in Q4, combined with the Atlas 950 SuperPoD ten-thousand-card interconnect — domestic compute now has the complete "training substitution" puzzle for the first time.
  2. Capacity is the biggest variable: order certainty is extremely high, but SMIC/Hua Hong advanced-process capacity, HBM supply, and advanced packaging remain ramp bottlenecks — the root of "scarce spot supply."
  3. Going overseas opens a second growth curve: bulk procurement from South Korea, Malaysia, Russia, and Latin America marks domestic compute's shift from "internal circulation" to "external circulation."

References​


Data in this article is based on official and major brokerage research; capacity/orders are dynamic figures and will be continuously updated.