Skip to main content

8 posts tagged with "AI Infrastructure"

Data center, power and networking infrastructure for AI

View all tags

Intelligent Compute to Hit 9,800 EFLOPS by 2030: Five Hard Targets in the ICT Industry "15th Five-Year" Plan

· 6 min read
Industry Research Team

This article is based on the MIIT "15th Five-Year" Plan for ICT Industry Development (issued in September 2026) and official interpretations by the People's Post and Telegraph News, The Paper, Huaxia Times, and other outlets.

In early September, the MIIT issued the "15th Five-Year" Plan for ICT Industry Development, drawing a roadmap for the ICT industry over the next five years with 13 major indicators and 26 key tasks. For the AI compute industry, this is the most substantive policy document — we have picked out five hard targets and break down their industry implications one by one.

Target 1: Intelligent Compute from 1,590 to 9,800 EFLOPS, More Than 5x Growth in Five Years​

This is the standout number in the entire plan. For reference:

  • As of the end of June 2026, national intelligent compute capacity had reached 2,185 EFLOPS, up 177% year over year — the first half alone overshot the year-end 2025 target (1,590 EFLOPS)
  • The country has built 42 10,000-card-class intelligent computing clusters; in July 2026, the first fully domestic 100,000-card-class AI compute cluster, the Sugon 8000, was completed in Zhengzhou
  • Huawei's rotating chairman Wang Tao's assessment: super-node clusters of more than 100,000 cards will become basic configuration by 2027

Going from 2,185 to 9,800 means that over the next four years the country must build 3.5 times the current installed base of intelligent computing facilities. The plan also calls for "orderly deployment of 10,000-card, 100,000-card and larger intelligent computing clusters" and "stepping up efforts to adapt domestic compute chips" — this is the most certain demand base for domestic AI chips over the next five years.

Target 2: Cumulative Information Infrastructure Investment of RMB 3.8 Trillion​

Some have compared this with the "14th Five-Year" figure of RMB 3.7 trillion and concluded that "investment has peaked." Official interpretations explicitly reject this claim: incremental capital is shifting from "scale expansion" to "quality-and-efficiency gains," precisely targeted at intelligent computing infrastructure, 10G optical networks, 6G, and other new tracks. The industry chain pull effect is changing accordingly — AI servers, high-speed optical interconnects, domestic compute chips, and new smart terminals were named as beneficiaries across the whole chain.

Target 3: PUE of New Large Compute Facilities Reduced Below 1.2​

At the end of the "14th Five-Year" period this figure was 1.25, and it must fall another 0.05 within five years — it doesn't sound like much, but with AI servers running at high load year-round and per-rack power density generally exceeding 20kW, every point of PUE reduction is a hard fight. The technical path laid out in the plan is very clear: liquid cooling.

  • The "15th Five-Year" Plan for ICT Industry Development: guide compute facilities to adopt high-efficiency energy-saving equipment and advanced technologies such as liquid cooling
  • The "15th Five-Year" Plan for Electronic Information Manufacturing Development (jointly issued by the MIIT and the NDRC on September 15): lists liquid cooling alongside high-bandwidth memory pooling and all-optical switching as key technologies to be broken through

Industry-side data confirms the trend: according to Omdia, liquid cooling's share of the global data center cooling market climbed from about 20% in 2024 to about 37% in 2025, with penetration expected to exceed 45% by 2029. The ceiling on compute expansion is shifting from chip supply to power supply, and the energy-efficiency constraint of "compute up, energy down" will be the norm for the next five years.

Target 4: Advanced Storage Capacity of 1,700 EB, More Than 2x Growth​

While intelligent compute grows 5x, advanced storage grows more than 2x — storage-compute coordination is given equal weight. The logic: training checkpoints for large models, and the intermediate states and contextual memory of agent inference, are all stored in layers within high-performance storage. As of the end of June, national storage capacity totaled about 2,021 EB, of which advanced storage accounted for about 32%; the 1,700 EB advanced storage target for 2030 means structural upgrading matters more than total-volume growth.

A direct implication for chip selection: KV cache and long context are eating the memory budget — when evaluating AI servers, memory capacity and memory bandwidth should carry more weight than peak compute.

Target 5: The Agent Interconnection Network Enters the Plan for the First Time​

This is the most forward-looking part of the plan: the "agent interconnection network" gets its own dedicated column, deploying four areas of work — building an agent network identifier system, accelerating the construction and application of agent network infrastructure, promoting global interconnection, and establishing a space governance system. It also explicitly states "launching 6G commercial use in a timely manner."

As AI shifts from "applications for people" to "agent infrastructure," the role of the communications network upgrades from connecting people to connecting agents — providing a national-level narrative for the distributed deployment of inference compute (edge inference, compute scheduling).

Three Judgments for Compute Practitioners​

JudgmentBasis
Domestic compute demand is highly certainThe plan explicitly states "stepping up efforts to adapt domestic compute chips" + the 100,000-card cluster build cycle has begun
Energy-efficiency targets become hard siting constraintsPUE below 1.2 + liquid cooling named as a key technology; high-density liquid cooling solutions take priority
Inference compute sinks to the edge on demand"Deploy inference compute facilities as needed for scenarios"; edge and regional compute centers enjoy a policy window

Under a 5x compute expansion target, what has always been scarce is not planning but chips, power, and delivery capability. For buyers, the supply window remains tight; for solution selection, we recommend using the TCO Calculator to convert PUE differences into electricity costs — the gap between 1.25 and 1.2 amounts to tens of millions on the electricity bill of a 10,000-card cluster.

Summary​

The "15th Five-Year" Plan writes intelligent compute into a national-level project: 5x compute, 2x storage, RMB 3.8 trillion in investment, PUE 1.2, and the agent internet. Looking back five years from now, the wave of domestic 100,000-card clusters in the second half of 2026 may well prove to be the starting point of this curve.

(Plan data is cited from MIIT documents and official media interpretations; market data is cited from statistics by CAICT, Omdia, and other institutions.)

GPU Rental Rates Climb Again: Nebius Raises H100/B200 On-Demand Prices About 20% From October, and the Build-vs-Rent Balance Is Shifting

· 5 min read
Industry Research Team

According to ZhiXun (media) reporting in late September, Nebius announced that on-demand rates for H100 / H200 / B200 will rise by an average of about 20% starting October 1. The same month, Cailian Press reported that as of August 11 CoreWeave's contracted power had grown to 4.2 GW and that it would "keep signing new compute at higher prices". This is the Nth consecutive price increase in the rental market this AI cycle — and for buyers, the TCO balance between building and renting is shifting. This article works through the math using the site's TCO calculator methodology.


1. Three Forces Behind the Price Hikes​

  • Exploding agentic inference demand: from MLPerf v6.1 to SemiAnalysis AgentX, agentic workloads have become the main engine of inference growth — token consumption is orders of magnitude beyond chat scenarios, and inference compute has gone from "enough" to "never enough";
  • Supply premium on the new platform generation: Vera Rubin NVL72 has debuted on CoreWeave, and early-ramp pricing on new platforms is naturally elevated — which also raises the anchor for renewal prices on the previous generation (H100/H200/B200);
  • Power has become a hard constraint: CoreWeave's 4.2GW of contracted power shows the data center supply bottleneck has shifted from GPUs to power and racks; enterprises holding power contracts have no reason to cut prices.

2. What a 20% Rent Increase Means for TCO​

The site's TCO calculator splits total cost of ownership into five parts: purchase + energy + operations + network + discounting. Rent increases don't enter that model directly, but they change the "rental baseline" — the opportunity cost a build option must beat:

ScenarioBefore 20% Rent HikeAfter 20% Hike
Short-term elastic workloads (under 1 year)Renting winsRenting still wins (build deployment lead times can't catch up)
Steady inference workloads (2-3 years)Crossover zoneBuild starts to win (opportunity cost rises)
Training clusters (3 years+)Build winsBuild's advantage widens

Rules of thumb:

  • For workloads running 24/7 at full load for 2 years or more, a 20% rent increase is usually enough for building (even at a 1.5x purchase premium and with a PUE 1.3 energy model) to overtake renting within 30 months;
  • Tidal workloads (daytime peaks, idle nights) should still rent — under the assumption that a build pays a continuous 15% idle power draw, building almost never breaks even below roughly 50% utilization;
  • A hybrid strategy (owned baseline + rented peaks) is least sensitive to rent increases and is the robust play in the current interest-rate and power-price environment.

You can re-derive these conclusions in the TCO calculator: enter the "cloud rental unit price" at +20% and compare scenarios across purchase price, electricity price, and PUE. The default discount rate is 8%; purchase costs are not discounted, and annual electricity is discounted by 1.08^-y.

3. Pricing Implications for Domestic Compute​

Rent increases have a layer of impact on domestic AI chip purchasing decisions that is easy to overlook:

  • The rental market for domestic chips is not yet mature: compute from Ascend / Cambricon / Moore Threads is delivered mostly as appliances, full racks, and intelligent computing center allocations; public on-demand rental prices are scarce — so procurement comparisons naturally lean toward "build / full-system" models;
  • Relative TCO shift: a 20% NVIDIA rent increase raises the "relative attractiveness" of every domestic alternative, especially for inference workloads (actual market transaction prices for 910C / 950PR / P800 on the primary market are not public, but the full-system price gap versus dollar-priced cards is widening);
  • Supply cadence: DeepSeek betting on Ascend training and rumors of 160,000-unit 950DT purchases — top demand eats capacity first, and queue times for later buyers are lengthening. Deciding early has value in itself.

4. Takeaways​

  • Nebius raises H100/H200/B200 on-demand rates +20% from October; CoreWeave's contracted power at 4.2GW keeps locking in volume at high prices — rental market supply and demand remain tight;
  • Agentic inference + Rubin supply premium + power constraints: three forces pushing rents up, with no near-term reversal in sight;
  • Decision framework: steady workloads 2 years+ → the build window opens; tidal workloads → keep renting; hybrid strategies are most price-hike resistant;
  • What to watch: whether other clouds follow with price moves, formal Rubin rental pricing, and whether a public rental market for domestic compute emerges.

Further Reading​

References​

  • ZhiXun (media): Nebius raises AI chip rental prices; from October 1, H100/H200/B200 on-demand rates rise about 20% on average (2026-09-21)
  • Cailian Press / East Money: CoreWeave deploys multi-rack Vera Rubin NVL72 cluster, contracted power grows to 4.2 GW (2026-09)

This article is compiled from public reports; TCO conclusions depend on assumptions such as utilization, electricity price, and discount rate. Please re-compute with your own parameters using the TCO calculator.

Alibaba Zhenwu V900 Unveiled at Apsara Conference: 3x M890 Performance, 216GB Memory, 500,000-Card Cluster, Mass Production in 2027 Q1

· 5 min read
Industry Research Team

Just five days after Huawei officially announced the Ascend 960 at HUAWEI CONNECT on September 17, Alibaba unveiled its new unified training-and-inference AI chip, Zhenwu V900, at the Apsara Conference in Hangzhou on September 22 — which Alibaba calls "the most powerful Chinese self-developed AI chip in terms of compute performance to date." This article is compiled from Alibaba's official announcements and reports from Sina Tech, C114, Huanqiu.com, and other sources.


1. Single Chip: 3x M890, 216GB + 1200GB/s​

MetricZhenwu V900 (this launch)Zhenwu M890 (previous generation)
Performance3x Zhenwu M890 (official figure; absolute value undisclosed)FP16 600 TFLOPS (as catalogued on this site)
Memory216 GB144 GB (HBM3)
Chip-to-chip interconnect1200 GB/s—
PrecisionNative FP8 / FP4 (including high-precision training)—
PositioningUnified training and inference (trillion-parameter-scale training + low-precision inference)Unified training and inference
Mass production2027 Q1, scaled deployment in Alibaba Cloud data centersAlready deployed at scale

Three takeaways:

  • Memory crosses into the 200GB+ tier: 216GB puts it in the same capacity class as the Ascend 960DT (288GB self-developed HBM) and NVIDIA Rubin (288GB HBM4), leveling the capacity threshold for long-context and very large MoE models;
  • 1200GB/s chip-to-chip interconnect: laying the foundation for the supernode's "memory-semantics interconnect" — the chip-level prerequisite that pairs with ICN Switch;
  • Precision coverage across all scenarios: officially described as "high-precision training, low-precision and ultra-low-precision inference across all scenarios," with native FP8/FP4 support aligned with the common spec of 2026's new cards.

The absolute single-chip compute figure has not been officially disclosed; this site records it using the official relative figure of "3x M890." See the Zhenwu M890 spec page for M890 details.

2. System Level: ICN Switch + Panjiu Supernode, a 500,000-Card Single Cluster​

V900's real selling point is not the single chip but system-level collaboration:

  • ICN Switch self-developed interconnect chip: once connected, the supernode gains native memory semantics and unified memory addressing, with a thousand cards interconnected at full bandwidth — over a thousand V900s can "work together like a single super chip";
  • Panjiu supernode server: integrates V900 (compute) + ICN Switch (interconnect) + Panmai intelligent NIC (network) + Zhenyue SSD controller (storage), a fully self-developed compute-storage-network stack;
  • 500,000-card single cluster: combined with Alibaba Cloud's next-generation intelligent computing center network architecture, a single cluster can scale up to 500,000 cards.

This mirrors Huawei's "11 key chips" approach: the unit of competition has shifted from the single chip to whole-system delivery capability — Zhenwu handles compute, Yitian handles general-purpose computing, Panmai handles networking, Zhenyue handles storage, and ICN Switch handles interconnect.

3. Business and Roadmap​

  • The Zhenwu family has served over 650+ enterprise customers (as of June 2026), spanning autonomous driving, finance, large models, embodied AI, energy, and manufacturing;
  • Supernodes based on the M890 are already deployed at scale, running models with over 2 trillion parameters such as Qwen3.8 and Kimi K3; Alibaba Cloud will add new serving nodes in Q4 to expand supernode supply;
  • Roadmap: V900 enters mass production and sales in 2027 Q1; Zhenwu J900 is planned for release in 2027 Q3;
  • CPU synergy: Yitian 720 / 730 arrive in 2027 (the 730 is the first to adopt T-Head's fully self-developed CPU microarchitecture, with single-core SPECint2017/GHz up to 1.4x that of Yitian 710); the 2029 Yitian 750 will interconnect directly with Zhenwu AI chips via the ICN bus;
  • Alibaba Group CEO Eddie Wu said T-Head's AI chip annual shipment volume will increase substantially.

4. Competitive Coordinates: Two Swords in One Week​

Viewing the two mid-September launches side by side, the landscape of "system-level competition" among domestic AI chips is now clear:

DimensionHuawei Ascend 960 (9-17)Alibaba Zhenwu V900 (9-22)
Launch eventHC2026Apsara Conference 2026
Per-card memory288GB (self-developed HBM)216GB
Memory bandwidth9.6 TB/sUndisclosed
InterconnectLinJu UnifiedBus / NPO optical interconnectICN Switch memory-semantics interconnect
SupernodeAscend 960 supernode (4,096 cards, 8 EFLOPS FP8)Panjiu supernode (500,000-card single cluster)
Availability2027 Q22027 Q1

Huawei is taking the "supernode + open-source CANN ecosystem" route, while Alibaba is taking the "full cloud stack + open-source model ecosystem" route; both still trail in per-card specs, but what they deliver are procurable, operable ultra-large-scale clusters. The substitution logic against NVIDIA's CUDA ecosystem is shifting from "performance parity" to "supply certainty + system efficiency."

5. Summary​

  • V900: 3x M890, 216GB, 1200GB/s chip-to-chip interconnect, native FP8/FP4, mass production in 2027 Q1;
  • System: ICN Switch with thousand-card full bandwidth, Panjiu supernode with a fully self-developed compute-storage-network stack, 500,000-card single cluster;
  • Cadence: J900 in 2027 Q3, Yitian 730 in 2027, Yitian 750 in 2029 — T-Head's first three-year generational roadmap;
  • What to watch: actual mass production and Alibaba Cloud deployment scale in 2027 Q1, disclosure of V900's absolute compute, and the possibility of external supply of ICN Switch.

Further Reading​

References​

  • Huanqiu.com: T-Head launches new unified training-and-inference AI chip Zhenwu V900 (2026-09-22)
  • Sina Tech / TechWeb: Alibaba T-Head launches new AI chip Zhenwu V900, mass production in the first quarter of 2027
  • C114: Alibaba launches new AI chip Zhenwu V900, mass production in 2027 Q1

This article is compiled from Alibaba's official announcements and public reports. The absolute single-chip compute of V900 has not been officially disclosed; supernode and cluster scale figures are official claims, subject to actual future delivery.

昇腾 960 官宣发布:FP4 算力翻倍、全球首个 NPO 超节点,华为确认一年一代

· 6 min read
Industry Research Team

2026 年 9 月 17 日,华为全联接大会(HC2026),华为常务董事、ICT 基础设施业务总裁汪涛正式发布昇腾 960。与 8 月数字中国峰会上的"路线图预告"不同,这次是完整的规格官宣——而且提前三个季度就绪。本文基于华为官网新闻稿及多方报道整理,逐项拆解 960 的规格、超节点与路线图含义。


1. 昇腾 960:双版本,训练先行​

昇腾 960 分为两个版本,节奏错开一个季度:

指标昇腾 960DT(训练版)昇腾 960PR(推理版)
FP8 算力2 PFLOPS(较 950 翻倍)待披露
FP4 算力4 PFLOPS待披露
显存288GB(自研 HBM)待披露
显存带宽9.6 TB/s待披露
配套超节点Atlas 860(风冷)Atlas 960(液冷)
就绪 / 上市2027 Q1 / 2027 Q22027 Q3

三个读数:

  • 提前三个季度就绪:960DT 原计划 2027 年 Q4,现在 2027 Q1 就绪、Q2 上市——与 950DT"提前上线华为云"一脉相承,路线图可信度在持续兑现;
  • 精度口径与英伟达对齐:FP8 / FP4 主口径(辅以自研 HiF8 / HiFP8),FP4 已是 2026 新卡的通用口径;
  • 288GB 自研 HBM + 9.6TB/s:显存容量与带宽同步翻倍,对训练长上下文大模型是实打实的容量红利。

单卡规格详见昇腾 960 规格页;上一代昇腾 950DT 规格页。

2. 昇腾 960 超节点:全球首个 NPO 超节点​

这是本场发布真正的"重器"——昇腾 960 超节点是全球第一个采用 NPO(近封装光学)的超节点:

指标昇腾 960 超节点
单节点规模4096 卡
系统算力8 EFLOPS FP8 / 16 EFLOPS FP4
HBM 总容量1 PB
互联 RTT低至 2μs
可用度99.8%

NPO 意味着什么:传统光模块挂在交换机面板上,信号要先出封装再转电—光;NPO 把光引擎挪到交换芯片封装近旁,大幅缩短 SerDes 距离、降低功耗与时延。华为的具体数字:

  • 5500 个自研 Hi-ONE 光引擎(业界首个量产 NPO,单引擎 7.2T);
  • 替代 4.8 万颗 800G 光模块;
  • 降低功耗 550kW 以上,同时支撑 2μs 级互联时延。

对比 8 月数字中国峰会的口径(Atlas 960 单超节点 15488 卡),官方最新口径为单超节点 4096 卡(1 个超集群可由多超节点组成)——以华为官网最新新闻稿为准。

⚠️ 华为还给出了"MFU 提升 2.75 倍""推理时延降低 70%"等收益数据,均出自华为马尔科夫实验室仿真,尚无独立第三方实测,阅读时注意口径。

3. 一年一代:970(2028)→ 980(2029)​

华为首次把"一年一代"从愿景变成官宣承诺:

年份产品状态
2026昇腾 950 系列950PR 已量产,950DT Q4 放量
2027昇腾 960本次官宣,提前就绪
2028昇腾 970HC2026 确认
2029昇腾 980HC2026 确认

支撑这一节奏的是华为提出的"韬定律":算力规格每代翻倍,访存带宽、访存容量、互联带宽同步大幅提升。

生态侧的同场数据也值得记录:CANN 外部开发者占比首次超过 61%、月活开发者 5200 人;910C 超节点部署已超 1000 套。

4. 竞争坐标:单卡有代差,系统级对打​

把 960DT 放到 2027 年的棋盘上看:

  • 对 NVIDIA Rubin(R200:288GB HBM4 / 50 PFLOPS FP4):单卡 FP4 算力约为 R200 的 8%(4 vs 50 PFLOPS),单卡代差客观存在;但 4096 卡超节点 + NPO 互联是系统级竞争——用"可交付的超大集群"对打单卡性能;
  • 对采购方:960 的价值主张是"在国产量产约束内拿到最大可用集群"。DeepSeek 拟采购 16 万颗 950DT 跑推理的订单已经证明:当供给与生态到位,头部实验室愿意主动选国产;
  • 对国产链条:950→960 的显存翻倍直接拉动国产 HBM 迭代,这是整条链的胜负手。

5. 小结​

  • 960DT:FP8 2P / FP4 4P、288GB HBM、9.6TB/s,2027 Q2 上市,提前三个季度;
  • 960 超节点:全球首个 NPO 超节点,4096 卡 / 8 EFLOPS FP8 / 1PB HBM / RTT 2μs;
  • 路线图:一年一代官宣至 2029,950、960 均提前兑现,规划可信度显著上升;
  • 看什么:2027 Q2 实际交付节奏、国产 HBM 产能、CANN 生态的第三方模型覆盖度。

相关阅读​

参考资料​

  • 华为官网:发布全球首个采用 NPO 的超节点(昇腾 960 超节点)(2026-09-17)
  • 电子工程专辑:国产 AI 芯片新突破!华为昇腾 960 芯片将提前发布
  • 环球网:华为全联接大会 2026 主题演讲报道

本文基于华为官网新闻稿与公开报道整理。MFU、时延等收益数据为华为实验室仿真口径;2028 / 2029 产品仅为路线图官宣,规格以未来发布为准。

AWS Adds Another 2 Million NVIDIA GPUs: Vera CPU Debuts on AWS, 100,000 GPUs Reserved for the U.S. Government AI Factory

· 5 min read
Industry Research Team

This article is based on official announcements from AWS and NVIDIA (September 5, 2026) and public statements by executives of both companies.

On September 5, 2026, AWS and NVIDIA announced an expanded strategic partnership: on top of the "1 million additional GPUs starting in 2026" plan announced at GTC 2026, they will deploy 2 million more NVIDIA GPUs, covering three architecture generations — Blackwell Ultra, Rubin, and Rubin Ultra — with a deployment window of 2027-2028. Demand growth exceeding all previous forecasts was the direct reason both companies cited.

This is no longer a simple "chip purchase" — it is a full-stack partnership spanning GPUs, CPUs, interconnect, memory, open-source models, and software. We break the key information into six points.

1. Composition and Timeline of the 2 Million GPUs​

  • Scale: 2 million GPUs (added on top of the original 1 million GPU plan, tripling the total committed scale)
  • Architectures: Blackwell Ultra, Rubin, Rubin Ultra
  • Timeline: 2027-2028, deployed across AWS global infrastructure (including newly built AI factories)
  • Context: Amazon's 2026 capital expenditure guidance has been raised from $200 billion to $220 billion, and CEO Andy Jassy has explicitly said it is still not enough to meet AI compute demand

For reference, the figures Jensen Huang gave at GTC 2026 put cumulative orders and demand for the Blackwell and Rubin platforms through 2027 on track to reach $1 trillion (at GTC 2025, the estimate for 2026 was roughly $500 billion). AWS's add-on order is one of the heaviest puzzle pieces in that big picture.

2. Vera CPU Comes to AWS for the First Time​

This is the most structurally significant change in the partnership: NVIDIA Vera CPU infrastructure will enter AWS.

The Vera CPU is the general-purpose processor in the Vera Rubin platform designed for agentic AI, positioned to efficiently turn AI resources into "completed agent tasks." As agentic workloads rise, CPU-side pressure on task orchestration, memory management, and data scheduling rises in step — pure GPU expansion is no longer enough, and AWS needs a matching high-performance CPU layer. This also aligns with AWS's strategy of "offering the broadest compute choices, from in-house chips (Graviton/Trainium) to partner chips": Vera does not replace Trainium, but fills in the CPU compute layer alongside the accelerated infrastructure.

At re:Invent 2025, AWS announced that its next-generation Trainium chip would support NVIDIA NVLink Fusion high-speed interconnect. This time, both companies took the partnership one step further:

  • Amazon Annapurna Labs will support NVIDIA's new custom high-bandwidth memory NVHBM (developed in collaboration with memory vendors)
  • Trainium thereby gains a faster, more power-efficient memory option
  • Trainium and GPUs can work together within the same rack-scale architecture, sharing scale-up interconnect

For the chip industry, this is a signal worth watching closely: NVLink Fusion + NVHBM means NVIDIA's interconnect and memory technologies have started "supplying" competing ASICs. The boundary between the in-house ASIC camp (Trainium, TPU, MTIA) and the NVIDIA GPU camp is shifting from "either/or" to "hybrid deployment."

4. 100,000 GPUs: The U.S. Government Sovereign AI Factory​

A dedicated public-sector business is carved out of the partnership: AWS and NVIDIA will build AI factories for the U.S. government, deploying 100,000 GPUs on secure AWS infrastructure to host federal and national security workloads (Impact Level 6 and above), supporting the development of advanced AI models within strict regulatory frameworks.

Sovereign AI turning from a slogan into concrete numbers is a defining feature of this cycle — government customers are becoming first-class buyers of AI compute.

5. Software and Physical AI Deepen in Parallel​

Beyond hardware, the software layer of the partnership is also strengthening:

  • NVIDIA Nemotron open-source models continue to arrive on Amazon Bedrock and SageMaker
  • Amazon EMR data processing and OpenSearch vector indexing are accelerated by cuDF / cuVS
  • Amazon Robotics officially adopts the NVIDIA physical AI platform (Jetson, Omniverse, Isaac) for warehouse automation and next-generation robotics

Physical AI (robotics, embodied intelligence) is becoming a new growth pole in cloud providers' compute narratives — the same trend as JD.com, Tesla, and others writing embodied intelligence into their compute procurement logic.

6. Implications for Compute Buyers​

ObservationImplication
Demand "beat all forecasts"Supply tightness is not a short-term phenomenon; the 2027-2028 compute window must be locked in now
Mixed procurement across three architecturesDuring the Rubin/Rubin Ultra production ramp-up, Blackwell Ultra remains the delivery mainstay; procurement needs cross-generation planning
CPU layer revaluedagentic AI pushes the bottleneck from GPU to CPU orchestration and memory bandwidth; do not focus only on accelerator cards when selecting
Sovereign AI landsGovernment-scale orders enter the market, further tightening the allocatable supply of high-end GPUs

For decision-makers weighing build versus rent, every massive add-on order from cloud giants reprices the future rental curve. If you are evaluating GPU purchase or rental options, we recommend running a quantitative calculation with the TCO Calculator: enter chip price, power draw, utilization, and rental rates to compare 3-year total cost of ownership.

Summary​

2 million GPUs, Vera CPU in the cloud, NVHBM opening up, 100,000 sovereign compute GPUs — this round of expansion between AWS and NVIDIA pushes the "AI factory" race into the full-stack era. Compute scarcity will most likely only tighten before 2027; whether you are buying, renting, or betting on domestic alternatives, locking in supply and cost curves early is the surest move right now.

(The data in this article comes from official AWS/NVIDIA announcements and public statements by executives of both companies; architecture performance figures are as released by the vendors.)

昇腾 950 超节点 Q4 上市,中兴、新华三、曙光全线跟进:国产超节点从概念走向交付

· 6 min read
Industry Research Team

芯片追不上,系统来补。这是国产算力过去一年最清晰的战略共识。2026 年下半年,国产超节点正式从"概念发布"进入"产品密集发布、联合适配和商业部署"阶段:华为 Atlas 950 SuperPoD 真机已亮相并计划 Q4 上市,中兴、新华三、中科曙光、壁仞、沐曦等一众厂商相继入局。


1. 华为 Atlas 950:业界最大 1024 卡超节点,Q4 上市​

继 7 月 WAIC 2026 真机首秀后,华为昇腾 950 超节点的商用脚步持续加快。核心规格:

指标Atlas 950 SuperPoD
互联规模最大 1024 × 昇腾 950DT(灵衢高速互联)
AI 算力1 EFLOPS FP8/mxFP8/HiF8、2 EFLOPS FP4
全局内存256TB 统一编址,片上内存最大 1024 × 96GB @ 4.0TB/s
互联带宽单柜最大 64 × 1.68 TB/s(双向),3μs 超低 RTT
形态满配 128 计算柜 + 32 互联柜,占地约 1000㎡,8192 颗 950DT
供电散热全液冷,100kW 供电,380V AC/336V DC/240V DC
上市时间2026 年第四季度

在 WAIC 展台之外,两个数字更能说明商业化进度:昇腾 384 超节点已商用落地 750 多套,规模应用于互联网、运营商、金融、教育、医疗等行业,并且是国内唯一训练出 SOTA 模型的超节点;更远期规划中,Atlas 950 SuperCluster(超 50 万颗昇腾芯片)同样定于 2026 Q4,Atlas 960 SuperPoD(15488 卡)与 Atlas 960 SuperCluster(超 100 万卡)将于 2027 Q4 接力。

华为的路线图背后是"Tau Scaling Law"——在 EUV 光刻机受限的前提下,通过系统级扩展突破单芯片物理极限,并计划到 2031 年将晶体管密度推升至 1.4nm 级工艺水平。

2. 超节点赛道全面开花:三条技术路线​

WAIC 之后,国产超节点形成了三类清晰的技术路线:

第一类:华为全栈自研。 同时掌握昇腾芯片、灵衢互联、Atlas 硬件、CANN 软件栈和 MindSpore 框架。CANN 已全面开源:社区上线 67 个项目、开源代码超 1244 万行、月活开发者突破 3500 人;昇腾开发者总数超 400 万。

第二类:中兴、新华三的兼容路线。 中兴发布 OEX 超节点新品,依托协议标准化与接口统一,全面兼容多元 GPU 生态;新华三 UniPoD S80000 系列可从 32 卡扩展至 1024 卡、最大支持 16384 卡互联,将 Scale-up/Scale-out 网络、液冷、供电、管理和故障恢复整合进同一套架构。

第三类:国产 GPU 联合创新。 中兴联合曦智、壁仞、沐曦、燧原、天数智芯等,基于 OEX+dOCS 架构打造国产高性能 Matrix 超节点;壁仞计划发布 BR20x 系列 GPU,基于自研 BLink 2.0 互联协议实现单超节点 1024 卡 Scale-up;沐曦推出"曦景"S 系列超节点。

在更大的 Scale-out 层面,中科曙光"曙光 8000(登峰)"全国产十万卡 AI 超集群已落成并接入国家超算互联网——它不是单体超节点,而是由大量计算节点、超节点和网络存储组成的超集群,检验的是国产算力的大规模调度、应用适配和工程交付能力。

3. 为什么"超节点"成了国产算力的主战场​

根本原因在于衡量标准变了:国产算力的评价体系,正从"比较单颗芯片的峰值性能"转为"一整套系统能够调动多少有效算力"。华为副董事长徐直军对此直言:单芯片上英伟达仍领先且短期难以追赶,但在超节点和集群层面,华为有信心。

这一转向的产业逻辑:

  • 出口管制下的现实选择:单卡制程受限 → 用高速互联把更多 NPU 组织成一台"大计算机";
  • 需求侧真实拉动:万亿参数 MoE 模型的训练与推理,天然需要大规模低时延互联,超节点恰是对症下药;
  • 生态护城河前移:CANN 开源 + MindSpore 兼容 PyTorch,把 CUDA 生态的竞争从"算子覆盖"层面拉回到"开发者规模"层面。

东兴证券指出,全球超节点赛道已聚集微软、Meta、亚马逊、三大运营商、阿里、字节、腾讯、百度、曙光、中兴、浪潮、新华三、海光、沐曦等数十家厂商——格局未定,但英伟达一家独大的局面正在被谷歌 TPU、AMD Helios、华为 Atlas 三股力量同时挑战。

4. 三道关卡:从"造出来"到"卖得动"​

业内人士普遍认为,国产超节点距离真正成熟还有三道坎:

  1. 单芯片性能与单柜算力密度:受制程限制,单卡性能差距仍存,现阶段靠多柜互联和规模扩展另辟蹊径;
  2. 规模扩展效率:超节点扩展到一定规模后性能提升边际递减,通信复杂度、能耗和容错成本持续上升;
  3. 可验证的商业价值:下一阶段的比拼是有效算力利用率、单位 token 成本、模型适配广度与客户复购意愿。

一句话总结:2026 年是"中国超节点元年",2027 年见真章——届时 Atlas 950 与 Vera Rubin NVL 系列将在全球两个平行市场各自接受规模化交付的检验。


相关链接​

参考资料​


本文基于 WAIC 2026 现场报道、昇腾社区官方规格与公开研报整理。超节点性能数据均为厂商公布口径,独立第三方评测结果尚待观察。

Vera Rubin 全面量产:100% 全液冷 + 800V 直流供电,AI 数据中心基础设施范式重构

· 6 min read
Industry Research Team

2026 年 9 月初,供应链信息确认:英伟达 Vera Rubin 平台已于 8 月正式量产、9 月启动批量出货,无延期、无卡顿。与 Blackwell 迭代初期的产能波折不同,这次量产节奏异常平稳——谷歌云、微软 Azure、CoreWeave、甲骨文云等头部云厂商已启动机架部署。但真正值得产业记住的,不是"又一代 GPU 量产了",而是 Rubin 把液冷从"可选配置"变成了"硬性前置条件"。


1. 量产节奏:史上最平稳的一次平台切换​

根据产业链调研与券商跟踪信息:

  • 2026 年 8 月:Vera Rubin 正式量产;
  • 2026 年 9 月:批量出货启动;
  • 2026 下半年:CoreWeave、谷歌云、微软 Azure、甲骨文云机架部署落地;
  • 2026 年:上代 GB 架构机柜出货量有望达 6 万台(同比翻倍);
  • 2027 年:GB 与 Rubin 两代平台合计出货体量有望接近 10 万台,Rubin 新机柜远期产能目标为每天 1000 个 NVL72 机柜。

需求侧同样在加码:华尔街报告披露,英伟达管理层表示 FY28 同比增长 70% 的目标并非需求上限——若供应不受限,增速可能超过 100%。当前主要约束已从需求端转向先进晶圆与 HBM 供应。

2. 单卡 2300W:风冷时代的终结​

Rubin 平台与前代最根本的差异不在算力,而在功耗密度:

指标H100GB300Rubin
单 GPU TDP700W~1400W2300W
机柜功耗~40kW~140kW190–230kW
散热方案风冷为主风液混合100% 全液冷
供电架构48V48V800V 高压直流

单芯片 TDP 从 700W 升至 2300W、单机柜功率密度突破风冷物理极限——这意味着 液冷不再是高端算力的选配升级,而是运行 Rubin 服务器的先决条件。英伟达官方将 Rubin 全液冷架构定义为"数据中心历史上最重要的能效突破之一",并已写入 DSX AI 工厂参考设计:所有跟随英伟达技术路线的云厂商和数据中心运营商,都必须采用全面液冷方案。

三个关键架构变化:

  1. 无风扇整机:GPU、CPU、交换机、DPU 全部器件强制采用直接冷板式液冷,45℃ 温水冷板成为出厂标配;
  2. 液冷边界延伸:散热覆盖范围从 GPU 冷板延伸至 CPU、DPU、交换机乃至光模块(液冷 Cage/鼠笼开始从"可选"变"刚需"),整套液冷系统价值量较 GB300 提升约 40%;
  3. 800VDC 供电:替代传统 48V 机架配电,整机电源 BOM 价值增长 30% 以上,PSU 电源模块从 5.5kW 向 18.3kW 迭代,固态变压器、高压直流 CDU 成为数据中心新增核心设备。

3. 对产业链的三重传导​

第一重:液冷从"配套"变"主角"。 2026 下半年以小规模部署验证为主,真正的放量窗口在 2027 年——Rubin 机架大规模铺货后,冷板、快速接头、CDU、液冷泵进入业绩兑现期。台系供应链 7 月数据已率先验证:AVC 奇鋐 7 月营收 185.9 亿新台币创历史新高(同比 +57.4%),双鸿、健策 7 月同比分别 +116.7%、+91.0%。

第二重:国产液冷供应链进入核心 BOM。 国内厂商由外围冷源和代工环节逐步进入芯片平台、服务器 ODM 和海外云厂商供应体系,替代路径从 Manifold、管路推进至高可靠快接头和冷板。英维克 26H1 海外收入占比 71.4%,飞龙股份液冷泵小功率平台订单超 5 万台——液冷全核心零部件自主可控正在成为现实。

第三重:供电与散热边界融合。 800VDC 架构下,电源模块、PDB 配电单元、高速交换芯片自身发热也达到很高水平,部分电源组件同样需要液冷辅助散热——电源与温控两条产业链正在合并成一条。

4. 需求矩阵扩容:云厂商之外,太空算力入场​

Rubin 的客户矩阵已从传统云厂商扩展至三个层次:

  • 全球云厂商:谷歌云、Azure、甲骨文云、CoreWeave;
  • AI 科技巨头:马斯克公开披露 2027 年 8GW 超大规模 IDC 建设规划;SpaceX 将 Vera Rubin 架构定义为"最优 AI 计算架构",计划地面与太空双向部署,支撑 "Starmind" 卫星算力项目;
  • 主权与边缘:远期 Rubin Ultra 及 2027 年后更高功耗机型单机柜有望冲击 600kW+。

普华永道预计全球数据中心累计投资到 2035 年将达 31.6 万亿美元。AI 基础设施建设的确定性,已经从"是否建设"变成"多快建设"。

5. 对采购方的启示​

  • 机房规划前置:2027 年起采购 Rubin 级算力,液冷改造(单千瓦改造成本较高)或按全液冷标准新建,必须在预算周期一开始就纳入;
  • 看 PUE 也看水温:45℃ 温水直冷允许更高进水温度,可利用自然冷源压低 PUE——选址时人工冷源依赖度成为新的评估维度;
  • 供应商组合即风险对冲:HBM 与先进封装供应是当前核心瓶颈(详见本站 HBM4 竞速分析),供应链多元化比单点性能更重要。

相关链接​

参考资料​


本文基于 2026 年 9 月初供应链调研、券商研报与英伟达官方披露整理。出货量与功耗数据为产业链预估口径,实际以英伟达及客户正式披露为准。

Hyperscaler Custom Silicon Wave 2026: OpenAI Jalapeno, Maia 200, MTIA, TPU v8 Together "De-NVIDIA-ize"

· 6 min read
Industry Research Team

The Tuesday-afternoon AI session at Hot Chips 2026 this August was the most historically significant of the conference — not because any single chip was so powerful, but because almost everything on stage was a "hyperscaler de-NVIDIA-ization" custom ASIC: Google's 8th-gen TPU, OpenAI's first self-designed chip, Microsoft Maia, Meta MTIA, and Cerebras wafer-scale racks, all on one stage. When the world's largest AI compute buyers start treating GPUs as "one of the options," the power structure of AI hardware is loosening.


1. OpenAI Jalapeno: Building a Chip in 9 Months​

On June 24, 2026, OpenAI, together with Broadcom, unveiled its first self-designed inference ASIC, Jalapeno — the fifth member of the "custom inference chip club."

DimensionJalapeno
PartnerBroadcom + TSMC manufacturing
PositioningInference-specific ASIC
Design cycle9 months end-to-end (Greg Brockman says aided by OpenAI's own models)
Cost target~50% lower token cost vs general-purpose GPU stack
Commercial timingFirst deployments by end-2026; long-term goal 10GW of self-designed chips
Deal scaleUp to $10B strategic partnership with Broadcom (accelerators + networking by 2029)

The talk title "You Can Just Build Things … Chips" is itself a signal: the largest AI compute buyer no longer defaults to GPU as the only path.


2. Google TPU v8: The Biggest Architectural Pivot in a Decade — Train/Infer Split​

Google has the longest custom-chip history (2016 to now), and its 8th-gen TPU for the first time splits the product line in two:

ModelCodenamePartnerPositioningKey Specs
TPU 8tSunfishBroadcomTraining9,600 cards per pod, 121 FP4 ExaFLOPS, 2PB shared HBM, 2× ICI bandwidth
TPU 8iZebrafishMediaTekInference288GB HBM, 384MB on-chip SRAM (3× prior gen), 19.2 Tb/s ICI

On capacity, Morgan Stanley estimates based on supply-chain interviews that Google TPU production in 2026 may exceed 3 million units (a brokerage estimate, not an official target). Google is also the only vendor to achieve large-scale custom-chip deployment and sell compute externally (Gemini runs on TPUs).


3. Meta MTIA: From Recommendation Systems to a GenAI Dual Mission​

Meta's custom journey has the clearest starting point — MTIA was originally built for recommendation ranking hardware and is being pulled toward a dual mission by generative AI.

  • MTIA 300 is deployed; 400 / 450 / 500 are planned at roughly one new model every 6 months through 2027;
  • Based on RISC-V, Meta claims up to 25× compute gain;
  • Node evolves with industry cadence: 100 (7nm) → 200 (5nm) → 300 series (3nm + CoWoS);
  • In partnership with Broadcom; another chip codenamed Iris reportedly passed testing in July 2026;
  • Meta plans to start volume production of one of them in September 2026, doubling its overall compute.

4. Microsoft Maia 200/300: Most Advanced Deployment​

Microsoft's Maia 200, released January 26, 2026, is the most advanced in deployment among the four:

DimensionMaia 200
ProcessTSMC 3nm, 140B+ transistors
Compute10+ PFLOPS FP4 / 5 PFLOPS FP8
Memory216GB HBM3E, 7 TB/s
Power750W
DeploymentAlready running in Des Moines data center, serving OpenAI GPT-5.2 and Microsoft 365 Copilot

Microsoft claims roughly 3× the performance of Amazon's Trainium on specific benchmarks. The short-term strategy is a dual track of "self-designed Maia + purchased NVIDIA" in parallel — self-designed chips need time from design to mass production, and NVIDIA's mature ecosystem cannot be replaced in the short term.


5. Amazon Trainium 3 and Anthropic's In-House Team​

  • Amazon: The Trainium series is already commercial, with 1.4 million units cumulatively deployed (officially disclosed) — a multi-billion-dollar business; its strength is the AWS customer base, letting enterprises choose between NVIDIA GPUs and self-designed chips. Trainium 3 continues this path.
  • Anthropic: In August 2026 announced the formation of an in-house chip team, with no tape-out or mass-production timeline yet; initially positioned as a complement (not a replacement) to existing partnerships with NVIDIA/AMD/AWS/Google Cloud, aiming to tailor-build for the Claude architecture and shed reliance on a single GPU.

6. NVIDIA's Answer: Not a Faster GPU, But Full-Stack​

It's easy to simplify the narrative to "four companies build chips, NVIDIA defends GPU." But NVIDIA took 6 slots at Hot Chips: a RISC-V tutorial, the Vera CPU, the Rubin GPU, the BlueField-4 DPU, the Spectrum-X multi-plane network, and an LPU accelerator.

A hyperscaler ASIC replaces only one of those five pillars. If the CPU, NIC, switching fabric, and software all come from the same vendor, what you save by swapping out the accelerator is far less than the accelerator line item on the bill suggests. Rubin's play is a full-stack AI factory platform spanning seven chips and five racks — the competitive answer is "full-stack positioning," not "a faster single chip."


7. Trend Judgment: Inference De-GPU-izes, Training Still GPU-Led​

  • Inference side: The CUDA moat visibly shallows. Inference is parallelizable and replaceable at the endpoint; custom ASICs trade away the generality tax (implementing only the operations LLMs actually execute) for lower cost/token. Groq LPU, Cerebras, and various TPU/ASIC players all compete on the same metric.
  • Training side: Foundation models are still trained on GPUs, with no serious challenger in the short term. NVIDIA's three training moats (fastest silicon + NVLink + CUDA) remain firm.
  • Conclusion: Custom chips are not "replacing NVIDIA," but giving buyers a credible external negotiation option in the largest and fastest-growing battlefield — inference. That alone is enough to reshape the economics of AI infrastructure.

References​


This article is compiled from August 2026 Hot Chips on-site reports, corporate announcements, and industry analysis. Some capacity and performance figures are brokerage estimates or vendor-disclosed figures; actual results are subject to mass-produced products.