Skip to main content

33 posts tagged with "Industry News"

Major events and market dynamics in the AI compute card industry

View all tags

NVIDIA's China Share Crashes to 8% as Ascend Rises to 50%: The Dramatic Reshuffle of China's AI Compute Market in a September Report

· 5 min read
Industry Research Team

This article is based on market analysis published by Bloomberg Intelligence and Bernstein in September 2026 and public reporting; market share and revenue figures are analyst estimates, not official disclosures.

On September 28, two Wall Street reports simultaneously painted a dramatic picture: the balance of market share in China's AI compute market completed an almost total flip within 12 months.

1. The Share Flip: 40% to 8%, Ascend to 50%​

Bernstein's latest forecast:

  • NVIDIA's share of the China AI chip market will fall from about 40% to about 8% by the end of this year
  • Huawei Ascend rises to nearly 50%
  • The immediate trigger: China's September ban on new H20 orders from NVIDIA, which shut the last channel for NVIDIA's return to the Chinese market
  • ByteDance, Alibaba, and Tencent have all placed large Ascend orders

On the timeline, this was not a sudden shock but the cumulative result of three years of tightening export controls: NVIDIA's share of the China AI chip market once peaked at 95%; as high-end chips were cut off, Chinese hyperscale cloud providers pivoted wholesale to Huawei and domestic alternatives. Even when Washington allowed H200 sales to China in January 2026, the gears of that shift had already meshed — there was no going back.

2. Ascend 950PR: The "Chinese H200" in Analysts' Eyes​

What underpins this share data is product strength itself:

  • The Ascend 950PR entered mass production in March this year, with FP4 compute of up to 2 PFLOPS and 128GB of domestic HBM
  • Analysts rate it as roughly on par with the NVIDIA H200
  • Huawei's AI chip revenue is expected to reach $12 billion this year, up 60% from $7.5 billion in 2025

For the full specifications of the 950 series, see our earlier in-depth articles: Ascend 950PR/950DT Dual-Configuration Analysis and Ascend 960 Official Launch — the 960 was ready three quarters ahead of schedule, confirming a one-generation-per-year cadence.

3. The Model Side: Chinese Open Models Take the OpenRouter Token Pie​

While chip shares flipped, the model usage data is equally striking (OpenRouter platform statistics):

MetricData
DeepSeek token share (June)16.3%, surpassing Google, Anthropic, and OpenAI individually to rank first
Combined token share of Chinese open-weight models (May)About 61% (DeepSeek, Qwen, MiniMax, Tencent Hunyuan, etc.)
Weekly token consumption of Chinese modelsAbout 18 trillion, versus about 5.5 trillion for US models — a gap of more than 3x, with the overtake completed within a year
China-US frontier model performance gap (June)Narrowed to 6%, a historic low (9% in May)

Compute and models form a positive feedback loop: cards that cannot be bought force domestic compute to scale up; scaled-up domestic compute produces cheap tokens; and cheap tokens let Chinese open-source models swallow the bulk of global inference traffic.

4. But the Stock Market Isn't Buying It​

Bloomberg Intelligence pointed out a contradiction in the same period: the valuation discount of the China Tech 8 relative to the US Magnificent Seven has widened to more than 50%, the widest of the year, and BI believes only a "true AI breakthrough" can close it. Year-to-date stock performance: Alibaba about -7%, Tencent about -23%, while US AI infrastructure leaders broadly gained more than 15%.

There are two readings of this divergence: either the market is underestimating China's AI fundamentals, or earnings execution (Alibaba's EPS missed expectations by 17.3% last quarter, Baidu's by 36%) has dragged down the delivery of the AI narrative. For industry observers, the notable point is: the compute-side data (share, revenue, shipments) has moved ahead, while application-side monetization is still on the way.

5. Three Takeaways for Compute Buyers​

  1. For deployments in China, the Ascend ecosystem is already the default option. The cost of migrating from CUDA to CANN is one-time, while supply uncertainty is persistent; after the September ban, there is no compliant path to procuring new NVIDIA cards in the Chinese market.
  2. Domestic compute "matching the H200" is a watershed. The implicit assumption of domestic substitution used to be "discounted performance, discounted price"; when single-card performance catches up to the H200 generation, procurement decisions return to pure TCO and supply-security calculations — we recommend using the TCO Calculator to put electricity, depreciation, and utilization into the same table.
  3. The token cost curve determines the application landscape. Chinese models' 61% share of OpenRouter tokens shows that the chain of compute self-sufficiency leading to token price drops leading to application prosperity has completed its first leg; the next question is whether inference costs can keep falling.

Summary​

From 40% to 8% is not a slogan but a market reshuffle driven jointly by the ban, product strength, and the model ecosystem. For the Chinese market, the mass-production cadence of Ascend 950/960 and the maturity of the CANN ecosystem will determine whether this share can be held; for the global market, the combination of China's compute "internal circulation + open-source models going overseas" is rewriting the geographic distribution of inference traffic.

(Market share and revenue forecasts come from Bernstein / Bloomberg Intelligence analyst reports; actual figures are subject to each company's financial reports.)

Intelligent Compute to Hit 9,800 EFLOPS by 2030: Five Hard Targets in the ICT Industry "15th Five-Year" Plan

· 6 min read
Industry Research Team

This article is based on the MIIT "15th Five-Year" Plan for ICT Industry Development (issued in September 2026) and official interpretations by the People's Post and Telegraph News, The Paper, Huaxia Times, and other outlets.

In early September, the MIIT issued the "15th Five-Year" Plan for ICT Industry Development, drawing a roadmap for the ICT industry over the next five years with 13 major indicators and 26 key tasks. For the AI compute industry, this is the most substantive policy document — we have picked out five hard targets and break down their industry implications one by one.

Target 1: Intelligent Compute from 1,590 to 9,800 EFLOPS, More Than 5x Growth in Five Years​

This is the standout number in the entire plan. For reference:

  • As of the end of June 2026, national intelligent compute capacity had reached 2,185 EFLOPS, up 177% year over year — the first half alone overshot the year-end 2025 target (1,590 EFLOPS)
  • The country has built 42 10,000-card-class intelligent computing clusters; in July 2026, the first fully domestic 100,000-card-class AI compute cluster, the Sugon 8000, was completed in Zhengzhou
  • Huawei's rotating chairman Wang Tao's assessment: super-node clusters of more than 100,000 cards will become basic configuration by 2027

Going from 2,185 to 9,800 means that over the next four years the country must build 3.5 times the current installed base of intelligent computing facilities. The plan also calls for "orderly deployment of 10,000-card, 100,000-card and larger intelligent computing clusters" and "stepping up efforts to adapt domestic compute chips" — this is the most certain demand base for domestic AI chips over the next five years.

Target 2: Cumulative Information Infrastructure Investment of RMB 3.8 Trillion​

Some have compared this with the "14th Five-Year" figure of RMB 3.7 trillion and concluded that "investment has peaked." Official interpretations explicitly reject this claim: incremental capital is shifting from "scale expansion" to "quality-and-efficiency gains," precisely targeted at intelligent computing infrastructure, 10G optical networks, 6G, and other new tracks. The industry chain pull effect is changing accordingly — AI servers, high-speed optical interconnects, domestic compute chips, and new smart terminals were named as beneficiaries across the whole chain.

Target 3: PUE of New Large Compute Facilities Reduced Below 1.2​

At the end of the "14th Five-Year" period this figure was 1.25, and it must fall another 0.05 within five years — it doesn't sound like much, but with AI servers running at high load year-round and per-rack power density generally exceeding 20kW, every point of PUE reduction is a hard fight. The technical path laid out in the plan is very clear: liquid cooling.

  • The "15th Five-Year" Plan for ICT Industry Development: guide compute facilities to adopt high-efficiency energy-saving equipment and advanced technologies such as liquid cooling
  • The "15th Five-Year" Plan for Electronic Information Manufacturing Development (jointly issued by the MIIT and the NDRC on September 15): lists liquid cooling alongside high-bandwidth memory pooling and all-optical switching as key technologies to be broken through

Industry-side data confirms the trend: according to Omdia, liquid cooling's share of the global data center cooling market climbed from about 20% in 2024 to about 37% in 2025, with penetration expected to exceed 45% by 2029. The ceiling on compute expansion is shifting from chip supply to power supply, and the energy-efficiency constraint of "compute up, energy down" will be the norm for the next five years.

Target 4: Advanced Storage Capacity of 1,700 EB, More Than 2x Growth​

While intelligent compute grows 5x, advanced storage grows more than 2x — storage-compute coordination is given equal weight. The logic: training checkpoints for large models, and the intermediate states and contextual memory of agent inference, are all stored in layers within high-performance storage. As of the end of June, national storage capacity totaled about 2,021 EB, of which advanced storage accounted for about 32%; the 1,700 EB advanced storage target for 2030 means structural upgrading matters more than total-volume growth.

A direct implication for chip selection: KV cache and long context are eating the memory budget — when evaluating AI servers, memory capacity and memory bandwidth should carry more weight than peak compute.

Target 5: The Agent Interconnection Network Enters the Plan for the First Time​

This is the most forward-looking part of the plan: the "agent interconnection network" gets its own dedicated column, deploying four areas of work — building an agent network identifier system, accelerating the construction and application of agent network infrastructure, promoting global interconnection, and establishing a space governance system. It also explicitly states "launching 6G commercial use in a timely manner."

As AI shifts from "applications for people" to "agent infrastructure," the role of the communications network upgrades from connecting people to connecting agents — providing a national-level narrative for the distributed deployment of inference compute (edge inference, compute scheduling).

Three Judgments for Compute Practitioners​

JudgmentBasis
Domestic compute demand is highly certainThe plan explicitly states "stepping up efforts to adapt domestic compute chips" + the 100,000-card cluster build cycle has begun
Energy-efficiency targets become hard siting constraintsPUE below 1.2 + liquid cooling named as a key technology; high-density liquid cooling solutions take priority
Inference compute sinks to the edge on demand"Deploy inference compute facilities as needed for scenarios"; edge and regional compute centers enjoy a policy window

Under a 5x compute expansion target, what has always been scarce is not planning but chips, power, and delivery capability. For buyers, the supply window remains tight; for solution selection, we recommend using the TCO Calculator to convert PUE differences into electricity costs — the gap between 1.25 and 1.2 amounts to tens of millions on the electricity bill of a 10,000-card cluster.

Summary​

The "15th Five-Year" Plan writes intelligent compute into a national-level project: 5x compute, 2x storage, RMB 3.8 trillion in investment, PUE 1.2, and the agent internet. Looking back five years from now, the wave of domestic 100,000-card clusters in the second half of 2026 may well prove to be the starting point of this curve.

(Plan data is cited from MIIT documents and official media interpretations; market data is cited from statistics by CAICT, Omdia, and other institutions.)

Groq 3 LPX Enters Full Mass Production: Samsung 4nm Foundry, 315 PFLOPS FP8 per Rack — NVIDIA Turns the LPU into the Seventh Chip of Vera Rubin

· 5 min read
Industry Research Team

This article is based on NVIDIA official product pages, the March 2026 GTC architecture blog, and third-party benchmark data; performance figures are vendor architecture comparisons, and actual gains should be verified with your own tests.

NVIDIA has confirmed that Groq 3 LPX — the "seventh chip" of the Vera Rubin platform — has entered full mass production, manufactured at Samsung's Pyeongtaek campus. This is the first time the deal has landed in the form of production silicon since December 2025, when NVIDIA spent roughly $20 billion to obtain a non-exclusive IP license to Groq plus its core engineering team.

For the inference hardware landscape, this is a paradigm-level event: for the first time, NVIDIA has incorporated a "specialized inference architecture defined by someone else" into its own flagship platform — and not as an add-on sale, but as deep co-design.

1. The Groq 3 LPU Chip and the LPX Rack: Specs at a Glance​

Single Groq 3 LPU (4nm, Samsung foundry):

ItemValue
Compiler-managed on-chip SRAM500 MB
SRAM bandwidth150 TB/s
Chip-to-chip scale-up bandwidth2.5 TB/s (96 chip-to-chip links, 112 Gbps each)

A single LPX rack (32 1U liquid-cooled trays, 8 LPUs per tray):

ItemValue
Total LPUs256
FP8 compute315 PFLOPS
Total on-chip SRAM128 GB
Aggregate SRAM bandwidth40 PB/s
In-rack scale-up bandwidth640 TB/s
DDR5 memory12 TB (hosting large model weights)

The 500 MB of SRAM per LPU may not look like much, but multiplied by 256 LPUs and stacked with 40 PB/s of aggregate bandwidth, it forms the physical foundation of "deterministic low-latency decoding" — the LPU's design philosophy is precisely to use compiler static scheduling of on-chip SRAM to completely eliminate memory-fetch stalls during the inference decoding phase, which is exactly the most typical bottleneck of GPU inference.

2. AFD: Attention and FFN Split Up, GPU and LPU Each Do Their Own Job​

LPX does not replace the GPU; instead, it forms a heterogeneous system with Vera Rubin NVL72, centered on AFD (Attention-FFN Disaggregation):

  1. Rubin GPUs handle prefill and attention — building the KV cache over large contexts and executing attention layers, consuming HBM capacity and high throughput;
  2. Groq 3 LPUs handle FFN / MoE expert-layer decoding — latency-sensitive and pattern-predictable, exactly the home turf of deterministic SRAM scheduling;
  3. Intermediate activations are exchanged between the two engines token by token, orchestrated and routed by NVIDIA Dynamo: the GPU computes every attention layer, the LPU computes every feed-forward layer, jointly producing each output token.

The logic of this division of labor is clear: agentic AI applications can consume 15x the tokens of traditional AI applications, and the bottleneck shifts from "can it finish computing" to "can it keep emitting tokens at stable low latency." Let the big HBM container run attention, and let the deterministic SRAM engine run decode — each plays to its strengths.

3. Measured Results and Official Figures​

  • Third-party benchmarks (Artificial Analysis): with Gemma 4 31B at a 100K token context, LPX output reaches roughly 3400 tokens/s — inter-token intervals below 1 millisecond, a qualitative leap in interactivity for long-context agentic scenarios
  • Official architecture comparisons: with LPX added, Vera Rubin NVL72 achieves up to 35x higher throughput per megawatt on trillion-parameter models; up to 10x more revenue opportunity per watt in "high-value token" scenarios
  • Launch customers: Nebius is the first to deploy LPX racks to expand its token capacity; CoreWeave already connects Vera Rubin racks in production with Spectrum-X Multiplane

It must be emphasized: the 35x/10x figures are architecture comparisons at specific high-interaction operating points, not universal conclusions. For training, high-throughput batch inference, and workloads that need CUDA ecosystem flexibility, the GPU remains the right answer; LPX's home turf is scenarios where "single-user interactive latency is the product" — agent loops, real-time coding assistants, long-context conversations.

4. Three Industry Signals​

1. Specialized inference chips get absorbed, not opposed. The old narrative of the LPU as a "GPU challenger" has become part of Vera Rubin. The endgame for inference hardware may not be one architecture winning, but the heterogeneous combination of "GPU + specialized decoding engines" becoming the standard. For other inference chip startups (Cerebras, Etched, etc.), this is both proof that the ceiling has risen and a warning: get integrated or find differentiation.

2. Samsung foundry lands a high-end AI order. Against the backdrop of TSMC's near-monopoly on AI main chips, Samsung 4nm taking on LPX mass production is highly significant — combined with Tesla's earlier AI6 2nm order, Samsung foundry has a shot at returning to full-year profitability in 2027.

3. Inference economics enters the "priced per MW" era. When vendors start telling their story with tokens/MW and revenue per watt, the core KPI of compute selection has completely shifted from peak compute (TFLOPS) to token output per unit of energy. This is consistent with the power-and-electricity cost model built into our TCO Calculator: the chip with the best-looking peak specs is not necessarily the chip with the lowest cost per token.

Summary​

Groq 3 LPX mass production marks the arrival of a heterogeneous era for inference hardware: "GPUs manage throughput, LPUs manage latency." For teams currently selecting inference clusters, the recommendation is to evaluate workloads separately: keep batch offline inference on GPUs, and separately calculate the unit cost of LPX-class solutions for interactive long-context agents. If in doubt, run the 3-year total cost of ownership of both architectures through the TCO Calculator before deciding.

(Performance data in this article comes from NVIDIA official architecture comparisons and Artificial Analysis third-party benchmarks; for actual deployments, please rely on your own testing.)

GPU Rental Rates Climb Again: Nebius Raises H100/B200 On-Demand Prices About 20% From October, and the Build-vs-Rent Balance Is Shifting

· 5 min read
Industry Research Team

According to ZhiXun (media) reporting in late September, Nebius announced that on-demand rates for H100 / H200 / B200 will rise by an average of about 20% starting October 1. The same month, Cailian Press reported that as of August 11 CoreWeave's contracted power had grown to 4.2 GW and that it would "keep signing new compute at higher prices". This is the Nth consecutive price increase in the rental market this AI cycle — and for buyers, the TCO balance between building and renting is shifting. This article works through the math using the site's TCO calculator methodology.


1. Three Forces Behind the Price Hikes​

  • Exploding agentic inference demand: from MLPerf v6.1 to SemiAnalysis AgentX, agentic workloads have become the main engine of inference growth — token consumption is orders of magnitude beyond chat scenarios, and inference compute has gone from "enough" to "never enough";
  • Supply premium on the new platform generation: Vera Rubin NVL72 has debuted on CoreWeave, and early-ramp pricing on new platforms is naturally elevated — which also raises the anchor for renewal prices on the previous generation (H100/H200/B200);
  • Power has become a hard constraint: CoreWeave's 4.2GW of contracted power shows the data center supply bottleneck has shifted from GPUs to power and racks; enterprises holding power contracts have no reason to cut prices.

2. What a 20% Rent Increase Means for TCO​

The site's TCO calculator splits total cost of ownership into five parts: purchase + energy + operations + network + discounting. Rent increases don't enter that model directly, but they change the "rental baseline" — the opportunity cost a build option must beat:

ScenarioBefore 20% Rent HikeAfter 20% Hike
Short-term elastic workloads (under 1 year)Renting winsRenting still wins (build deployment lead times can't catch up)
Steady inference workloads (2-3 years)Crossover zoneBuild starts to win (opportunity cost rises)
Training clusters (3 years+)Build winsBuild's advantage widens

Rules of thumb:

  • For workloads running 24/7 at full load for 2 years or more, a 20% rent increase is usually enough for building (even at a 1.5x purchase premium and with a PUE 1.3 energy model) to overtake renting within 30 months;
  • Tidal workloads (daytime peaks, idle nights) should still rent — under the assumption that a build pays a continuous 15% idle power draw, building almost never breaks even below roughly 50% utilization;
  • A hybrid strategy (owned baseline + rented peaks) is least sensitive to rent increases and is the robust play in the current interest-rate and power-price environment.

You can re-derive these conclusions in the TCO calculator: enter the "cloud rental unit price" at +20% and compare scenarios across purchase price, electricity price, and PUE. The default discount rate is 8%; purchase costs are not discounted, and annual electricity is discounted by 1.08^-y.

3. Pricing Implications for Domestic Compute​

Rent increases have a layer of impact on domestic AI chip purchasing decisions that is easy to overlook:

  • The rental market for domestic chips is not yet mature: compute from Ascend / Cambricon / Moore Threads is delivered mostly as appliances, full racks, and intelligent computing center allocations; public on-demand rental prices are scarce — so procurement comparisons naturally lean toward "build / full-system" models;
  • Relative TCO shift: a 20% NVIDIA rent increase raises the "relative attractiveness" of every domestic alternative, especially for inference workloads (actual market transaction prices for 910C / 950PR / P800 on the primary market are not public, but the full-system price gap versus dollar-priced cards is widening);
  • Supply cadence: DeepSeek betting on Ascend training and rumors of 160,000-unit 950DT purchases — top demand eats capacity first, and queue times for later buyers are lengthening. Deciding early has value in itself.

4. Takeaways​

  • Nebius raises H100/H200/B200 on-demand rates +20% from October; CoreWeave's contracted power at 4.2GW keeps locking in volume at high prices — rental market supply and demand remain tight;
  • Agentic inference + Rubin supply premium + power constraints: three forces pushing rents up, with no near-term reversal in sight;
  • Decision framework: steady workloads 2 years+ → the build window opens; tidal workloads → keep renting; hybrid strategies are most price-hike resistant;
  • What to watch: whether other clouds follow with price moves, formal Rubin rental pricing, and whether a public rental market for domestic compute emerges.

Further Reading​

References​

  • ZhiXun (media): Nebius raises AI chip rental prices; from October 1, H100/H200/B200 on-demand rates rise about 20% on average (2026-09-21)
  • Cailian Press / East Money: CoreWeave deploys multi-rack Vera Rubin NVL72 cluster, contracted power grows to 4.2 GW (2026-09)

This article is compiled from public reports; TCO conclusions depend on assumptions such as utilization, electricity price, and discount rate. Please re-compute with your own parameters using the TCO calculator.

DeepSeek Confirms Betting on Huawei Chips for LLM Training: From the 160,000-Unit 950DT Rumor to a "Must Succeed" Commitment

· 5 min read
Industry Research Team

On September 22, 2026, multiple media outlets reported that DeepSeek's CEO explicitly stated the company is betting heavily on Huawei chips, expects to receive a new batch of Huawei chips for large-model training, and stressed that "this choice must succeed." Following earlier reports that "DeepSeek planned to procure 160,000 Ascend 950DT units for inference," this is the most significant alignment signal yet in the domestic compute ecosystem — an upgrade from inference procurement to a training bet.


1. The Signal Chain: Four Steps to a "Training Bet"​

Piecing together the public information from the past six months, the DeepSeek–Ascend collaboration shows a clear escalation path:

TimeSignalNature
Mid-2026DeepSeek V4 completes ecosystem migration from CUDA to CANNSoftware adaptation
August 2026Report: DeepSeek plans to procure 160,000 Ascend 950DT units for inference deploymentInference procurement intent
2026-09-17HC2026: 960DT ready three quarters ahead of schedule, 950DT ramping in Q4Supply delivery
2026-09-22DeepSeek CEO: betting on Huawei chips for LLM training, "must succeed"Training bet

The key change is the workload tier: inference deployment means "cutting costs with off-the-shelf compute," while a training bet means "staking the existence of next-generation models on domestic chips" — training clusters demand an order of magnitude more in stability, interconnect efficiency, and software-stack maturity.

2. Huawei's Ability to Deliver​

DeepSeek's willingness to commit rests on a series of verifiable progress points on Huawei's supply side (all official figures):

  • One generation per year, delivered: the 950PR is in mass production, the 950DT ramps in 2026 Q4, and the 960DT moved three quarters earlier than its original 2027 Q4 plan to ready in 2027 Q1 / launch in Q2 — the roadmap's credibility validated twice in a row;
  • Deployment scale: Ascend supernodes have been commercially deployed at scale in over 1,000 sets, covering internet, finance, healthcare, and manufacturing;
  • Ecosystem maturity: CANN has entered routine open-source operation, with external developers exceeding 61% for the first time and 5,200 monthly active developers; there are over 40 Ascend-native training models, making it the only domestic technology route supporting pretraining;
  • Training evidence chain: China Telecom's Xing 4.0-29B-A4B agentic MoE model, open-sourced in September, was announced as trained end-to-end on Ascend (company claim) — "Ascend can train large models" is no longer just Huawei's self-attestation.

See the Ascend 950DT and Ascend 960 spec pages for details; for supernode analysis, see Ascend 960 Official Launch.

3. Why DeepSeek?​

DeepSeek's choice has strong structural drivers:

  1. Supply certainty: under export-control constraints, NVIDIA's flagship supply to China keeps tightening; domestic compute is the only plannable large-scale training supply;
  2. Cost structure: DeepSeek has always been known for extreme engineering efficiency (the V3 training cost set the industry benchmark), and domestic compute plus supernode system efficiency fits its approach;
  3. Betting on ecosystem dividends: CANN open-sourcing plus deep binding with leading model vendors means model vendors can participate in shaping the toolchain's direction — something impossible within CUDA's closed system;
  4. Self-fulfilling demonstration effect: a leading lab's public commitment pulls back on Huawei's production scheduling and upstream HBM and system investment, making "must succeed" a rational commitment rather than a slogan.

4. Risks and Open Questions​

Viewed coolly, this route still has three items that need time to verify:

  • Actual training scale: the quantity, model type (950DT or 960DT), and delivery schedule of the new batch of chips are all undisclosed;
  • MFU methodology: the MFU / latency gains Huawei cites all come from Markov-lab simulations, with no independent third-party measurements yet; the real effective compute of a training cluster depends on long-term data from large-scale production environments;
  • Per-card generation gap: the 960DT's 4 PFLOPS (FP4) versus Rubin R200's 50 PFLOPS (FP4) — the per-card gap objectively exists, and training efficiency depends on whether supernode scale and software optimization can compensate; that is precisely the decisive battleground of "system-level competition."

5. Summary​

  • DeepSeek confirms betting on Huawei chips for training large models — the first training-grade commitment from a leading lab in domestic compute;
  • Signal chain: CUDA→CANN migration → 160,000-unit 950DT inference procurement rumor → 960 ready ahead of schedule → training bet;
  • Supporting factors: one-generation-per-year delivery, 1,000+ supernode sets, and 61% external developers in the CANN open-source ecosystem;
  • What to watch: actual arrival of the new chips and training-cluster scale, third-party MFU data, and training-compute disclosure in DeepSeek's next release.

Further Reading​

References​

  • Toutiao Tech Morning Report: DeepSeek confirms betting on Huawei chips for large-model training (2026-09-22)
  • HUAWEI CONNECT 2026 official announcements (2026-09-17)
  • Earlier report: DeepSeek plans to procure 160,000 Ascend 950DT units (2026-08)

This article is compiled from public reports and vendors' official statements. Details such as chip quantities and training-cluster scale are subject to subsequent disclosures from DeepSeek and Huawei.

Three Signals for the Domestic Compute Ecosystem in One Week: China Telecom Open-Sources Ascend-Trained Model, Inspur 128-Chip Inference Appliance, EVAS Raises RMB 2 Billion

· 5 min read
Industry Research Team

Beyond the official Ascend 960 unveiling and the Zhenwu V900 launch, late September also delivered three ecosystem signals that headline coverage easily drowned out, yet whose structural significance rivals flagship chips: a carrier open-sourcing a model trained on Ascend, an OEM delivering a hundred-chip-class domestic inference system, and a cloud-compute chip startup closing a new funding round. This article breaks down each one.


1. China Telecom Open-Sources Xing4.0: Carrier-Grade Endorsement for End-to-End Ascend Training​

On September 23 (as reported), China Telecom open-sourced Xing4.0-29B-A4B — an agentic mixture-of-experts (MoE) model that the company says was trained end-to-end on Huawei Ascend accelerators.

Why it matters:

  • Third-party proof that "it can train": until now, evidence that "Ascend can train large models" came mainly from Huawei itself (40+ natively trained models); a carrier running training with its own data, engineering teams, and clusters — and open-sourcing it — is the first scaled training case outside the Huawei ecosystem;
  • Agentic positioning: the 29B-A4B (29B total / 4B active) MoE form factor targets agent workloads directly — aligned with the open-source cadence of DeepSeek and Qwen, rather than a small experimental model;
  • Open-source spillover: publishing model weights plus the training recipe means Ascend training engineering know-how can be reused by other institutions — the standard playbook for ecosystem diffusion.

Combined with Huawei's disclosure that "CANN has entered routine open-source operations, with external developers at 61%", the Ascend ecosystem is shifting from "vendor-led" to "community-operated".

2. Inspur MetaBrain SD200 Ultra: Measured Throughput Claims for a 128-Chip Domestic Inference System​

Inspur launched the MetaBrain SD200 Ultra inference system around the same time:

MetricValue (company figures)
Domestic AI chips128 chips (vendor did not disclose specific models)
Model servedKimi K3
Token throughput2.8T tokens (system throughput figure under that methodology)
Latency5.85 ms

Three readings:

  • Hundred-chip-class system integration: 128 chips cooperating across a heterogeneous setup to run very large MoE inference is a test of interconnect topology, scheduling, and fault tolerance — OEMs are now demonstrably capable of assembling domestic chips into large systems;
  • Aimed at top open-source models: targeting Kimi K3 (already live on AWS Bedrock) as the benchmark workload shows the acceptance standard for domestic inference systems is "runs today's most popular models", not a self-referential demo;
  • Mind the methodology: the 2.8T token throughput and 5.85ms latency are both vendor claims; the tested model version, concurrency, and batch size were not disclosed, so procurement evaluations should re-test under real workloads.

3. EVAS Closes RMB 2 Billion B+ Round: The Cloud Compute Chip Race Still Attracts Capital​

EVAS (Yixing Intelligence) completed a B+ round of RMB 2 billion, at a post-money valuation near RMB 15 billion, with Oriza Capital following on; earlier in the first half of the year it had closed a RMB 1.5 billion Series B and brought in China Mobile as a strategic investor. The company focuses on next-generation cloud compute chips.

In 2026, when the flagship chip lineup (Ascend / Cambricon / Moore Threads / Hygon / Zhenwu) looks settled, a cloud compute chip newcomer still raising RMB 2 billion suggests:

  • Primary-market conviction in the "second tier of domestic compute": with top vendors' capacity booked into 2027, overflow demand gives newcomers a window;
  • The role of carrier strategic investment: China Mobile is both a buyer and an investor — the demand side of domestic compute is using capital to lock in the supply side;
  • Differentiated room remains for cloud inference DSA routes (see the ecosystem positioning of routes like Tsingmicro and Houmo).

4. Assembling the Week's Signals: The Loop Is Taking Shape​

Put the three signals together with this month's flagship launches, and every segment of the domestic compute loop now has players filling the gaps:

SegmentThis Month's Evidence
Flagship chipsAscend 960 early official unveiling (9-17), Zhenwu V900 launch (9-22)
Training validationChina Telecom's Xing4.0 end-to-end trained on Ascend and open-sourced, DeepSeek betting on Ascend training
Inference systemsInspur SD200 Ultra running Kimi K3 on 128 chips
Software ecosystemCANN routine open-sourcing, external developers at 61%
Capital supplyEVAS B+ round of RMB 2 billion (post-money near RMB 15 billion)
Demand sideDeepSeek's alignment, Kimi K3 live on overseas cloud platforms

Independent players are investing across all four segments — "chips, systems, models, capital" — the biggest difference from the "isolated breakthroughs" of 2024-2025.

5. Takeaways​

  • China Telecom Xing4.0-29B: the first open-source model end-to-end trained on Ascend outside the Huawei ecosystem;
  • Inspur SD200 Ultra: a 128-chip domestic inference system at 2.8T token throughput / 5.85ms (company figures, pending re-testing);
  • EVAS RMB 2 billion B+ round: the cloud compute chip second tier is still getting real money;
  • What to watch: Xing4.0's training cluster scale and MFU, third-party re-tests of the SD200 Ultra, and EVAS's tape-out cadence.

Further Reading​

References​

  • The GPU Daily (2026-09-24): China Telecom open-sources Xing4.0-29B; Inspur MetaBrain SD200 Ultra
  • Toutiao (2026-09-21): EVAS completes RMB 2 billion B+ round, post-money valuation near RMB 15 billion
  • CSDN AI Daily (2026-09-21): Huawei CANN enters routine open-source operations, OceanStor M900 launched

This article is compiled from public reports. Xing4.0 training details and SD200 Ultra performance figures are company statements; testing conditions are subject to subsequent disclosures.

Alibaba Zhenwu V900 Unveiled at Apsara Conference: 3x M890 Performance, 216GB Memory, 500,000-Card Cluster, Mass Production in 2027 Q1

· 5 min read
Industry Research Team

Just five days after Huawei officially announced the Ascend 960 at HUAWEI CONNECT on September 17, Alibaba unveiled its new unified training-and-inference AI chip, Zhenwu V900, at the Apsara Conference in Hangzhou on September 22 — which Alibaba calls "the most powerful Chinese self-developed AI chip in terms of compute performance to date." This article is compiled from Alibaba's official announcements and reports from Sina Tech, C114, Huanqiu.com, and other sources.


1. Single Chip: 3x M890, 216GB + 1200GB/s​

MetricZhenwu V900 (this launch)Zhenwu M890 (previous generation)
Performance3x Zhenwu M890 (official figure; absolute value undisclosed)FP16 600 TFLOPS (as catalogued on this site)
Memory216 GB144 GB (HBM3)
Chip-to-chip interconnect1200 GB/s—
PrecisionNative FP8 / FP4 (including high-precision training)—
PositioningUnified training and inference (trillion-parameter-scale training + low-precision inference)Unified training and inference
Mass production2027 Q1, scaled deployment in Alibaba Cloud data centersAlready deployed at scale

Three takeaways:

  • Memory crosses into the 200GB+ tier: 216GB puts it in the same capacity class as the Ascend 960DT (288GB self-developed HBM) and NVIDIA Rubin (288GB HBM4), leveling the capacity threshold for long-context and very large MoE models;
  • 1200GB/s chip-to-chip interconnect: laying the foundation for the supernode's "memory-semantics interconnect" — the chip-level prerequisite that pairs with ICN Switch;
  • Precision coverage across all scenarios: officially described as "high-precision training, low-precision and ultra-low-precision inference across all scenarios," with native FP8/FP4 support aligned with the common spec of 2026's new cards.

The absolute single-chip compute figure has not been officially disclosed; this site records it using the official relative figure of "3x M890." See the Zhenwu M890 spec page for M890 details.

2. System Level: ICN Switch + Panjiu Supernode, a 500,000-Card Single Cluster​

V900's real selling point is not the single chip but system-level collaboration:

  • ICN Switch self-developed interconnect chip: once connected, the supernode gains native memory semantics and unified memory addressing, with a thousand cards interconnected at full bandwidth — over a thousand V900s can "work together like a single super chip";
  • Panjiu supernode server: integrates V900 (compute) + ICN Switch (interconnect) + Panmai intelligent NIC (network) + Zhenyue SSD controller (storage), a fully self-developed compute-storage-network stack;
  • 500,000-card single cluster: combined with Alibaba Cloud's next-generation intelligent computing center network architecture, a single cluster can scale up to 500,000 cards.

This mirrors Huawei's "11 key chips" approach: the unit of competition has shifted from the single chip to whole-system delivery capability — Zhenwu handles compute, Yitian handles general-purpose computing, Panmai handles networking, Zhenyue handles storage, and ICN Switch handles interconnect.

3. Business and Roadmap​

  • The Zhenwu family has served over 650+ enterprise customers (as of June 2026), spanning autonomous driving, finance, large models, embodied AI, energy, and manufacturing;
  • Supernodes based on the M890 are already deployed at scale, running models with over 2 trillion parameters such as Qwen3.8 and Kimi K3; Alibaba Cloud will add new serving nodes in Q4 to expand supernode supply;
  • Roadmap: V900 enters mass production and sales in 2027 Q1; Zhenwu J900 is planned for release in 2027 Q3;
  • CPU synergy: Yitian 720 / 730 arrive in 2027 (the 730 is the first to adopt T-Head's fully self-developed CPU microarchitecture, with single-core SPECint2017/GHz up to 1.4x that of Yitian 710); the 2029 Yitian 750 will interconnect directly with Zhenwu AI chips via the ICN bus;
  • Alibaba Group CEO Eddie Wu said T-Head's AI chip annual shipment volume will increase substantially.

4. Competitive Coordinates: Two Swords in One Week​

Viewing the two mid-September launches side by side, the landscape of "system-level competition" among domestic AI chips is now clear:

DimensionHuawei Ascend 960 (9-17)Alibaba Zhenwu V900 (9-22)
Launch eventHC2026Apsara Conference 2026
Per-card memory288GB (self-developed HBM)216GB
Memory bandwidth9.6 TB/sUndisclosed
InterconnectLinJu UnifiedBus / NPO optical interconnectICN Switch memory-semantics interconnect
SupernodeAscend 960 supernode (4,096 cards, 8 EFLOPS FP8)Panjiu supernode (500,000-card single cluster)
Availability2027 Q22027 Q1

Huawei is taking the "supernode + open-source CANN ecosystem" route, while Alibaba is taking the "full cloud stack + open-source model ecosystem" route; both still trail in per-card specs, but what they deliver are procurable, operable ultra-large-scale clusters. The substitution logic against NVIDIA's CUDA ecosystem is shifting from "performance parity" to "supply certainty + system efficiency."

5. Summary​

  • V900: 3x M890, 216GB, 1200GB/s chip-to-chip interconnect, native FP8/FP4, mass production in 2027 Q1;
  • System: ICN Switch with thousand-card full bandwidth, Panjiu supernode with a fully self-developed compute-storage-network stack, 500,000-card single cluster;
  • Cadence: J900 in 2027 Q3, Yitian 730 in 2027, Yitian 750 in 2029 — T-Head's first three-year generational roadmap;
  • What to watch: actual mass production and Alibaba Cloud deployment scale in 2027 Q1, disclosure of V900's absolute compute, and the possibility of external supply of ICN Switch.

Further Reading​

References​

  • Huanqiu.com: T-Head launches new unified training-and-inference AI chip Zhenwu V900 (2026-09-22)
  • Sina Tech / TechWeb: Alibaba T-Head launches new AI chip Zhenwu V900, mass production in the first quarter of 2027
  • C114: Alibaba launches new AI chip Zhenwu V900, mass production in 2027 Q1

This article is compiled from Alibaba's official announcements and public reports. The absolute single-chip compute of V900 has not been officially disclosed; supernode and cluster scale figures are official claims, subject to actual future delivery.

JD Cloud's 100,000-Card Cluster Bets on Moore Threads: Domestic GPUs Enter a Top AI Cloud's Core Compute Base for the First Time

· 6 min read
Industry Research Team

This article is based on official announcements from the 2026 JD Global Technology Explorer Conference (September 9) and public market information; order values and revenue forecasts are brokerage/media estimates, not company announcements.

On September 9, at the 2026 JD Global Technology Explorer Conference, JD Cloud announced a milestone decision: partnering with Moore Threads to build a 100,000-card full-function GPU cluster, creating hyperscale domestic intelligent computing infrastructure.

Two keywords in this sentence deserve amplification: "full-function GPU" and "100,000-card-class core cluster." The former means it must run not just inference but the front lines of large model training, inference, and embodied AI; the latter means a domestic GPU has, for the first time, been placed at the core of a top AI cloud provider's compute base — not a pilot, not an adaptation, not an all-in-one appliance, but 100,000 cards.

1. Partnership Details: From 10,000 to 100,000 Cards — How Big Is the Order?​

  • Prior foundation: JD Cloud has already built a domestic 10,000-card cluster with Moore Threads and other partners; this move is a magnitude leap from 10,000 to 100,000 cards
  • Order scale: According to market and brokerage information, GPU modules are supplied exclusively by Moore Threads, with a total order value of about RMB 20-30 billion; revenue can be recognized as early as next year according to delivery cadence. Some brokerages have raised Moore Threads' revenue expectation for next year to RMB 15-20 billion (this estimate is not a company announcement; refer to official disclosures)
  • Partnership depth: Full-stack coordination from chips and cloud platform to model training, supporting the iteration of JD's JoyAI model family and forming a closed loop of "data, training, simulation, deployment"
  • Openness: Compute is open to all industries, focused on large model training, inference, and embodied AI

Moore Threads founder Zhang Jianzhong was direct on stage: "The Scaling Law still holds — 100,000-card clusters are an inevitable trend." JD Cloud President Cao Peng positioned domestic compute as a core pillar of JD's physical AI strategy.

2. Why Moore Threads? Three Calculations Behind the Procurement Logic​

Tech companies buy cards with no sentiment involved. JD's choice of Moore Threads comes down to three calculations that all add up:

1. The stability calculation: Moore Threads has commercially deployed thousand-card and 10,000-card large clusters under a single network, and has achieved breakthroughs in core training scenarios such as foundation models, embodied brains, and world models — the engineering validation of a 10,000-card cluster is the prerequisite for 100,000 cards; this is not a cold start.

2. The integration calculation: Full-function GPUs can plug into existing IT systems and cloud platform scheduling, and the MUSA software stack's adaptation to large model frameworks has already passed training-grade workloads.

3. The ROI calculation: This is the most critical shift. Buyers of domestic GPUs used to be mostly "policy-friendly" projects; JD writing a 100,000-card cluster into its capital expenditure means the product's return on investment can now stand up to investor scrutiny — the buyer structure shifting from "daring to use" to "rushing to use" is the hallmark of commercial maturity for domestic GPUs.

3. S6000: Next-Gen Chip Taped Out and Back, with a Dual-Supply Safeguard​

According to market information, Moore Threads' next-generation chip, the S6000, has successfully come back from the fab and been distributed to vendors for testing, with ample FAB and memory supply guarantees. Note that the S6000 has not been officially released and its specifications are not public; this article makes no speculation. The on-sale flagship MTT S5000 has on-site data of 400 TFLOPS FP16 and 80GB of memory (the 1.6TB/s bandwidth is HBM-class).

Regarding HBM supply constraints, Moore Threads' response strategy is reportedly "next-generation product iteration + multi-source supply chain safeguards" — until domestic HBM capacity ramp-up is complete, this is the same problem every domestic GPU vendor must solve.

4. "100,000 Cards" Is Not One Company's Game: The Domestic GPU Cluster Landscape​

It is worth widening the view — 100,000 cards is now a collective goal for domestic compute:

Player100,000-Card MoveCompute Base
JD CloudAnnounced co-built 100,000-card cluster on September 9Moore Threads full-function GPU
Sugon 8000Released at WAIC in July, completed in Zhengzhou; a fully domestic 100,000-card AI compute clusterHygon DCU (of the Shensuan BW1000 family)
HuaweiAtlas 950 SuperPoD super node (8,192 cards); 100,000-card super node in 2027Ascend 950/960

Two chip routes (full-function GPU vs DCU vs NPU) and two paths (commercial cloud vs national supercomputing) point to the same validation question: can domestic compute reliably run real training workloads at 100,000-card scale. It is worth emphasizing that there is currently no public third-party benchmark comparison for the Moore Threads x JD cluster, and no completion timetable — going from "usable" to "running well" at 100,000 cards still requires engineering validation.

5. The Capital Markets Perspective​

Moore Threads listed on the STAR Market in December 2025: an IPO price of RMB 114.28, closing at RMB 600.5 on the first day; 2025 revenue of RMB 1.505 billion (up 243.37% year over year), with a net loss of about RMB 1.001 billion. If the brokerage-raised revenue expectation of RMB 15-20 billion materializes, 2027 will be the key inflection point from "high-growth loss-making" to "profitable at scale" — which would also become the first complete answer sheet for the domestic GPU business model.

Summary​

From debuting with DeepSeek all-in-one appliances in 2025 to entering a 100,000-card core cluster in 2026, domestic GPUs completed the cognitive leap from "usable" to "commercially viable at scale" in under two years. The value of JD Cloud's order lies not in its amount but in its significance as a sample: when top internet buyers begin procuring domestic GPUs on ROI logic, substitution is no longer a policy narrative but a commercial fact.

(Order values, revenue forecasts, and supply chain status come from public market information and brokerage research estimates; refer to company announcements; chip specification data is available in the on-site full comparison table.)

昇腾 960 官宣发布:FP4 算力翻倍、全球首个 NPO 超节点,华为确认一年一代

· 6 min read
Industry Research Team

2026 年 9 月 17 日,华为全联接大会(HC2026),华为常务董事、ICT 基础设施业务总裁汪涛正式发布昇腾 960。与 8 月数字中国峰会上的"路线图预告"不同,这次是完整的规格官宣——而且提前三个季度就绪。本文基于华为官网新闻稿及多方报道整理,逐项拆解 960 的规格、超节点与路线图含义。


1. 昇腾 960:双版本,训练先行​

昇腾 960 分为两个版本,节奏错开一个季度:

指标昇腾 960DT(训练版)昇腾 960PR(推理版)
FP8 算力2 PFLOPS(较 950 翻倍)待披露
FP4 算力4 PFLOPS待披露
显存288GB(自研 HBM)待披露
显存带宽9.6 TB/s待披露
配套超节点Atlas 860(风冷)Atlas 960(液冷)
就绪 / 上市2027 Q1 / 2027 Q22027 Q3

三个读数:

  • 提前三个季度就绪:960DT 原计划 2027 年 Q4,现在 2027 Q1 就绪、Q2 上市——与 950DT"提前上线华为云"一脉相承,路线图可信度在持续兑现;
  • 精度口径与英伟达对齐:FP8 / FP4 主口径(辅以自研 HiF8 / HiFP8),FP4 已是 2026 新卡的通用口径;
  • 288GB 自研 HBM + 9.6TB/s:显存容量与带宽同步翻倍,对训练长上下文大模型是实打实的容量红利。

单卡规格详见昇腾 960 规格页;上一代昇腾 950DT 规格页。

2. 昇腾 960 超节点:全球首个 NPO 超节点​

这是本场发布真正的"重器"——昇腾 960 超节点是全球第一个采用 NPO(近封装光学)的超节点:

指标昇腾 960 超节点
单节点规模4096 卡
系统算力8 EFLOPS FP8 / 16 EFLOPS FP4
HBM 总容量1 PB
互联 RTT低至 2μs
可用度99.8%

NPO 意味着什么:传统光模块挂在交换机面板上,信号要先出封装再转电—光;NPO 把光引擎挪到交换芯片封装近旁,大幅缩短 SerDes 距离、降低功耗与时延。华为的具体数字:

  • 5500 个自研 Hi-ONE 光引擎(业界首个量产 NPO,单引擎 7.2T);
  • 替代 4.8 万颗 800G 光模块;
  • 降低功耗 550kW 以上,同时支撑 2μs 级互联时延。

对比 8 月数字中国峰会的口径(Atlas 960 单超节点 15488 卡),官方最新口径为单超节点 4096 卡(1 个超集群可由多超节点组成)——以华为官网最新新闻稿为准。

⚠️ 华为还给出了"MFU 提升 2.75 倍""推理时延降低 70%"等收益数据,均出自华为马尔科夫实验室仿真,尚无独立第三方实测,阅读时注意口径。

3. 一年一代:970(2028)→ 980(2029)​

华为首次把"一年一代"从愿景变成官宣承诺:

年份产品状态
2026昇腾 950 系列950PR 已量产,950DT Q4 放量
2027昇腾 960本次官宣,提前就绪
2028昇腾 970HC2026 确认
2029昇腾 980HC2026 确认

支撑这一节奏的是华为提出的"韬定律":算力规格每代翻倍,访存带宽、访存容量、互联带宽同步大幅提升。

生态侧的同场数据也值得记录:CANN 外部开发者占比首次超过 61%、月活开发者 5200 人;910C 超节点部署已超 1000 套。

4. 竞争坐标:单卡有代差,系统级对打​

把 960DT 放到 2027 年的棋盘上看:

  • 对 NVIDIA Rubin(R200:288GB HBM4 / 50 PFLOPS FP4):单卡 FP4 算力约为 R200 的 8%(4 vs 50 PFLOPS),单卡代差客观存在;但 4096 卡超节点 + NPO 互联是系统级竞争——用"可交付的超大集群"对打单卡性能;
  • 对采购方:960 的价值主张是"在国产量产约束内拿到最大可用集群"。DeepSeek 拟采购 16 万颗 950DT 跑推理的订单已经证明:当供给与生态到位,头部实验室愿意主动选国产;
  • 对国产链条:950→960 的显存翻倍直接拉动国产 HBM 迭代,这是整条链的胜负手。

5. 小结​

  • 960DT:FP8 2P / FP4 4P、288GB HBM、9.6TB/s,2027 Q2 上市,提前三个季度;
  • 960 超节点:全球首个 NPO 超节点,4096 卡 / 8 EFLOPS FP8 / 1PB HBM / RTT 2μs;
  • 路线图:一年一代官宣至 2029,950、960 均提前兑现,规划可信度显著上升;
  • 看什么:2027 Q2 实际交付节奏、国产 HBM 产能、CANN 生态的第三方模型覆盖度。

相关阅读​

参考资料​

  • 华为官网:发布全球首个采用 NPO 的超节点(昇腾 960 超节点)(2026-09-17)
  • 电子工程专辑:国产 AI 芯片新突破!华为昇腾 960 芯片将提前发布
  • 环球网:华为全联接大会 2026 主题演讲报道

本文基于华为官网新闻稿与公开报道整理。MFU、时延等收益数据为华为实验室仿真口径;2028 / 2029 产品仅为路线图官宣,规格以未来发布为准。

Vera Rubin NVL72 MLPerf v6.1 Debut: Qwen3-VL Throughput 3.7x GB300, CoreWeave Brings Multi-Rack Cluster Online Same Day

· 5 min read
Industry Research Team

On September 16, MLCommons released the MLPerf Inference v6.1 Closed Division results, and NVIDIA's Vera Rubin NVL72 completed its benchmark debut. The same day, CoreWeave announced that a multi-rack Vera Rubin cluster was live on its cloud — a "rent the new hardware on launch day" cadence that is compressing the cycle from next-generation compute announcement to billable output down to a quarter.


1. Debut Results: 3.7x and 2.5x​

NVIDIA, via preview submissions (entries 6.1-0106 / 6.1-0074), provided an apples-to-apples comparison of its two flagship generations:

Benchmark ModelVera Rubin NVL72 vs GB300 NVL72Software Stack
Qwen3-VL-235B-A22B (235B multimodal MoE)Throughput up to 3.7x (across offline / server / interactive scenarios)vLLM + NVIDIA Dynamo
DeepSeek-R1-671B (671B inference model)Throughput up to 2.5xTensorRT-LLM

The three technologies behind these numbers:

  • NVFP4 precision: compresses the memory footprint of model weights, attention tensors, and the KV Cache, enabling significantly larger effective batch sizes;
  • Disaggregated serving: prefill (compute-intensive) and decode (bandwidth-intensive) are split across separate GPU pools, each independently optimized;
  • Expert parallelism: MoE expert layers route across chips in a distributed fashion, paired with sixth-generation NVLink / NVLink Switch (NVIDIA claims 10x the packet rate of commodity Ethernet at 3x lower latency).

2. Same-Session Highlights: Scaling Efficiency and the Software Dividend​

  • GB300 four racks, 288 chips, 99% scaling efficiency: in the DeepSeek-R1 offline scenario, scaling from a single rack of 72 chips to 4 racks of 288 chips shows almost no loss (entries 6.1-0073 / 6.1-0074) — the first time rack-scale interconnect scaling efficiency has been validated at the 288-chip level;
  • The software dividend is still being paid out: on the same GB300 hardware, v6.1 software optimizations deliver up to 1.6x Qwen3-VL gains over v6.0; NVIDIA also reported post-submission gains for GPT-OSS-120B and DLRMv3 (not verified by MLCommons);
  • 19 partners submitted, 8 of them using multi-node Blackwell NVL72 configurations (ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell, Fujitsu, HPE, Lambda, Nebius, OCI, Supermicro, and others); Nebius also submitted Vera Rubin preview results.

3. CoreWeave Online the Same Day: Rentable at Launch​

On the same day the MLPerf results were published, CoreWeave announced that its multi-rack Vera Rubin NVL72 cluster was live on its cloud:

  • Spectrum-X networking links hundreds of Rubin GPUs into a single scale-out cluster, with each GPU paired with 2 ConnectX-9 SuperNICs (1.6Tb/s scale-out bandwidth);
  • The AI object store LOTA cuts read latency by 8x;
  • On the operations side: Valvey software-defined liquid cooling (NVIDIA spec: 45°C liquid inlet), the Racky rack control layer, and Rack LifeCycle Controller manage an entire NVL72 rack as a single programmable entity;
  • Back in July, CoreWeave's own testing put Vera Rubin at 10x tokens/s/MW versus GB200 NVL72 on DeepSeek R1 (company figures, not MLCommons-verified).

4. The Competitive Picture​

  • AMD + Crusoe: set MLPerf's all-time high aggregate throughput of 5.75M tokens/s on 512 chips (AMD had the broadest submission coverage this round);
  • SemiAnalysis AgentX preview: Vera Rubin up to 30x over GB300 on agentic workloads (preview figures); the MLPerf Endpoints benchmark is forthcoming and will bring agentic inference into standardized measurement;
  • Pending validation: Vera Rubin results are preview submissions awaiting independent reproduction; no quantified quality metrics were given for NVFP4's "virtually lossless" claim; per-token cost and power-consumption scaling comparisons have not been published. Matching submissions from Google TPU v7 and AMD MI455X UALoE72 are expected within 2026.

On-site spec comparisons: Rubin R200 (288GB / 22TB/s), B300 Ultra (GB300-class core).

5. Takeaways​

  • Vera Rubin NVL72 debut: Qwen3-VL 3.7x, DeepSeek-R1 2.5x over GB300 NVL72 (preview figures);
  • GB300 across four racks of 288 chips at 99% scaling efficiency — rack-scale interconnects have entered the usable range;
  • CoreWeave brought a multi-rack cluster online the same day — the cycle from next-gen compute launch to billable output is now measured in quarters;
  • What to watch: independent reproductions and formal (non-preview) submissions, TPU v7 / MI455X comparison results, and measured per-MW token throughput.

Further Reading​

References​

  • NVIDIA Blog: Vera Rubin NVL72 Makes Its First Appearance in MLPerf Inference v6.1 (2026-09-16)
  • MLCommons: MLPerf Inference v6.1 Closed Division results (entries 6.1-0073 / 6.1-0074 / 6.1-0106)
  • CoreWeave Investor Relations: multi-rack Vera Rubin NVL72 availability announcement (2026-09-16)
  • AMD Blogs: MLPerf Inference v6.1 submission results

This article is compiled from public MLCommons results and vendor official statements. Vera Rubin results are preview submissions; some gain figures are not MLCommons-verified and are flagged item by item.