Skip to main content

7 posts tagged with "AI Compute"

AI compute supply, demand and industry trends

View all tags

Groq 3 LPX Enters Full Mass Production: Samsung 4nm Foundry, 315 PFLOPS FP8 per Rack — NVIDIA Turns the LPU into the Seventh Chip of Vera Rubin

· 5 min read
Industry Research Team

This article is based on NVIDIA official product pages, the March 2026 GTC architecture blog, and third-party benchmark data; performance figures are vendor architecture comparisons, and actual gains should be verified with your own tests.

NVIDIA has confirmed that Groq 3 LPX — the "seventh chip" of the Vera Rubin platform — has entered full mass production, manufactured at Samsung's Pyeongtaek campus. This is the first time the deal has landed in the form of production silicon since December 2025, when NVIDIA spent roughly $20 billion to obtain a non-exclusive IP license to Groq plus its core engineering team.

For the inference hardware landscape, this is a paradigm-level event: for the first time, NVIDIA has incorporated a "specialized inference architecture defined by someone else" into its own flagship platform — and not as an add-on sale, but as deep co-design.

1. The Groq 3 LPU Chip and the LPX Rack: Specs at a Glance​

Single Groq 3 LPU (4nm, Samsung foundry):

ItemValue
Compiler-managed on-chip SRAM500 MB
SRAM bandwidth150 TB/s
Chip-to-chip scale-up bandwidth2.5 TB/s (96 chip-to-chip links, 112 Gbps each)

A single LPX rack (32 1U liquid-cooled trays, 8 LPUs per tray):

ItemValue
Total LPUs256
FP8 compute315 PFLOPS
Total on-chip SRAM128 GB
Aggregate SRAM bandwidth40 PB/s
In-rack scale-up bandwidth640 TB/s
DDR5 memory12 TB (hosting large model weights)

The 500 MB of SRAM per LPU may not look like much, but multiplied by 256 LPUs and stacked with 40 PB/s of aggregate bandwidth, it forms the physical foundation of "deterministic low-latency decoding" — the LPU's design philosophy is precisely to use compiler static scheduling of on-chip SRAM to completely eliminate memory-fetch stalls during the inference decoding phase, which is exactly the most typical bottleneck of GPU inference.

2. AFD: Attention and FFN Split Up, GPU and LPU Each Do Their Own Job​

LPX does not replace the GPU; instead, it forms a heterogeneous system with Vera Rubin NVL72, centered on AFD (Attention-FFN Disaggregation):

  1. Rubin GPUs handle prefill and attention — building the KV cache over large contexts and executing attention layers, consuming HBM capacity and high throughput;
  2. Groq 3 LPUs handle FFN / MoE expert-layer decoding — latency-sensitive and pattern-predictable, exactly the home turf of deterministic SRAM scheduling;
  3. Intermediate activations are exchanged between the two engines token by token, orchestrated and routed by NVIDIA Dynamo: the GPU computes every attention layer, the LPU computes every feed-forward layer, jointly producing each output token.

The logic of this division of labor is clear: agentic AI applications can consume 15x the tokens of traditional AI applications, and the bottleneck shifts from "can it finish computing" to "can it keep emitting tokens at stable low latency." Let the big HBM container run attention, and let the deterministic SRAM engine run decode — each plays to its strengths.

3. Measured Results and Official Figures​

  • Third-party benchmarks (Artificial Analysis): with Gemma 4 31B at a 100K token context, LPX output reaches roughly 3400 tokens/s — inter-token intervals below 1 millisecond, a qualitative leap in interactivity for long-context agentic scenarios
  • Official architecture comparisons: with LPX added, Vera Rubin NVL72 achieves up to 35x higher throughput per megawatt on trillion-parameter models; up to 10x more revenue opportunity per watt in "high-value token" scenarios
  • Launch customers: Nebius is the first to deploy LPX racks to expand its token capacity; CoreWeave already connects Vera Rubin racks in production with Spectrum-X Multiplane

It must be emphasized: the 35x/10x figures are architecture comparisons at specific high-interaction operating points, not universal conclusions. For training, high-throughput batch inference, and workloads that need CUDA ecosystem flexibility, the GPU remains the right answer; LPX's home turf is scenarios where "single-user interactive latency is the product" — agent loops, real-time coding assistants, long-context conversations.

4. Three Industry Signals​

1. Specialized inference chips get absorbed, not opposed. The old narrative of the LPU as a "GPU challenger" has become part of Vera Rubin. The endgame for inference hardware may not be one architecture winning, but the heterogeneous combination of "GPU + specialized decoding engines" becoming the standard. For other inference chip startups (Cerebras, Etched, etc.), this is both proof that the ceiling has risen and a warning: get integrated or find differentiation.

2. Samsung foundry lands a high-end AI order. Against the backdrop of TSMC's near-monopoly on AI main chips, Samsung 4nm taking on LPX mass production is highly significant — combined with Tesla's earlier AI6 2nm order, Samsung foundry has a shot at returning to full-year profitability in 2027.

3. Inference economics enters the "priced per MW" era. When vendors start telling their story with tokens/MW and revenue per watt, the core KPI of compute selection has completely shifted from peak compute (TFLOPS) to token output per unit of energy. This is consistent with the power-and-electricity cost model built into our TCO Calculator: the chip with the best-looking peak specs is not necessarily the chip with the lowest cost per token.

Summary​

Groq 3 LPX mass production marks the arrival of a heterogeneous era for inference hardware: "GPUs manage throughput, LPUs manage latency." For teams currently selecting inference clusters, the recommendation is to evaluate workloads separately: keep batch offline inference on GPUs, and separately calculate the unit cost of LPX-class solutions for interactive long-context agents. If in doubt, run the 3-year total cost of ownership of both architectures through the TCO Calculator before deciding.

(Performance data in this article comes from NVIDIA official architecture comparisons and Artificial Analysis third-party benchmarks; for actual deployments, please rely on your own testing.)

GPU Rental Rates Climb Again: Nebius Raises H100/B200 On-Demand Prices About 20% From October, and the Build-vs-Rent Balance Is Shifting

· 5 min read
Industry Research Team

According to ZhiXun (media) reporting in late September, Nebius announced that on-demand rates for H100 / H200 / B200 will rise by an average of about 20% starting October 1. The same month, Cailian Press reported that as of August 11 CoreWeave's contracted power had grown to 4.2 GW and that it would "keep signing new compute at higher prices". This is the Nth consecutive price increase in the rental market this AI cycle — and for buyers, the TCO balance between building and renting is shifting. This article works through the math using the site's TCO calculator methodology.


1. Three Forces Behind the Price Hikes​

  • Exploding agentic inference demand: from MLPerf v6.1 to SemiAnalysis AgentX, agentic workloads have become the main engine of inference growth — token consumption is orders of magnitude beyond chat scenarios, and inference compute has gone from "enough" to "never enough";
  • Supply premium on the new platform generation: Vera Rubin NVL72 has debuted on CoreWeave, and early-ramp pricing on new platforms is naturally elevated — which also raises the anchor for renewal prices on the previous generation (H100/H200/B200);
  • Power has become a hard constraint: CoreWeave's 4.2GW of contracted power shows the data center supply bottleneck has shifted from GPUs to power and racks; enterprises holding power contracts have no reason to cut prices.

2. What a 20% Rent Increase Means for TCO​

The site's TCO calculator splits total cost of ownership into five parts: purchase + energy + operations + network + discounting. Rent increases don't enter that model directly, but they change the "rental baseline" — the opportunity cost a build option must beat:

ScenarioBefore 20% Rent HikeAfter 20% Hike
Short-term elastic workloads (under 1 year)Renting winsRenting still wins (build deployment lead times can't catch up)
Steady inference workloads (2-3 years)Crossover zoneBuild starts to win (opportunity cost rises)
Training clusters (3 years+)Build winsBuild's advantage widens

Rules of thumb:

  • For workloads running 24/7 at full load for 2 years or more, a 20% rent increase is usually enough for building (even at a 1.5x purchase premium and with a PUE 1.3 energy model) to overtake renting within 30 months;
  • Tidal workloads (daytime peaks, idle nights) should still rent — under the assumption that a build pays a continuous 15% idle power draw, building almost never breaks even below roughly 50% utilization;
  • A hybrid strategy (owned baseline + rented peaks) is least sensitive to rent increases and is the robust play in the current interest-rate and power-price environment.

You can re-derive these conclusions in the TCO calculator: enter the "cloud rental unit price" at +20% and compare scenarios across purchase price, electricity price, and PUE. The default discount rate is 8%; purchase costs are not discounted, and annual electricity is discounted by 1.08^-y.

3. Pricing Implications for Domestic Compute​

Rent increases have a layer of impact on domestic AI chip purchasing decisions that is easy to overlook:

  • The rental market for domestic chips is not yet mature: compute from Ascend / Cambricon / Moore Threads is delivered mostly as appliances, full racks, and intelligent computing center allocations; public on-demand rental prices are scarce — so procurement comparisons naturally lean toward "build / full-system" models;
  • Relative TCO shift: a 20% NVIDIA rent increase raises the "relative attractiveness" of every domestic alternative, especially for inference workloads (actual market transaction prices for 910C / 950PR / P800 on the primary market are not public, but the full-system price gap versus dollar-priced cards is widening);
  • Supply cadence: DeepSeek betting on Ascend training and rumors of 160,000-unit 950DT purchases — top demand eats capacity first, and queue times for later buyers are lengthening. Deciding early has value in itself.

4. Takeaways​

  • Nebius raises H100/H200/B200 on-demand rates +20% from October; CoreWeave's contracted power at 4.2GW keeps locking in volume at high prices — rental market supply and demand remain tight;
  • Agentic inference + Rubin supply premium + power constraints: three forces pushing rents up, with no near-term reversal in sight;
  • Decision framework: steady workloads 2 years+ → the build window opens; tidal workloads → keep renting; hybrid strategies are most price-hike resistant;
  • What to watch: whether other clouds follow with price moves, formal Rubin rental pricing, and whether a public rental market for domestic compute emerges.

Further Reading​

References​

  • ZhiXun (media): Nebius raises AI chip rental prices; from October 1, H100/H200/B200 on-demand rates rise about 20% on average (2026-09-21)
  • Cailian Press / East Money: CoreWeave deploys multi-rack Vera Rubin NVL72 cluster, contracted power grows to 4.2 GW (2026-09)

This article is compiled from public reports; TCO conclusions depend on assumptions such as utilization, electricity price, and discount rate. Please re-compute with your own parameters using the TCO calculator.

JD Cloud's 100,000-Card Cluster Bets on Moore Threads: Domestic GPUs Enter a Top AI Cloud's Core Compute Base for the First Time

· 6 min read
Industry Research Team

This article is based on official announcements from the 2026 JD Global Technology Explorer Conference (September 9) and public market information; order values and revenue forecasts are brokerage/media estimates, not company announcements.

On September 9, at the 2026 JD Global Technology Explorer Conference, JD Cloud announced a milestone decision: partnering with Moore Threads to build a 100,000-card full-function GPU cluster, creating hyperscale domestic intelligent computing infrastructure.

Two keywords in this sentence deserve amplification: "full-function GPU" and "100,000-card-class core cluster." The former means it must run not just inference but the front lines of large model training, inference, and embodied AI; the latter means a domestic GPU has, for the first time, been placed at the core of a top AI cloud provider's compute base — not a pilot, not an adaptation, not an all-in-one appliance, but 100,000 cards.

1. Partnership Details: From 10,000 to 100,000 Cards — How Big Is the Order?​

  • Prior foundation: JD Cloud has already built a domestic 10,000-card cluster with Moore Threads and other partners; this move is a magnitude leap from 10,000 to 100,000 cards
  • Order scale: According to market and brokerage information, GPU modules are supplied exclusively by Moore Threads, with a total order value of about RMB 20-30 billion; revenue can be recognized as early as next year according to delivery cadence. Some brokerages have raised Moore Threads' revenue expectation for next year to RMB 15-20 billion (this estimate is not a company announcement; refer to official disclosures)
  • Partnership depth: Full-stack coordination from chips and cloud platform to model training, supporting the iteration of JD's JoyAI model family and forming a closed loop of "data, training, simulation, deployment"
  • Openness: Compute is open to all industries, focused on large model training, inference, and embodied AI

Moore Threads founder Zhang Jianzhong was direct on stage: "The Scaling Law still holds — 100,000-card clusters are an inevitable trend." JD Cloud President Cao Peng positioned domestic compute as a core pillar of JD's physical AI strategy.

2. Why Moore Threads? Three Calculations Behind the Procurement Logic​

Tech companies buy cards with no sentiment involved. JD's choice of Moore Threads comes down to three calculations that all add up:

1. The stability calculation: Moore Threads has commercially deployed thousand-card and 10,000-card large clusters under a single network, and has achieved breakthroughs in core training scenarios such as foundation models, embodied brains, and world models — the engineering validation of a 10,000-card cluster is the prerequisite for 100,000 cards; this is not a cold start.

2. The integration calculation: Full-function GPUs can plug into existing IT systems and cloud platform scheduling, and the MUSA software stack's adaptation to large model frameworks has already passed training-grade workloads.

3. The ROI calculation: This is the most critical shift. Buyers of domestic GPUs used to be mostly "policy-friendly" projects; JD writing a 100,000-card cluster into its capital expenditure means the product's return on investment can now stand up to investor scrutiny — the buyer structure shifting from "daring to use" to "rushing to use" is the hallmark of commercial maturity for domestic GPUs.

3. S6000: Next-Gen Chip Taped Out and Back, with a Dual-Supply Safeguard​

According to market information, Moore Threads' next-generation chip, the S6000, has successfully come back from the fab and been distributed to vendors for testing, with ample FAB and memory supply guarantees. Note that the S6000 has not been officially released and its specifications are not public; this article makes no speculation. The on-sale flagship MTT S5000 has on-site data of 400 TFLOPS FP16 and 80GB of memory (the 1.6TB/s bandwidth is HBM-class).

Regarding HBM supply constraints, Moore Threads' response strategy is reportedly "next-generation product iteration + multi-source supply chain safeguards" — until domestic HBM capacity ramp-up is complete, this is the same problem every domestic GPU vendor must solve.

4. "100,000 Cards" Is Not One Company's Game: The Domestic GPU Cluster Landscape​

It is worth widening the view — 100,000 cards is now a collective goal for domestic compute:

Player100,000-Card MoveCompute Base
JD CloudAnnounced co-built 100,000-card cluster on September 9Moore Threads full-function GPU
Sugon 8000Released at WAIC in July, completed in Zhengzhou; a fully domestic 100,000-card AI compute clusterHygon DCU (of the Shensuan BW1000 family)
HuaweiAtlas 950 SuperPoD super node (8,192 cards); 100,000-card super node in 2027Ascend 950/960

Two chip routes (full-function GPU vs DCU vs NPU) and two paths (commercial cloud vs national supercomputing) point to the same validation question: can domestic compute reliably run real training workloads at 100,000-card scale. It is worth emphasizing that there is currently no public third-party benchmark comparison for the Moore Threads x JD cluster, and no completion timetable — going from "usable" to "running well" at 100,000 cards still requires engineering validation.

5. The Capital Markets Perspective​

Moore Threads listed on the STAR Market in December 2025: an IPO price of RMB 114.28, closing at RMB 600.5 on the first day; 2025 revenue of RMB 1.505 billion (up 243.37% year over year), with a net loss of about RMB 1.001 billion. If the brokerage-raised revenue expectation of RMB 15-20 billion materializes, 2027 will be the key inflection point from "high-growth loss-making" to "profitable at scale" — which would also become the first complete answer sheet for the domestic GPU business model.

Summary​

From debuting with DeepSeek all-in-one appliances in 2025 to entering a 100,000-card core cluster in 2026, domestic GPUs completed the cognitive leap from "usable" to "commercially viable at scale" in under two years. The value of JD Cloud's order lies not in its amount but in its significance as a sample: when top internet buyers begin procuring domestic GPUs on ROI logic, substitution is no longer a policy narrative but a commercial fact.

(Order values, revenue forecasts, and supply chain status come from public market information and brokerage research estimates; refer to company announcements; chip specification data is available in the on-site full comparison table.)

Vera Rubin 全面量产:100% 全液冷 + 800V 直流供电,AI 数据中心基础设施范式重构

· 6 min read
Industry Research Team

2026 年 9 月初,供应链信息确认:英伟达 Vera Rubin 平台已于 8 月正式量产、9 月启动批量出货,无延期、无卡顿。与 Blackwell 迭代初期的产能波折不同,这次量产节奏异常平稳——谷歌云、微软 Azure、CoreWeave、甲骨文云等头部云厂商已启动机架部署。但真正值得产业记住的,不是"又一代 GPU 量产了",而是 Rubin 把液冷从"可选配置"变成了"硬性前置条件"。


1. 量产节奏:史上最平稳的一次平台切换​

根据产业链调研与券商跟踪信息:

  • 2026 年 8 月:Vera Rubin 正式量产;
  • 2026 年 9 月:批量出货启动;
  • 2026 下半年:CoreWeave、谷歌云、微软 Azure、甲骨文云机架部署落地;
  • 2026 年:上代 GB 架构机柜出货量有望达 6 万台(同比翻倍);
  • 2027 年:GB 与 Rubin 两代平台合计出货体量有望接近 10 万台,Rubin 新机柜远期产能目标为每天 1000 个 NVL72 机柜。

需求侧同样在加码:华尔街报告披露,英伟达管理层表示 FY28 同比增长 70% 的目标并非需求上限——若供应不受限,增速可能超过 100%。当前主要约束已从需求端转向先进晶圆与 HBM 供应。

2. 单卡 2300W:风冷时代的终结​

Rubin 平台与前代最根本的差异不在算力,而在功耗密度:

指标H100GB300Rubin
单 GPU TDP700W~1400W2300W
机柜功耗~40kW~140kW190–230kW
散热方案风冷为主风液混合100% 全液冷
供电架构48V48V800V 高压直流

单芯片 TDP 从 700W 升至 2300W、单机柜功率密度突破风冷物理极限——这意味着 液冷不再是高端算力的选配升级,而是运行 Rubin 服务器的先决条件。英伟达官方将 Rubin 全液冷架构定义为"数据中心历史上最重要的能效突破之一",并已写入 DSX AI 工厂参考设计:所有跟随英伟达技术路线的云厂商和数据中心运营商,都必须采用全面液冷方案。

三个关键架构变化:

  1. 无风扇整机:GPU、CPU、交换机、DPU 全部器件强制采用直接冷板式液冷,45℃ 温水冷板成为出厂标配;
  2. 液冷边界延伸:散热覆盖范围从 GPU 冷板延伸至 CPU、DPU、交换机乃至光模块(液冷 Cage/鼠笼开始从"可选"变"刚需"),整套液冷系统价值量较 GB300 提升约 40%;
  3. 800VDC 供电:替代传统 48V 机架配电,整机电源 BOM 价值增长 30% 以上,PSU 电源模块从 5.5kW 向 18.3kW 迭代,固态变压器、高压直流 CDU 成为数据中心新增核心设备。

3. 对产业链的三重传导​

第一重:液冷从"配套"变"主角"。 2026 下半年以小规模部署验证为主,真正的放量窗口在 2027 年——Rubin 机架大规模铺货后,冷板、快速接头、CDU、液冷泵进入业绩兑现期。台系供应链 7 月数据已率先验证:AVC 奇鋐 7 月营收 185.9 亿新台币创历史新高(同比 +57.4%),双鸿、健策 7 月同比分别 +116.7%、+91.0%。

第二重:国产液冷供应链进入核心 BOM。 国内厂商由外围冷源和代工环节逐步进入芯片平台、服务器 ODM 和海外云厂商供应体系,替代路径从 Manifold、管路推进至高可靠快接头和冷板。英维克 26H1 海外收入占比 71.4%,飞龙股份液冷泵小功率平台订单超 5 万台——液冷全核心零部件自主可控正在成为现实。

第三重:供电与散热边界融合。 800VDC 架构下,电源模块、PDB 配电单元、高速交换芯片自身发热也达到很高水平,部分电源组件同样需要液冷辅助散热——电源与温控两条产业链正在合并成一条。

4. 需求矩阵扩容:云厂商之外,太空算力入场​

Rubin 的客户矩阵已从传统云厂商扩展至三个层次:

  • 全球云厂商:谷歌云、Azure、甲骨文云、CoreWeave;
  • AI 科技巨头:马斯克公开披露 2027 年 8GW 超大规模 IDC 建设规划;SpaceX 将 Vera Rubin 架构定义为"最优 AI 计算架构",计划地面与太空双向部署,支撑 "Starmind" 卫星算力项目;
  • 主权与边缘:远期 Rubin Ultra 及 2027 年后更高功耗机型单机柜有望冲击 600kW+。

普华永道预计全球数据中心累计投资到 2035 年将达 31.6 万亿美元。AI 基础设施建设的确定性,已经从"是否建设"变成"多快建设"。

5. 对采购方的启示​

  • 机房规划前置:2027 年起采购 Rubin 级算力,液冷改造(单千瓦改造成本较高)或按全液冷标准新建,必须在预算周期一开始就纳入;
  • 看 PUE 也看水温:45℃ 温水直冷允许更高进水温度,可利用自然冷源压低 PUE——选址时人工冷源依赖度成为新的评估维度;
  • 供应商组合即风险对冲:HBM 与先进封装供应是当前核心瓶颈(详见本站 HBM4 竞速分析),供应链多元化比单点性能更重要。

相关链接​

参考资料​


本文基于 2026 年 9 月初供应链调研、券商研报与英伟达官方披露整理。出货量与功耗数据为产业链预估口径,实际以英伟达及客户正式披露为准。

HBM4 Mass-Production Year One: Samsung Yield Breaks 80%, Three Giants Pass NVIDIA Certification, the Last Bottleneck of AI Compute Supply

· 6 min read
Industry Research Team

If 2025 was the year of HBM3E capacity ramp-up, then 2026 is year one of HBM4 mass production. With NVIDIA Vera Rubin and AMD MI400 — two generations of flagship — both betting on HBM4, this "memory on the AI chip" has for the first time become a strategic commodity that dictates the delivery pace of entire racks. The yield and certification data disclosed densely in August is rewriting the global HBM supply map.


1. Golden Yield Breakthrough: Samsung Jumps from Under 60% to 80% in Six Months​

Per South Korea's Seoul Economic Daily on August 9, Samsung Electronics' HBM4 yield officially crossed the 80% "golden yield" threshold in early August — more than four months ahead of its original year-end target.

TimelineSamsung HBM4 YieldNotes
Feb 2026 (mass production start)Under 60%Line ramp-up period
Early Aug 2026~80%Crosses the mass-production / stable-profit watershed

The semiconductor industry has long held that "80% yield is the golden yield" — it is both a yardstick of foundry competitiveness and the financial break-even point for large-scale commercial supply. The key to this leap was Samsung's breakthrough in Thermal Compression Non-Conductive Film (TC-NCF) bonding, plus the stable base of its underlying 1c DRAM yield, already above 80%. In the same period, Samsung's HBM4E reliability test yield also broke 70%.

Industry assessments suggest SK Hynix's HBM4 yield has likewise entered the 80% range. The gap between the two giants in production quality is being rapidly erased.


2. Supply Map: SK Hynix Holds 60–70% of Rubin Allocation​

At a Seoul event on June 5, Jensen Huang publicly confirmed: Samsung, SK Hynix, and Micron have all passed HBM4 certification for Vera Rubin — the first time three memory makers have simultaneously received public certification for the same platform.

But certification is just the "entry ticket" — allocation share is where the real voice lies:

Vendor2026 Rubin HBM4 Allocation (est.)Notes
SK Hynix60%–70%Based on HBM3/3E-era customer relationships and MR-MUF packaging
Samsung25%–30%Rapid share gains after yield leap
MicronRemainderLimited HBM4 exposure, relatively stable share

Counterpoint Research forecasts the 2026 HBM4 market as SK Hynix 54% / Samsung 28% / Micron 18%. Samsung has set staged catch-up targets: Q3 HBM4 revenue up 3× QoQ, HBM4 exceeding 60% of total HBM revenue in H2, and year-end overall HBM market share approaching 38%.


3. The Real Bottleneck: From Wafers to "Back-End Stacking"​

As front-end yield stabilizes, the rhythm of the AI accelerator supply chain no longer depends on "how many wafers can be made," but on the speed of back-end stacking, bonding, testing, and shipment.

  • Industry analysts rank HBM stacking as the second-most severe bottleneck in the AI chip supply chain, second only to TSMC's CoWoS advanced packaging capacity.
  • HBM accounts for roughly 25% of 2026 DRAM wafer output; each HBM wafer consumes about 3–4× the resources of a standard DRAM wafer (extra TSV and stacking steps), so every wafer redirected pulls 3–4 units of commodity memory off the spot market.
  • Samsung is considering relocating part of its legacy memory back-end lines (Cheonan, Onyang) to Vietnam to free up HBM back-end capacity — a side confirmation that back-end throughput is now the tightest link in the chain.

4. HBM4 Spec Snapshot: Generational Leap in Bandwidth and Efficiency​

SpecHBM4 (12-Hi / 16-Hi)HBM4E
Per-stack capacity36 GB / 48 GB—
Pin rate11.7–13.0 Gbps16 Gbps
Per-stack bandwidthup to 3.3 TB/sup to 3.6 TB/s
Bus width2048-bit—
Energy efficiency+40% vs HBM3E—
Thermal resistance / cooling+10% improvement / +30%—

Samsung HBM4 entered mass production in Feb 2026; its 11.7 Gbps pin rate already exceeds the 8 Gbps industry baseline required for Vera Rubin compatibility; HBM4E samples were first shipped to major customers on May 29.


5. Pricing Power Extends Into 2027: Supply Remains Tight Balance​

TrendForce judges that HBM suppliers' pricing power will run through 2027, because supply remains constrained:

  • 2027 HBM bit shipments are expected to grow 50%–60% YoY, but will still lag demand growth, keeping the market tight;
  • The industry already anticipates significant price increases;
  • For NVIDIA and AMD, a stronger Samsung means more supply options and more comfortable lead times — in a market where memory is the tightest link in AI servers, the mere existence of second and third suppliers is itself a buffer.

For entire racks, HBM cost is already the biggest driver: the Rubin Ultra rack carries an estimated price tag as high as $21 million, with HBM making up a substantial portion.


6. Lessons for China: HBM Export Controls Accelerate Domestic Iteration​

HBM is one of the core fronts of current AI chip controls. As the overseas HBM4 arms race intensifies, domestic HBM technology iteration is being pushed forward in sync — Huawei's Ascend roadmap has explicitly written "drive domestic HBM technology iteration" into its product cadence (the 950 series advances domestic HBM pairing, with the 960/970 series planned for gradual rollout in 2027–2028).

In the short term, HBM4 scarcity will directly transmit to the delivery cadence of Rubin / MI400; in the long term, whoever can lock in stable HBM4 supply holds the valve on 2027 AI compute expansion.

References​


This article is compiled from August 2026 public reports by TrendForce, Seoul Economic Daily, TechTimes, etc. HBM allocation shares and market shares are third-party estimates, not official vendor-confirmed data.

Inference Accelerator Market 2026: 60%–70% of the Accelerator Market, GPU vs ASIC Share Inverts, Five Schools Clash

· 5 min read
Industry Research Team

For the past three years, the entire AI hardware story was "training": who had the most H100s, who could connect a hundred thousand GPUs into a cluster. That race is essentially settled — NVIDIA won. But the next battlefield, "inference," is being fought under completely different rules: the measure is no longer peak FLOPS, but cost-per-token, latency, and power. In 2026, inference chips overtake training in scale for the first time, becoming the main battlefield of AI accelerators.


1. Inference Becomes the Main Battlefield: 80%–90% of Compute Spent on Inference​

Training a large model costs hundreds of millions of dollars — once. But once the model goes live, it must answer billions of queries day after day. A popular consumer model may need tens of thousands of accelerators running 7×24 to keep up with demand. Therefore:

  • Inference accounts for roughly 80%–90% of a model's lifecycle compute;
  • Inference chips will make up about 60%–70% of the ~$400B AI accelerator market in 2026, up from only ~40% in 2023;
  • Inference chip growth (estimated +52.7% YoY) significantly outpaces training chips (+28.4%); the share of inference-side compute demand exceeded training-side for the first time in 2026, reaching 54% (~$1010B).

The economics of inference are straightforward: training cost is amortized to near-zero, while inference cost becomes the entire bill. Every 1% cut in inference cost flows directly to profit — for a company whose inference traffic reaches hyperscale like OpenAI, the half of the bill is a number followed by a string of zeros.


2. Market Size: Structural Growth Inflection Point Has Arrived​

Market2026 SizeGrowthNotes
Global dedicated inference chips$412.7B+38.4%14.2 pct higher growth than training chips
China dedicated inference chips$118.6B (28.7% of global)+44.1%Strongest single market in APAC by growth
Global AI training/inference chips (incl. GPU/NPU)exceeds $1850B+40.2%GPU ~62%

China's domestic substitution is accelerating, with domestic inference chips reaching 34.6% of shipments, up 9.8 pct from 2025.


3. Technology-Axis Share Inverts: GPU Slows, ASIC Soars​

Axis2026 Shipment ShareTrend
GPU52.6%Still leads, but growth slows to 22.7%
ASIC custom chips41.3%Up sharply from 17.8% in 2022
FPGAStableSpecific low-latency scenarios

Thanks to ecosystem maturity, GPU remains the mainstay, but NPU/ASIC already holds a 1.8× advantage over same-generation GPUs in energy efficiency, driving rapid adoption at the edge and on-device. Shipments of inference-optimized ASICs are expected to reach 11.5 million units, with unit cost about 35% lower than GPUs.


4. Five Schools Clash​

SchoolRepresentative ProductsCore StrengthUse Cases
General-purpose GPUNVIDIA Rubin / B200 / H200Mature ecosystem, train+infer unifiedFrontier training + highly interactive inference
LPU (Language Processing Unit)Groq LPUUltra-low latency, deterministic throughputReal-time dialogue, high-concurrency inference
TPU (inference-specific)Google TPU 8i (Zebrafish)288GB HBM, 384MB on-chip SRAM, 19.2 Tb/s ICIGoogle's scaled inference
Custom ASICOpenAI Jalapeno, Microsoft Maia 200, Meta MTIAStrip generality tax for own models, ~50% lower cost/tokenHyperscaler's own workloads
Air-cooled inference cardIntel Crescent Island350W air-cooled, 480GB LPDDR5X, tokens/wattCost-sensitive mid/long-tail inference

OpenAI's Jalapeno, co-developed with Broadcom, aims to cut inference token cost by roughly 50% versus a general-purpose GPU stack — the fifth member to join the "custom inference chip club" (after Google TPU, Amazon Inferentia/Trainium, Microsoft Maia, and Meta MTIA).


5. Core Metric Shifts: cost-per-token and tokens/watt​

The fundamental difference between the inference race and the training race is the low switching cost:

  • Training requires a 100k-GPU cluster + NVLink + CUDA, with extremely high migration cost;
  • Inference is "embarrassingly parallel" at the endpoint level — no million-GPU cluster needed; a node that produces tokens fast and cheaply suffices, and is replaceable per endpoint.

This means NVIDIA's three moats (fastest silicon, NVLink scale-out, CUDA) are no longer absolute on the inference side. When the largest AI buyer (OpenAI) starts treating GPUs as "one of the options," the GPU premium begins to erode — pricing power relies on scarcity, and custom chips attack that scarcity from two directions at once: both reducing merchant-chip demand and giving buyers a credible external negotiation option.


6. Edge and On-Device Explosion: Long-Tail Signal​

Demand shows significant long-tail and fragmentation:

Scenario2026 Demand SizeGrowth
Cloud inference$198.2B (48%)+24.5% (slowing)
Edge inference$126.5B (30.7%)+52.3%
On-device inference$88.0B (21.3%)+68.9%
Autonomous-driving inference$67.3B+58.2%
Industrial QA / robotics inference$42.1B+63.7%

The latency sensitivity and power constraints of inference workloads are reshaping chip architecture design priorities — which also explains why "air-cooled, large-memory" solutions like Crescent Island can find a niche.

References​


This article is compiled from publicly available 2026 market research, brokerage reports, and industry analysis. Market sizes and shares are third-party estimates with inconsistent methodologies and are for reference only.

2026 Global AI Computing Report & Ten Major Computing Industry Trends Released

· 7 min read
Industry Research Team

On May 29, 2026, during the World Intelligence Expo 2026 in Tianjin, the China Intelligent Computing Industry Alliance, National Supercomputing Center in Tianjin, Tianjin Artificial Intelligence Society, Shenzhen Artificial Intelligence Industry Association,ZDNET, and ZDNET ThinkTank jointly released the "2026 Global AI Computing Development Research Report."

The report analyzes the current state and future trends of the global AI computing industry, revealing that the sector has entered a new stage of "intelligence-driven, system-reconstruction."

Core Viewpoints​

1. Computing power becomes a national strategic element​

The global computing industry is entering a new stage of "intelligence-driven, system-reconstruction." With the rise of the "token economy," computing power has become a key foundational element supporting national technological breakthroughs, industrial competition, and strategic positioning.

Computing is evolving from traditional IT support into a strategic bedrock driving scientific innovation and the industrial revolution.

2. AI computing development covers the full chain​

AI computing development must upgrade the full chain of chip, system, and compute cluster, while matching the differentiated computing needs of model training, inference, and data preparation.

  • Training: pre-training of super-large models needs ten-thousand-card-scale compute
  • Inference: super-large models need thousand-card-scale compute
  • Data preparation: needs tens to hundreds of cards

Compute demand at both training and inference ends will keep growing.

3. Domestic AI chip industry's distinctive path​

The domestic AI chip industry follows a route of "autonomy + cluster breakthrough + hardware-software integration + cost-performance advantage," distinct from the foreign pursuit of absolute single-chip compute — better suited to large-scale deployment.

4. Energy challenges for computing centers and solutions​

Computing centers have become the fastest-growing source of global electricity demand. The future requires a diversified energy supply of "short-term wind-solar-storage integration, mid-term nuclear, long-term hydrogen."

Meanwhile, space computing will become a new direction to solve ground-based computing bottlenecks.

5. Compute-network convergence as a core direction​

Future computing will move toward "compute-network convergence," making compute as on-demand as water and electricity — a core part of the national modern infrastructure system.

The computing network has been included in the national "15th Five-Year Plan" major engineering projects, ranked alongside public infrastructure such as hydro power.

Key Data​

Compute performance evolution​

MetricEvolution trend
Chip computefrom TFLOPS scale up to tens of PFLOPS
System formfrom single 8-card machine to thousand-card super-node architecture
Cluster scalefrom thousand-card clusters to hundreds-of-thousands-card clusters
Cluster powerfrom kilowatt to gigawatt scale

Global computing center capacity & energy forecast​

  • Global computing center total capacity: expected to grow from 102GW (2026) to 220GW (2030)

    • AI load capacity from 62GW to 156GW, share rising to 71%
  • U.S. computing center annual electricity: expected to grow from 292TWh to 606TWh, share of national demand rising to 11%

  • China computing center total capacity: ~60GW by 2030, AI load share rising to 48%

  • Global computing center electricity: per IEA base scenario, from ~415TWh (2024) to ~945TWh (2030), ~15% CAGR

Embodied intelligence compute support data​

  • Cloud compute: can generate PB-scale interaction data daily; large-model training cycle shortened from months to weeks
  • Edge compute: tens-to-hundreds of TOPS enables 10–50ms low-latency real-time perception & decision

Industry Trend Analysis​

1. Heterogeneous architecture upgrade​

From traditional CPU+GPU to a new GPU+LPU+CPU+DPU heterogeneous inference architecture.

CPU plays the core role of task scheduling, data pre-processing, serial tasks, and system interconnection in heterogeneous architectures. In 2010, "Tianhe-1A" pioneered large-scale CPU+GPU deployment, leading the global intelligent-computing underlying architecture direction.

2. Clear scale-up / scale-out paths​

  • Scale Up: pursue extreme performance by raising single-node hardware config
  • Scale Out: add nodes for load sharing and high availability

Together they form the core support of computing system capability.

3. Super-node servers become mainstream​

With ultra-high interconnect bandwidth and low communication latency, they shorten model training cycles.

Representative products:

  • Huawei Ascend 384 super-node
  • Sugon scaleX640 super-node
  • Alibaba Cloud Panjiu AL128 super-node
  • Inspur YuanNao SD200
  • Kunlunxin super-node solution

4. Long-context processing optimization​

Through Compressed Sparse Attention (CSA), Heavy-Compressed Attention (HCA) and sliding-window mechanisms, build a "coarse + fine, sparse + dense" long-context modeling system to improve compute efficiency.

Representative application: DeepSeek-V4 attention architecture design.

AI chips​

International vendors:

  • NVIDIA: leads high-end training/inference with Blackwell and Rubin architectures
    • GTC 2026 Taipei (June 1) major releases:
      • Vera Rubin platform in full mass production: NVL72 rack system, agent throughput 10x over Grace Blackwell
      • Vera CPU released: 88-core Olympus in-house Armv9.2, LPDDR5X 1.5TB, 1.2 TB/s, world's first CPU with native FP8
      • RTX Spark AI PC chip: co-developed with MediaTek and Microsoft (codename N1X), Blackwell GPU 1 PFLOP, 128GB unified memory, TSMC 3nm
      • Nemotron 3 Ultra open model: SSM+MoE hybrid, 5x inference speed, 30% lower cost
    • Expanding advantage via CUDA ecosystem
  • Google: deepens vertical HW/SW integration via in-house TPU
  • AWS: Trainium (training) + Inferentia (inference) for cost-effective cloud compute

Domestic vendors: a product matrix represented by Huawei Ascend 910C, Kunlunxin P800, Moore Threads MTT S5000, MetaX XiYun C600.

In 2026 Huawei proposed the "Tao (τ) Law," aiming to systematically reduce the time constant and raise transistor density via logic folding, driving domestic chip evolution.

AI workstations​

  • Form factors: tower, mobile, mini — for different deployment scenarios
  • Compute tiers: entry, professional, enterprise — covering personal dev to enterprise deployment

AI servers​

  • By function: training AI servers and inference AI servers
  • By deployment: cloud AI servers and edge AI servers

With high compute output, high memory bandwidth, and high-speed interconnect, suited to large-scale parallel tasks.

AI computing centers​

  • Trending toward "high AI share, high power density, high electricity consumption"
  • Ultra-large AI computing centers become the construction focus
  • Energy supply moving toward diversified clean sources

Space computing is a new direction, leveraging space's continuous sunlight, extreme cold/vacuum, and interference-free environment to solve ground centers' energy, cooling, and interconnect bottlenecks. Starcloud and Guoxing Weiyu have begun exploration.

1. Scientific research paradigm shift​

The "dry-wet closed loop" research paradigm becomes mainstream, forming a loop between AI-driven "dry experiments" and automated "wet experiments" via data feedback — shifting science from experience-driven to model-driven.

2. Synthetic biology empowerment​

AI's multi-task learning and unknown-space exploration can decode biology's complex "sequence–structure–function" mapping, enabling breakthroughs in protein synthesis, gene editing, and nucleic-acid vaccines. E.g., the AlphaFold series revolutionized protein structure prediction.

3. Embodied intelligence support​

Efficient cloud-edge compute coordination provides full-stack support for embodied intelligence — covering massive data processing, high-fidelity simulation, model training, and edge real-time perception/decision in a closed loop.

Compute-network convergence is the core direction, evolving from "interconnect first, then network" toward a national integrated computing network. The three major telecom operators have begun interconnecting their own compute with dispersed social compute nationwide, promoting ubiquitous compute supply.

The domestic computing ecosystem keeps improving, with deeper government-industry-academia-research coordination. The China Intelligent Computing Industry Alliance, National Supercomputing Center in Tianjin, regional AI societies, industry associations, and service institutions jointly build exchange platforms — driving R&D, standard-setting, technology transfer, and talent cultivation for high-quality domestic computing development.

Conclusions & Outlook​

  1. Computing power is a core element of national strategic competitiveness — major countries are increasing infrastructure investment to seize the AI-era high ground.
  2. The domestic AI chip industry follows a distinctive path — via cluster breakthrough, HW/SW integration, and cost-performance, forming advantage in large-scale deployment.
  3. Computing architecture keeps evolving — heterogeneous computing, super-node servers, and long-context processing are key directions.
  4. Application scenarios keep expanding — from research paradigm shifts to synthetic biology and embodied intelligence, AI compute deeply empowers frontier fields.
  5. Computing infrastructure evolves toward compute-network convergence — future compute will be ubiquitous public infrastructure, on-demand like water and electricity.

References:

  • "2026 Global AI Computing Development Research Report" (China Intelligent Computing Industry Alliance et al.)
  • World Intelligence Expo 2026 (Tianjin, May 29, 2026)