Skip to main content

33 posts tagged with "Industry News"

Major events and market dynamics in the AI compute card industry

View all tags

AMD Advancing AI 2026 Opens Tomorrow: Three CDNA5 MI400 Models, Helios Rack Hits 3 exaFLOPS, OpenAI + Meta Lock 12GW Deal

· 4 min read
AI Hardware Analyst

AMD has confirmed its flagship AI event Advancing AI 2026 will be held July 22-23, 2026 at the Moscone Center in San Francisco, with the keynote on July 23 hosted by Chair and CEO Lisa Su. The event will complete the Instinct MI400 series availability timeline, pricing, and independent benchmark data.

1. Instinct MI400 family: three CDNA 5 accelerators​

AMD fully revealed the MI400 matrix at CES 2026; all three accelerators use CDNA 5 architecture, TSMC 2nm process, differentiated by precision and scenario:

ModelPositioningKey specs
MI455X (flagship)Large-scale train/inference (rack-scale)320B transistors, 12 chiplets, 432 GB HBM4 (12×36GB), 19.6 TB/s, FP4 40 PFLOPS / FP8 20 PFLOPS
MI440X (enterprise)Local enterprise AI (8-card node)Low-precision AI (FP4/FP8/BF16), direct MI300/MI350 replacement, compatible with existing power/cooling
MI430X (HPC/sovereign AI)High-precision scientific computing + AIFull FP32/FP64, already deployed at Oak Ridge Discovery and France's first exascale Alice Recoque

MI455X and MI440X target low-precision AI (FP4/FP8/BF16); MI430X fills traditional HPC high-precision needs — improving energy efficiency and cost-performance by "trimming execution units by precision." All three support UALink (among the first accelerators compatible with the standard) and Infinity Fabric die-to-die interconnect; rack scaling uses Ultra Ethernet.

Lisa Su confirmed on the Q1 2026 earnings call: MI455X samples have been sent to core customers, with demand "exceeding the company's internal expectations for 2027."

2. Helios rack: 3 exaFLOPS per cabinet​

AMD enters the hyperscale market with the Helios rack-scale platform:

MetricHelios rack
Accelerators72 × MI455X
Aggregate HBM431 TB
Total memory bandwidth1.4 PB/s
Per-cabinet computeUp to 3 AI exaFLOPS (Q3 delivery target)
Target customersHyperscale train/inference clusters

Helios uses AMD's in-house Zen 6 EPYC Venice CPU (18 per rack) + Pensando Vulcano 800G NIC, integrated via the open ROCm software stack; AMD also plans a double-width 128-card Helios variant, pushing per-cabinet compute to the 3 AI exaFLOPS ceiling. Further out, the MI500 series (CDNA 6, 2nm, HBM4E) is planned for 2027, with official claims of up to 1000× AI performance vs MI300X.

3. 12GW deal: OpenAI + Meta dual endorsement​

AMD holds two historic-scale compute agreements totaling about 12 GW, with lifetime potential revenue possibly reaching $100B:

CustomerScaleFirst deploymentStructure
OpenAI6 GW (multi-gen products)First 1 GW, H2 2026 on MI450"compute-for-upside": up to 160M warrants, vesting by milestone and stock-price targets
Meta6 GWCustom MI450 chips, from H2 2026Deployed in next-gen data centers

Financial expectations​

Metric2026 forecast
MI400 series revenue~$7.2B (about 25% of data-center sales)
Data-center GPU revenue~$15B (up +114% YoY)
Total data-center revenuePossibly $28.7B (up +73% YoY)

⚠️ Execution risk: AMD has flagged that MI450's Q3 mass production will weigh on gross margin (new products below company average); advanced process and advanced packaging (TSMC CoWoS) capacity remain the main constraint.

Industry interpretation​

  1. CUDA moat being pried open: when companies building the world's largest training clusters — Meta, OpenAI — bet on AMD silicon, AMD's long-standing 5-7% GPU share ceiling is being broken.
  2. Memory advantage as differentiation: 432 GB HBM4 / 19.6 TB/s vs NVIDIA Rubin's 288 GB offers capacity advantage, critical for large-model inference (KV Cache-constrained scenarios).
  3. Tight benchmarking pace: MI450 and NVIDIA Vera Rubin both ramp in H2 2026, with the two giants competing head-on over HBM4 supply and CoWoS capacity.

References​


This article was written on the eve of Advancing AI 2026 (July 22-23, opening tomorrow); the keynote is July 23 hosted by Lisa Su, where MI400's final availability, pricing, and independent benchmarks will be revealed — we will update in sync.

NVIDIA Vera Rubin Officially Ships: First VR200 NVL72 Delivered, Samsung HBM4 Mass Production, Rubin Ultra Cabinet Sky-High Price

· 5 min read
Industry Research Team

July 2026, NVIDIA's next-gen AI compute platform Vera Rubin officially began its first shipments, succeeding the Blackwell architecture, with large-scale mass production planned for H2 2026. First customers include Microsoft, Google, Amazon, Meta, Oracle, and other large cloud providers.

1. World's First VR200 NVL72 Delivered (Milestone)​

CoreWeave jointly with Dell announced that the world's first NVIDIA Vera Rubin VR200 NVL72 cabinet has been officially delivered and passed the L11 full-cabinet hardware diagnostics on the first try. This marks Rubin's move from roadmap to physical product, with no major bottlenecks in core supply-chain links (HBM4, advanced packaging, liquid cooling, ultra-high-power power supply).

VR200 NVL72 Core Configuration​

MetricVera Rubin VR200 NVL72
Cabinet codenameOberon
GPU72 Rubin GPUs
CPU36 Vera CPUs
Per-GPU memory288 GB HBM4
Per-CPU memory1.5 TB LPDDR5X
Total cabinet HBM420.7 TB (20,736 GB)
Total cabinet LPDDR5X54 TB
InterconnectNVLink 6 full mesh
Inference performance~3.6 exaFLOPS class
CoolingLiquid cooling
Generational improvement~3.5× per-GPU compute, ~2.8× memory bandwidth (vs Blackwell)

Vera CPU integrates 88 custom Olympus ARM cores, with 1.8 TB/s interconnect to the GPU, usable as a GPU memory expansion pool. NVIDIA completed its first Vera CPU deliveries to Anthropic, OpenAI, xAI, and Oracle Cloud in May.

2. Samsung HBM4 Mass Production: Key Bottleneck Eases​

July 8, 2026, Samsung Electronics officially started HBM4 mass production for the Vera Rubin platform, with reported HBM4 mass-production yield reaching 70% (above the initial 60-65% expectation). Confirmation of this key supply-chain link clears obstacles for Rubin's large-scale deployment.

HBM Supply Landscape (2026 Q1)Share
SK hynix45%
Samsung40%
Micron15%

HBM4 uses 8-layer stacking (12-layer design planned for 2028), priced at about 2.8× HBM3e. TrendForce predicts HBM supply will grow 65% annually, with HBM4 reaching 35% of total output by 2027 Q4.

3. Rubin Ultra Sky-High Price: HBM Cost Dominates​

Per BofA Global Research estimates, the Rubin generation will push single-server cost to a new high:

Cost ItemRubin VR200 (Oberon)Comparison
Cabinet HBM4 usage20,736 GB—
HBM4 unit price~$18.40 / GBBlackwell (HBM3e) ~$11.26 / GB
HBM4 cost alone~$382KExcluding LPDDR5X
Rubin Ultra cabinet estimated price~$21MITHome / BofA estimate

4. Rubin Ultra Design Change: Original 4-die Cancelled (per SemiAnalysis)​

Semiconductor research firm SemiAnalysis (2026-06-30) disclosed that the original 4-die Rubin Ultra GPU unveiled at GTC 2026 has been cancelled; the version actually shipping in 2027 is roughly halved in scale and performance:

  • Reason for cancellation: The original integrated 4 compute dies + 16 HBM4E in a single CoWoS-L package; the substrate warped under the 4-die config, causing compute-die-to-substrate contact failure and yield collapse; the alternative CoPoS won't reach mass production until after late 2028, missing the 2027 node.
  • New approach: Changed to dual-die (same construction as standard Rubin) + HBM4E, ~384 GB HBM4E per GPU (higher than standard Rubin's 288 GB), but total compute and bandwidth only half the original; to approach the original's aggregate compute, NVIDIA plans to assemble "2+2" board-level configs within the Kyber rack to reach four-die equivalent scale.
  • Kyber rack delay: The companion Kyber NVL144 rack is delayed 12+ months to 2028 due to midplane PCB manufacturing difficulties; the 800V DC power scheme is likewise delayed to 2028.

⚠️ Note: NVIDIA has not commented officially on the above design change; some on X argue "the chip count hasn't changed, it's old news reheated." This section is compiled from SemiAnalysis public reports, subject to final NVIDIA disclosure. We have marked "specs pending official confirmation" on the Rubin Ultra preview card.

Industry Interpretation​

  1. "Never doubt" moment realized: Rubin's first delivery passed L11 on the first try, dispelling market doubts about "Rubin delay," locking in H2 2026 AI compute supply certainty ahead of time.
  2. Designed for Agentic AI: Rubin targets agentic workflows and ultra-long-context inference, further lowering the training/inference cost curve for trillion-parameter models.
  3. HBM is the full-chain winner: 20.7 TB HBM4 per cabinet is enormous usage; SK hynix, Samsung, Micron, advanced packaging (CoWoS-L), liquid cooling, and power retrofitting all benefit across the chain, while also becoming the biggest cost and capacity constraint.

References​


This article continuously tracks Vera Rubin mass-production ramp and HBM4 supply-chain dynamics.

June 2026 AI Chip Major Events Roundup: Ascend 910C Trains Trillion-Parameter Model, OpenAI Custom Chip, RTX Spark Launch

· 6 min read
Industry Research Team

June 2026 saw multiple milestone events in the AI chip field, marking acceleration of two major trends: "domestic substitution" and "de-NVIDIA-ization."

1. Huawei Ascend 910C Completes 1.6-Trillion-Parameter DeepSeek V4 Pro Training (2026-06-05)​

Event Overview​

June 5, 2026, Shenzhen Hetao College, together with Harbin Institute of Technology (Shenzhen), Shenzhen Big Data Research Institute, Huawei, and other teams, relied on an Ascend 910C domestic AI compute cluster to successfully complete full-parameter post-training of the 1.6-trillion-parameter DeepSeek V4 Pro large model.

Technical Significance​

MetricValue
Model parameters1.6 trillion
Training chipAscend 910C cluster
Training typeFull Parameter Post-Training
SignificanceFirst time domestic AI chips complete trillion-parameter-level model training

Industry Impact​

  1. Breaks technology blockade: Proves domestic AI chips can train trillion-parameter models
  2. Accelerates "farewell to NVIDIA": DeepSeek fully switches to Huawei Ascend, reducing dependence on H100
  3. Domestic substitution inflection point: From "inference substitution" to "training substitution"

2. OpenAI Launches First Custom AI Inference Chip Jalapeño (2026-06-24)​

Event Overview​

June 24, 2026, OpenAI and Broadcom jointly launched the first custom AI inference chip Jalapeño, with a design cycle of only 9 months (industry average 18 months), using TSMC 3nm process.

Key Metrics​

MetricJalapeñoComparison (Blackwell)
ProcessTSMC 3nmTSMC 4nm
ArchitectureSystolic ArrayBlackwell GPU
Design cycle9 months~18 months
Inference cost-50%Baseline
AI-assisted design✅ First❌ No
DeploymentEnd of 2026Shipped

Strategic Significance​

  1. First AI chip with AI-assisted design: OpenAI used models like GPT-5.3-Codex-Spark to assist architecture exploration
  2. Accelerates "de-NVIDIA-ization": Tech giants (Google, Amazon, Microsoft, Meta, OpenAI) collectively develop custom chips
  3. Inference cost revolution: For OpenAI processing hundreds of millions of API calls daily, a 50% cost reduction is significant

3. NVIDIA Launches RTX Spark AI PC Superchip at Computex 2026 (2026-06-01)​

Event Overview​

June 1, 2026, NVIDIA CEO Jensen Huang launched the RTX Spark AI PC superchip at Computex 2026 / GTC Taipei, in collaboration with MediaTek, using an Arm CPU + Blackwell GPU unified-memory architecture.

Key Metrics​

MetricRTX Spark
CPUUp to 20-core Arm (with MediaTek)
GPU6,144 CUDA cores (Blackwell)
Unified memory128GB LPDDR5X (shared CPU+GPU)
Memory bandwidth300 GB/s
AI compute~1 PFLOPS (est.)
Model capacityCan run 120B-parameter models
ContextUp to 1 million tokens
TDP~100W (est.)
AvailabilityFall 2026

Industry Impact​

  1. NVIDIA enters PC chip market: Challenges Intel's dominance in personal computers
  2. New AI PC standard: Run 120B-parameter models locally, 1M-token context
  3. Windows transforms into AI Agent platform: Deep collaboration with Microsoft OpenShell framework

4. MIIT Publishes "2026 AI Chip Industry Development White Paper" (2026-06-09)​

Event Overview​

June 9, 2026, China's Ministry of Industry and Information Technology published the "2026 AI Chip Industry Development White Paper," predicting the domestic AI chip market will exceed 200 billion RMB in 2026.

Key Predictions​

Metric2026 Prediction
Market sizeExceed 200 billion RMB
Domestic chip share>50% (41% in 2025)
Edge inference chipsSignificant progress
Shipment growthMore than double (vs 2025)

Industry Significance​

  1. Domestic AI chip capitalization accelerates: Cambricon, Enflame, Moore Threads, etc. accelerate IPOs
  2. Edge inference becomes the breakthrough: Easier to achieve domestic substitution than training chips
  3. Policy dividend continues: Domestic substitution upgraded from "market behavior" to "national strategy"

5. ByteDance in Talks to Procure 50K Iluvatar Inference Chips (2026-06-17)​

Event Overview​

June 17, 2026, Reuters reported that ByteDance is in talks with Shanghai AI chip firm Iluvatar to procure at least 50,000 AI chips, mainly for inference tasks.

Deal Details​

ItemContent
BuyerByteDance
SupplierIluvatar
Chip modelZhiKai series (inference GPU)
QuantityAt least 50,000
UseInference workloads
Training chipTianTai series

Industry Significance​

  1. Domestic GPU top player "adds a member": Iluvatar enters a top internet company's supply chain for the first time
  2. ByteDance 2026 capex raised over 200B RMB: Mainly for AI compute and datacenters
  3. "Domestic substitution" extends from government/SOEs to private tech giants

Trend 1: "Domestic Substitution" Moves from Inference to Training​

  • Ascend 910C completes 1.6-trillion-parameter model training → Proves domestic chips have training capability
  • DeepSeek fully switches to Ascend → Leading AI companies first to "farewell to NVIDIA"
  • ByteDance procures Iluvatar → Private tech giants follow

Trend 2: "De-NVIDIA-ization" from Slogan to Action​

  • OpenAI Jalapeño → First custom chip, inference cost -50%
  • Google TPU, Amazon Trainium, Microsoft Maia → Continuous iteration
  • Meta MTIA, Apple M5 Ultra → Increased investment

Trend 3: AI PC and Edge Inference Become New Battlefield​

  • NVIDIA RTX Spark → New AI PC standard, launches Fall 2026
  • Edge inference chip localization accelerates → Key mention in MIIT white paper
  • "Local trillion-parameter model execution" → New consumer market selling point

Looking Ahead (2026 H2)​

  1. Ascend 950DT full scale-up (2026 Q4) → Huawei's latest-gen training chip
  2. NVIDIA Rubin R200 shipment (2026 H2) → Next-gen flagship
  3. AMD MI400 Helios rack (2026 H2) → Targets NVIDIA GB200
  4. OpenAI Jalapeño deployment (end of 2026) → Gigawatt-scale datacenters
  5. Domestic AI chip shipments more than double → CITIC Securities prediction

References​


This article is continuously updated. Please provide the latest developments.

Computex 2026 Wrap-Up: AI PC Chip War Begins, NVIDIA RTX Spark Arrives Fall 2026

· 3 min read
Industry Research Team

June 6, 2026 — COMPUTEX 2026 concluded yesterday in Taipei. Under the theme "AI Together," this year's event set records with 1,500+ exhibitors and 6,000 booths. The head-to-head battle between NVIDIA, Intel, and AMD in the AI PC space was the defining story of the show.

1. NVIDIA RTX Spark: June Launch at $1,399​

Less than a week after its COMPUTEX debut, the NVIDIA-MediaTek RTX Spark Superchip confirmed its commercial timeline:

DetailInfo
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Starting PriceNot yet announced (analysts estimate $3,000-4,000)
Core SpecsArm CPU (up to 20 cores) + Blackwell GPU (6,144 CUDA cores)
Unified Memory128 GB LPDDR5X (300 GB/s)
Model CapacityRuns 120B parameter models, up to 1M token context

Market Reaction: AMD, Intel, and Qualcomm shares fell following the announcement. Analysts believe RTX Spark will reshape the market across three fronts — Windows AI PCs, creator workstations, and edge inference nodes.


2. Intel 18A in Full Production: Clearwater Forest + Crescent Island​

Intel CEO Lip-Bu Tan delivered his first COMPUTEX keynote with two key updates:

Clearwater Forest (Xeon 6+)​

  • 288 cores, Darkmont architecture
  • First Intel 18A process node data center CPU
  • Foveros Direct 3D packaging
  • Now in full production

Crescent Island AI GPU​

  • 480 GB LPDDR5x memory
  • 350 W air-cooled PCIe form factor
  • Native FP4 support, targeting agentic inference
  • Shipping H2 2026

"As AI moves into the agentic era, the CPU returns to the center of modern AI infrastructure." — Lip-Bu Tan


3. AMD Ryzen AI 400 Series Now Shipping​

AMD showcased the Ryzen AI 400 series (Zen 5 + Zen 5C hybrid + XDNA2 NPU) at COMPUTEX:

  • NPU performance: 60 TOPS, the highest in x86
  • 7 consumer SKUs + commercial PRO series
  • Multiple OEM models already available or launching soon
  • Advancing AI 2026 summit set for July in San Francisco

4. Chinese Domestic Chips Gaining Momentum​

VendorProductStatus
HuaweiAscend 950PR/950DTIn production, self-developed HBM
CambriconMLU6902 PFLOPS FP8, shipping
Moore ThreadsMTT S50001,000 TFLOPS, specs public

5. The AI PC Era: Three-Way Roadmap Comparison​

DimensionNVIDIA RTX SparkIntel Clearwater Forest + Crescent IslandAMD Ryzen AI 400
CPU Cores20-core Grace (Arm)288-core Darkmont (x86)Up to 12-core Zen5+5C
GPU/NPUBlackwell GPUCrescent Island (discrete GPU)XDNA2 NPU (60 TOPS)
AI Compute1 PFLOPSTBD60 TOPS NPU
TargetPersonal AI agentsDual-track: DC + AI PCCopilot+ PC
ProcessTSMC 4NPIntel 18ATSMC 4nm
AvailabilityJune 2026H2 2026Shipping now

This Week in AI Compute (6/1 – 6/6)​

DateEvent
Jun 1NVIDIA GTC Taipei: RTX Spark, Vera Rubin production, DGX Station for Windows
Jun 1Intel unveils Crescent Island, Clearwater Forest
Jun 2COMPUTEX 2026 opens: "AI Together"
Jun 5COMPUTEX closes: 1,500+ exhibitors, record scale
Jun 6RTX Spark confirmed June launch at $1,399

Sources: COMPUTEX Daily, Tencent News, Phoenix Technology, Xueqiu, The Silicon Review.

Computex 2026 AI Compute Card Major Events: DGX Station for Windows, Intel Crescent Island, and More Major Launches

· 4 min read
Industry Research Team

June 1-5, 2026, Taipei — Computex 2026 (Taipei International Information Technology Show) wrapped up successfully this week. With the theme "AI Together," industry giants including NVIDIA, Intel, AMD, and Qualcomm unveiled numerous AI compute products in rapid succession. Below, MirrorFrog brings you a roundup of the most noteworthy developments in the compute card space this week.

① NVIDIA DGX Station for Windows: A Desktop AI Supercomputer​

NVIDIA officially launched the DGX Station for Windows during its Computex 2026 keynote, calling it "the world's most powerful desktop AI supercomputer."

Core Specifications​

ItemSpecification
ChipGB300 Grace Blackwell Ultra Desktop Superchip
GPU Memory252 GB HBM3e (7.1 TB/s)
CPU Memory496 GB LPDDR5X (396 GB/s)
Unified Memory748 GB (NVLink-C2C interconnect)
FP4 Compute20 PFLOPS (sparse)
FP8 Compute10 PFLOPS (sparse)
NetworkConnectX-8 SuperNIC, up to 800 Gb/s
Model CapacityCan run 1 trillion parameter models
System Power1,600 W
Operating SystemMicrosoft Windows
ShippingQ4 2026

Significance: DGX Station compresses AI compute power (20 PFLOPS FP4) that previously required datacenter-class clusters into a single desktop workstation. 748GB of unified memory means developers can run models with hundreds of billions or even trillions of parameters locally, without cloud dependency.


② Intel Crescent Island: Inference-Specialized AI GPU​

At Computex, Intel disclosed detailed specifications for its next-generation datacenter AI inference GPU, Crescent Island.

ItemSpecification
MemoryUp to 480 GB LPDDR5x
Power350 W (PCIe form factor)
Precision SupportFP4/MXFP4 → FP64 (full precision coverage)
TargetAI inference workloads (Agentic Inference)
PositioningBetter price-performance than HBM solutions
ShippingH2 2026

Significance: Crescent Island represents Intel's key strategic move in the AI inference market. 480GB of massive LPDDR5x memory (non-HBM) means significantly lower cost compared to NVIDIA H200/B200 and other competing products, targeting enterprise inference deployment scenarios.


③ Intel Xeon 6+ (Clearwater Forest): First Intel 18A Datacenter CPU​

Intel also unveiled the new Xeon 6+ processor, codenamed Clearwater Forest, its first datacenter CPU built on the 18A process:

  • 288 Darkmont architecture cores
  • L2 288MB + L3 576MB cache
  • 12-channel DDR5-8000 memory
  • Foveros Direct 3D advanced packaging
  • AI Agent Era: CPU returns to the center of infrastructure

④ NVIDIA RTX Spark Ecosystem Takes Shape​

This week, the RTX Spark super chip developed in collaboration between NVIDIA and MediaTek continued to generate buzz. Multiple OEMs showcased RTX Spark-based laptop and compact desktop prototypes:

  • ASUS, Dell, HP, Lenovo, Microsoft Surface, MSI all confirmed as launch partners
  • Equipped with 20-core Grace CPU + Blackwell GPU (6144 CUDA cores)
  • AI compute 1 PFLOPS
  • Retail availability Fall 2026

⑤ Intel × Foxconn AI Infrastructure Partnership​

Intel and Foxconn announced a joint AI infrastructure initiative, covering the complete chain from chip → server → rack-scale system, targeting the datacenter market opportunity driven by surging AI inference demand.


⑥ Domestic AI Chip Developments​

According to the IDC 2025 annual report, total AI accelerator card shipments in China reached approximately 4 million units, with domestic vendors shipping approximately 1.65 million units, capturing a market share exceeding 41%. Huawei's Ascend 950 series has entered mass production and delivery, while Cambricon's MLU690 has begun shipping to internet customers.


This Week's Compute Roundup​

VendorProductHighlightTimeline
NVIDIADGX Station for Windows20 PFLOPS, 748GB unified memoryQ4 2026
NVIDIARTX Spark1 PFLOPS AI PC chipFall 2026
IntelCrescent Island GPU480GB LPDDR5x, 350WH2 2026
IntelXeon 6+ (Clearwater Forest)288 cores, Intel 18AH2 2026
Intel + FoxconnAI infrastructure partnershipChip→rack full chainStrategic partnership
HuaweiAscend 950PR/DT1 PFLOPS FP8, self-developed HBMIn mass production
CambriconMLU6902 PFLOPS FP8, 192GB HBM3EShipping

Sources: NVIDIA GTC Taipei 2026 / Computex 2026 official announcements, Intel press releases, ifeng Tech, IT Home.

Huawei Ascend 950 Mass Production and the Full Picture of China's AI Chip Ecosystem

· 4 min read
Industry Research Team

June 2026 — Huawei's Ascend 950 series (950PR / 950DT) has entered formal mass production and delivery, a landmark event for China's AI chip industry in 2026. Meanwhile, Cambricon's MLU690 has begun shipping and Moore Threads has announced MTT S5000 specifications, formally establishing China's tri-polar AI chip landscape.

Ascend 950 Series: A Historic Breakthrough with Self-Developed HBM​

Huawei HiSilicon's Ascend 950 series is the fourth-generation Ascend AI chip, first revealed at Huawei Connect 2025 in September and entering mass production in Q1 2026.

950PR (Prefill Inference Specialized)​

ItemSpecification
ArchitectureDa Vinci v5 (SIMD + SIMT dual-model)
ProcessN+2 (SMIC domestic)
HBMHiBL 1.0 (Huawei self-developed) , 128 GB
FP8 Compute1 PFLOPS (HiF8 format)
TDP~400 W
TargetInference Prefill (video recommendation, real-time interaction)

950DT (Decode + Training Specialized)​

ItemSpecification
ArchitectureDa Vinci v5 (SIMD + SIMT dual-model)
ProcessN+2 (SMIC domestic)
HBMHiZQ 2.0 (Huawei self-developed) , 144 GB, 4 TB/s
FP8 Compute1 PFLOPS (HiF8 format)
TDP~500 W
TargetInference Decode + Model Training

Historical Significance​

Self-developed HBM (HiBL 1.0 / HiZQ 2.0) represents the most important technical breakthrough of Huawei Ascend 950 — this is the first time a Chinese enterprise has achieved self-developed mass production of HBM memory, completely eliminating dependence on SK Hynix / Samsung HBM supply. Combined with the domestic N+2 process, Ascend 950 has achieved full-chain domestic production from HBM → Compute Die → Packaging → System.

Cambricon MLU690: China's Only Native FP8 Support​

Cambricon's seventh-generation AI chip MLU 690 (Siyuan 690) began volume production and shipping in H1 2026. This is the first domestic AI chip with native FP8 precision support.

ItemMLU 690
Process5nm (TSMC / SMIC)
FP8 dense2 PFLOPS
HBM192GB HBM3E, 5 TB/s
TDP~500 W
Unit Price (OAM)~$8,000-12,000

MLU 690's FP8 compute power (2 PFLOPS dense) is on paper comparable to NVIDIA Blackwell (B200 FP8 4.5 PFLOPS sparse). Leveraging its financing advantage as a STAR Market listed company, Cambricon targets 2026 revenue of ¥15-20B (2025: ¥7.2B).

Moore Threads MTT S5000: From Graphics to Training-Inference Unified​

Moore Threads publicly disclosed detailed specifications of the MTT S5000 in February 2026, featuring the fourth-generation MUSA "Pinghu" architecture, single-card AI compute of 1,000 TFLOPS, 80GB GDDR6X memory, 1.6 TB/s bandwidth.

Moore Threads pursues a full-function GPU path (graphics rendering + AI compute + general-purpose compute), closest to NVIDIA's strategy. The founding team comes from former NVIDIA China, and the MUSIFY toolchain helps auto-migrate CUDA code to the MUSA platform, lowering ecosystem migration costs.

China's Tri-Polar AI Chip Landscape​

DimensionHuawei AscendCambriconMoore Threads
Core ArchitectureDa Vinci v5MLUv07MUSA 4th Gen
ProcessN+2 domestic5nm6nm
FP8 Compute~1 PFLOPS2 PFLOPS0.5 PFLOPS (estimated)
HBM Self-Sufficiency✅ Self-developed HiBL/HiZQ❌ Purchased❌ Purchased
EcosystemCANN + MindSporeNeuWare + MindSporeMUSA + MUSIFY
AdvantageFull-chain domesticHighest FP8 computeFull-function + CUDA migration
2025 Revenue(Huawei internal)¥7.2B¥2.2B

Global Market Comparison (Q2 2026 Update)​

TierVendorFlagship ChipFP8/PFLOPSHBMMass Production
Tier 1NVIDIARubin R20025 PF (sparse)288GB HBM42026 H2
Tier 2AMDMI40020 PF (dense)432GB HBM42026
HuaweiAscend 950DT1 PF (dense)144GB self-developed HBM2026 Q1
CambriconMLU6902 PF (dense)192GB HBM3E2026 H1
AWSTrainium 35.7 PF (dense)144GB HBM2025 Q4 GA
Tier 3IntelGaudi 31.8 PF128GB HBM2eIn production
GoogleTPU v74.6 PF(TFLOPS)192GB HBM2025
Moore ThreadsMTT S50001 PF80GB GDDR6X2025 Q1

Note: NVIDIA uses sparse compute as standard, while AMD / Huawei / Cambricon use dense — not directly comparable.

Outlook for H2 2026​

  • NVIDIA Rubin R200: Official shipment in H2 2026, 288GB HBM4, 6-chip CoWoS-L packaging
  • Huawei Ascend 960: Roadmap H2 2027, expected FP8 compute doubled to 2 PFLOPS
  • Cambricon MLU790: Expected 2027, 3nm, 384GB HBM4, 2.5 PFLOPS
  • Moore Threads: Next-gen GPU expected with HBM3, 2× MTT S5000 compute

By 2026, China's AI chip industry has formed a complete product matrix from Training (Cambricon MLU690 / Ascend 950DT) → Inference (Ascend 950PR / Moore Threads S5000) → Systems (CloudMatrix / Distributed Clusters).


This article is based on public information from Huawei Connect 2025 (2025-09-18), industry analysis reports from April 2026, and the latest market data as of June 2026.

NVIDIA Launches RTX Spark: AI Compute Enters the Personal Computer Era

· 3 min read
Industry Research Team

June 1, 2026, Taipei — During the Computex 2026 opening keynote, NVIDIA CEO Jensen Huang officially unveiled the RTX Spark super chip, marking NVIDIA's formal entry into the personal computer processor market dominated by Intel, AMD, Qualcomm, and Apple.

RTX Spark: The "Heart" of the Personal AI Computer​

RTX Spark was developed in collaboration between NVIDIA and MediaTek, featuring a heterogeneous package with a 20-core Grace CPU + Blackwell RTX GPU, equipped with 6144 CUDA cores. AI compute reaches 1 PFLOPS (one quadrillion floating-point operations per second), meaning personal computers now possess computing power comparable to a datacenter-class H100 GPU for the first time.

SpecificationRTX Spark
CPU20-core Grace (MediaTek collaboration, Arm architecture)
GPUBlackwell RTX (6144 CUDA cores)
AI Compute1 PFLOPS
TargetPersonal AI Agent, local LLM inference
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Form FactorLaptop SoC + compact desktop workstation

Jensen Huang's "Full-Stack AI" Strategy​

The launch of RTX Spark is a key step in NVIDIA's "full-stack AI" strategy. Jensen Huang stated during the keynote: "AI should not only run in the cloud. Everyone's computer should have the ability to run AI agents."

RTX Spark transforms NVIDIA from a datacenter GPU monopolist into a full competitor in the personal computing market. Following the announcement, shares of AMD, Intel, and Qualcomm fell accordingly.

Market Impact​

  • Intel: Personal computer AI processor business faces direct threat
  • AMD: Ryzen AI series must compete at the same level
  • Qualcomm: Snapdragon X Elite's Copilot+ PC positioning challenged
  • Apple: M-series chips are no longer the only high-performance AI PC option

Vera Rubin Platform Enters Full Mass Production​

During the same keynote, Jensen Huang also announced that the NVIDIA Vera Rubin platform has entered full mass production. Rubin R200 features a 6-chip CoWoS-L package (1× Vera CPU + 2× Rubin GPU die + I/O/HBM die), equipped with 288GB HBM4, 22 TB/s bandwidth, and 50 PFLOPS FP4 compute (sparse).

The Rubin NVL72 rack (72 Rubin GPUs + 36 Vera CPUs) will begin shipping in H2 2026.

Other Highlights from Computex 2026​

  • AMD: Showcased the MI350 series (192GB HBM3e, 5 PFLOPS FP8 dense), officially launching in June
  • Intel: Jaguar Shores publicly unveiled for the first time
  • Qualcomm: AI 200 / 300 series inference card roadmap updated
  • Domestic AI Chip Zone: Huawei, Cambricon, Moore Threads, and others showcased their latest products

Industry Significance​

The launch of RTX Spark means AI compute is no longer confined to datacenters. Individual developers, designers, and researchers will be able to run large model tasks locally that previously required cloud GPUs, potentially redefining the market landscape for personal AI computing.

The mass production of Vera Rubin further consolidates NVIDIA's absolute leadership in datacenter AI training. Together, both product lines form NVIDIA's full-stack AI computing landscape of "cloud training + personal inference."


This report is based on official NVIDIA announcements from Computex 2026 / GTC Taipei on June 1, 2026.

AI Cluster Power Crisis: 1MW Racks, Nuclear Plants, SMRs, and Green AI

· 8 min read
Industry Research Team

In 2026, AI compute growth has hit a hard constraint — electric power. With NVIDIA Rubin NVL576 single-rack power consumption at 1 MW, the xAI Colossus cluster at 200 MW, and OpenAI's planned Stargate campus at 5 GW, power supply is becoming the biggest bottleneck for AI development. This article provides an in-depth analysis of this "power crisis" and the solutions.

AI Chip Startup Survival Report: Tenstorrent / SambaNova / Graphcore in 2026

· 8 min read
Industry Research Team

2026 AI chip market enters a "winner takes all" phase. NVIDIA holds 90%+ market share, AMD struggles at 10%, and Google/AWS/Huawei/Cerebras each occupy niche segments. But a group of AI chip startups are fighting to survive in the cracks — this article analyzes the 2026 status and future of Tenstorrent, SambaNova, Graphcore, Cambricon, Moore Threads, Biren, and Iluvatar.

Intel Cancels Falcon Shores, Pivots to Jaguar Shores: From Single-Chip Competition to Rack-Scale Systems

· 5 min read
Industry Research Team

May 14, 2026, Intel disclosed in its Q1 earnings report that it has formally cancelled the Falcon Shores single-chip GPU project and confirmed a new rack-scale AI system project named Jaguar Shores to launch in 2027-2028. This is a major strategic adjustment in Intel's AI roadmap. This article provides an in-depth analysis of the reasons and future implications.