Skip to main content

7 posts tagged with "AI Chip"

AI chip industry dynamics and trends

View all tags

HBM4 从送样竞速进入量产爬坡竞速:SK 海力士率先交付 Rubin、美光产能翻倍、三星直逼王座

· 6 min read
Industry Research Team

9 月第一周,存储行业传来三组关键信号:SK 海力士面向英伟达 Vera Rubin 平台量产交付全球首款 12 层 HBM4;三星 Q2 HBM 份额飙升至 33%、直逼王座;美光宣布年底 HBM 月产能翻倍至约 10 万片晶圆。 HBM4 的竞争,已经从"谁能造出来"变成"谁能最快量产爬坡、谁绑定的客户最深"。


1. 三巨头战报:领跑者、追赶者与搅局者​

维度SK 海力士三星电子美光
HBM4 关键节点2026 年中向 Vera Rubin 量产交付 12 层 HBM4(全球首款完整质量认证)2026 年 2 月全球首家大规模量产出货2026 Q2 量产 12 层 HBM4,批量出货
DRAM 核心裸片1b(第五代 10nm 级)1c(第六代)—
Base Die 代工台积电 12nm自家 4nm拟转台积电(1γ 工艺代)
2026 HBM4 份额预测~54%~28%~18%
Q2 HBM 营收份额50%(环比 -14pct)33%(环比 +12pct)18%
HBM 产能~15–20 万片/月~15–20 万片/月年底冲刺 10 万片/月(现为 4–5 万)

三个关键读数:

  • 三星的反扑是真的:从去年 Q2 份额低至 15%(落后美光 6 个百分点)到今年 Q2 的 33%,三星用一年时间把与 SK 海力士的差距缩小到 17 个百分点。三星预计 HBM4 将占其下半年 HBM 收入的 60% 以上,且现有 HBM4 产能已被客户预订售罄,2026 年 HBM 总产能计划提升约 50%;
  • 美光的追赶也是真的:CEO 梅赫罗特拉透露 12 层 HBM4 爬坡速度约为 HBM3E 的两倍、累计收入已超 10 亿美元;年底若如期达成 10 万片月产能,与两强的差距将缩小一半;
  • SK 海力士的护城河依然深:与英伟达签署多年期 HBM/DRAM 供应协议、率先完成 2027 年 HBM4 价格谈判,在英伟达 HBM4 采购中占比约 60%(乐观情形 70%)。

2. 变局核心:封装比 DRAM 制程更值钱​

HBM4 与 HBM3E 时代最大的不同,是胜负手从存储颗粒转向了封装与 Base Die:

  • I/O 通道数从 1024 条翻倍至 2048 条,单颗 12 层封装吞吐超 2TB/s,能效较前代提升逾 40%;
  • Base Die 首次引入晶圆代工厂先进逻辑工艺——SK 海力士用台积电 12nm,三星用自家 4nm,美光计划在 1γ(首次导入 EUV)世代转由台积电代工;
  • 存储原厂与代工厂的协作深度,第一次成为决定出货速度的关键变量——这正是 Hot Chips 2026 上"三星 HBM Base Die 逻辑工艺化"议题的产业注脚。

SEMI 数据显示,2026 年全球 HBM 市场规模预计增长 58% 至 546 亿美元,约占 DRAM 市场四成。行业预计 HBM4 销售占比将在 2026 年 Q4 正式超越 HBM3E,成为市场主流。

3. 客户端信号:Rubin 加码,Rubin CPX 重生​

需求端同样在 9 月出现重要变化:

  • Vera Rubin 全面量产(详见本站专文):每颗 Rubin GPU 搭载 288GB HBM4、带宽 22TB/s,NVL72 整柜 75TB 快速内存——HBM4 供给直接决定 Rubin 出货上限,这也是英伟达把"先进晶圆与 HBM 供应"列为当前第一约束的原因;
  • Rubin CPX 项目重启:据 CFM 闪存市场 9 月 2 日简讯,英伟达重启此前搁置的 Rubin CPX 机柜项目,内存规格由 128GB GDDR7 改为 168GB HBM4,机架改为独立 MGX ETL 设计,客户可选 64/128/192/256 颗 GPU 配置——推理专用卡也全面转向 HBM4,进一步放大了 HBM4 需求;
  • 供给协议长期化:英伟达已与 SK 海力士、美光签署多年期 HBM 及 DRAM 供应协议,锁供给、锁价格、锁产能成为巨头标配动作。

4. 下半场:HBM4E 已经鸣枪​

下一代产品的竞争在 2026 年同步开启:

  • SK 海力士:HBM4E 送样提前至 2026 年年中(原计划下半年),COMPUTEX 2026 已展出 12 层样品,单引脚速率最高 16Gbps、单堆栈带宽约 4TB/s,计划 2027 年量产;
  • 三星:2026 年 Q2 起向头部客户交付业界首批 HBM4E 样品(48GB),目标 2027 年拿下 HBM4E 市场 50% 以上份额;
  • 美光:HBM4E 预计 2027 年量产,首批样品采用 1γ 制程 DRAM。

5. 对采购方与投资观察者的启示​

  • 供应商组合成为第一优先级采购变量:同样买 Rubin 平台,采用哪家 HBM 供应商的整机,交付周期可能相差数月;
  • 国产链条的机会窗口:HBM 供不应求叠加出口管制,长鑫等国产存储玩家的 HBM 进展值得持续跟踪——国产 AI 芯片(昇腾 950 系列等)对 HBM 的需求同样在放量;
  • 跟踪三个领先指标:头部云厂商 Rubin 机架采购量、HBM 设备供应商订单(美光已加大 PO)、HBM4E 送样→量产的时间差。

相关链接​

参考资料​


本文基于 2026 年 9 月初 Counterpoint Research、CFM 闪存市场及韩媒报道整理。份额与产能数据均为研究机构/供应链预估口径,实际以各原厂财报为准。

Huawei Ascend 910C Deep Dive: Specs, Deployment, and Full Performance Overview

· 8 min read
Industry Research Team

Huawei Ascend 910C (Ascend 910C), Huawei's third-generation Ascend AI chip, adopts innovative dual-die (Chiplet) packaging and began mass supply in May 2025, becoming the backbone of domestic AI compute.

This article comprehensively analyzes this domestic flagship AI chip across four dimensions: technical specs, deployment cases, performance comparison, and market positioning.


1. Core Technical Specifications​

1.1 Chip architecture and process​

ItemParameter
ArchitectureDa Vinci (dual-die packaging)
ProcessSMIC N+2 (7nm-class)
PackagingChiplet (2× Ascend 910B compute dies)
Transistors~53 billion
Die size~800mm² (estimated)

Technology highlights:

  • Dual-die Chiplet packaging integrates two 910B chips, breaking the single-die yield bottleneck
  • Centerless I/O die design lets the two compute dies interconnect directly, reducing communication latency
  • SMIC N+2 process delivers 7nm-class performance with a controllable, autonomous supply chain

1.2 Compute performance​

PrecisionComputeReference
BF16800 TFLOPS~60% of NVIDIA H100
FP16~800 TFLOPSClose to H100 at same precision
INT8~1600 TOPSClear inference advantage
FP32Not disclosedTraining mainly uses BF16/FP16

Performance characteristics:

  • 800 TFLOPS at BF16, a new domestic AI chip compute benchmark
  • ~2× compute over 910B (dual-die stacking + architecture optimization)
  • No FP8 precision support (NVIDIA Blackwell's strength)

1.3 Memory and interconnect​

ItemParameter
HBM typeHBM2E (8 stacks)
Memory capacity~128 GB (combined dual-die)
Memory bandwidth784 GB/s
Interconnect protocolHuawei AscendLink (in-house)
Interconnect bandwidth400 GB/s unidirectional (800 GB/s bidirectional)

Memory advantages:

  • 128GB capacity supports full-pipeline training of hundred-billion-parameter models
  • 784 GB/s is a high-end configuration among HBM2E solutions
  • In-house AscendLink protocol supports 384-chip all-optical interconnect

1.4 Power and energy efficiency​

ItemParameter
TDP (dual-die)~310 W
Energy efficiency (BF16)~2.58 TFLOPS/W
vs. H100~45% of H100's power, comparable energy efficiency

Energy efficiency advantages:

  • At equal compute, significantly lower power than NVIDIA H100 (700W)
  • 7nm-class process, ~30% better energy efficiency than 910B
  • Suited to large-scale cluster deployment, reducing data center PUE pressure

2. Key Deployment Cases​

2.1 CloudMatrix 384 SuperNode​

System specs:

ItemConfiguration
Chip count384 Ascend 910C
Cabinets16 (12 compute + 4 network)
Total HBM~49 TB (128GB × 384)
InterconnectAll-optical mesh network
Optical modules6,912 LPO optical modules
System BF16 compute~300 PFLOPS

Performance comparison:

  • CloudMatrix 384's total BF16 compute exceeds NVIDIA GB200 NVL72 (72× B200)
  • In large-model training, 384-chip 910C linear scaling efficiency reaches 85%+
  • Supports smooth scaling to ten-thousand-card clusters for ultra-large training

Deployment progress:

  • As of June 2026, over 500 CloudMatrix 384 SuperNodes deployed
  • Key customers: China Telecom, China Mobile, China Unicom, Huawei Cloud, iFlytek
  • Scenarios: large-model training, smart customer service, autonomous-driving simulation, scientific computing

2.2 DeepSeek-V4-Pro full-parameter post-training​

Breakthrough significance:

On June 5, 2026, the AI training platform of Shenzhen Hetao College — together with Harbin Institute of Technology (Shenzhen), Shenzhen Big Data Research Institute, Huawei, and Shenzhen Zhicheng AI Compute Platform — completed full-parameter post-training of the 1.6-trillion-parameter DeepSeek-V4-Pro large model on an Ascend 910C compute cluster.

Technical highlights:

  • Among the world's first to run full-parameter post-training of a trillion-parameter model on a domestic compute platform
  • Validates Ascend 910C maturity in ultra-large-model training
  • Proves domestic AI chips now have the capability to replace imported chips

Performance data (official disclosure):

  • Training throughput: ~60% of an H100 cluster (BF16 precision)
  • Memory utilization: 92% (128GB HBM2E capacity advantage)
  • Interconnect efficiency: 384-chip linear scaling efficiency 85%+
  • Stability: 30 consecutive days of training with no failures

2.3 Commercial deployment cases​

Case 1: A provincial big-data center (300 P FLOPS compute center)​

  • Scale: 300 P FLOPS AI compute (~1,000× 910C)
  • Scenarios: government large model, city brain, smart transportation
  • Deployment: September 2025
  • Investment: ~¥200M (120 servers)

Case 2: Huawei Cloud AI training platform​

  • Chips: over 10,000 Ascend 910C
  • Customers served: over 500 enterprises
  • Model support: Pangu large model, third-party open-source models (LLaMA, ChatGLM, etc.)
  • Global deployment: China, Southeast Asia, Middle East, Latin America

Case 3: iFlytek smart education​

  • Scale: 256 Ascend 910C
  • Scenarios: smart-education large model, speech recognition, machine translation
  • Performance: 90% faster training than 910B

3. Performance Comparison Analysis​

3.1 vs. NVIDIA H100​

ItemAscend 910CNVIDIA H100Notes
BF16 compute800 TFLOPS~1,300 TFLOPS910C ~60% of H100
HBM capacity128 GB80 GB910C +60%
HBM bandwidth784 GB/s3.35 TB/sH100 clear bandwidth lead
TDP310 W700 W910C only 45% of H100 power
Process7nm (SMIC N+2)4nm (TSMC)H100 more advanced
Software ecosystemCANN (CUDA-compatible)CUDAH100 more mature
SupplyChina autonomousExport-controlled910C no supply-chain risk

Conclusion:

  • In raw compute, 910C is ~60% of H100
  • In memory capacity, 910C leads by 60%, suited to large-model training
  • In energy efficiency, 910C clearly outperforms H100
  • In supply chain security, 910C wins outright

3.2 vs. Ascend 910B​

ItemAscend 910CAscend 910BImprovement
ArchitectureDual-die ChipletSingle die—
BF16 compute800 TFLOPS~400 TFLOPS+100%
HBM capacity128 GB64 GB+100%
TDP310 W310 WFlat (single-die power)
ProcessSMIC N+2SMIC N+2Same
Yield~40%~30%+33%

Conclusion:

  • 910C's dual-die packaging doubles compute and memory capacity
  • Yield up from 910B's 30% to 40%, lowering manufacturing cost
  • At equal power, 100% performance gain, significantly better energy efficiency

3.3 Inference performance (DeepSeek measured)​

Test environment:

  • Model: DeepSeek-V3 (671B parameters)
  • Hardware: Ascend 910C vs NVIDIA H100
  • Precision: BF16
  • Batch size: 64

Results:

MetricAscend 910CNVIDIA H100Ratio
Inference speed (tokens/s)8,50014,20060%
First-token latency (ms)12085141%
Power (W)31070044%
Cost (¥10k/card)~10~1856%

Conclusion:

  • 910C inference speed is 60% of H100, but power only 44%
  • In cost-sensitive scenarios, 910C's cost-performance advantage is clear
  • For China-market localization needs, 910C is the only option

4. Market Positioning and Competitive Advantages​

4.1 Target markets​

Core markets:

  1. Chinese government and SOEs: localization, data security, autonomy
  2. Large-model startups: cost-sensitive, high compute demand
  3. Operators and cloud providers: large-scale deployment, high efficiency requirements
  4. Research and education: ultra-large-scale computing, talent development

Edge markets:

  1. Autonomous driving: end-to-end large-model training
  2. Smart healthcare: medical imaging, drug discovery
  3. Fintech: risk control, robo-advisory

4.2 Competitive advantages​

AdvantageDescription
AutonomySMIC N+2 process + Huawei in-house architecture, no supply-chain risk
Large memory128GB HBM2E, full-pipeline training of hundred-billion-parameter models
High energy efficiency310W TDP delivers 800 TFLOPS, close to H100 efficiency
System scalingCloudMatrix 384 SuperNode, total compute exceeds GB200 NVL72
Software ecosystemCANN CUDA-compatible, lower migration cost
Cost advantage~¥100k/card, ~44% cheaper than H100

4.3 Weaknesses and improvement directions​

WeaknessImprovement direction
Single-chip computeNext-gen 910D to adopt 3nm, target doubling
HBM bandwidth950 series to adopt in-house HBM (HiBL 1.0), bandwidth to 4 TB/s
Software ecosystemContinued CANN + MindSpore investment, expand developer community
ProcessDeep cooperation with SMIC to ramp N+3 (5nm-class)

5. 2026 Shipment Plan and Market Forecast​

5.1 Shipment plan​

PeriodShipmentsCumulativeKey customers
2025 Q2-Q4200k200kHuawei Cloud, China Telecom
2026 Q1-Q2300k500kChina Mobile, China Unicom, iFlytek
2026 Q3-Q4300k800kGovernment projects, large-model startups
20271,000k1,800kGlobal market (Southeast Asia, Middle East, Latin America)

Capacity bottleneck:

  • SMIC N+2 capacity ~100k wafers/month, Ascend 910C ~30% of that
  • 2026 plan of 800k chips needs ~400k wafers, requiring 80%+ utilization
  • Huawei prioritizes 910C capacity via deep SMIC cooperation

5.2 Market forecast​

China AI chip market (2026):

  • Total: ~¥50B
  • Domestic share: ~35% (¥17.5B)
  • Ascend 910C share: ~60% (¥10.5B, ~800k chips)

Global AI chip market (2026):

  • Total: ~$200B
  • Huawei share: ~5% ($10B)
  • Growth drivers: China-market localization + Belt and Road exports

6. Summary and Outlook​

6.1 Core conclusions​

  1. Ascend 910C is a milestone domestic AI chip, with comprehensive breakthroughs in compute, memory, energy efficiency, and system scaling
  2. CloudMatrix 384 SuperNode proves domestic chips can replace imported ones
  3. DeepSeek-V4-Pro training success validates 910C maturity in ultra-large-model training
  4. 800k chips shipped in 2026, projected 60% of China's AI chip market

6.2 Future outlook​

Short term (2026-2027):

  • 910C continues ramping, shipments exceed 1,000k
  • CloudMatrix 384 deployments over 1,000 units
  • Software ecosystem (CANN + MindSpore) maturity approaches 70% of CUDA

Medium term (2028-2029):

  • Next-gen 910D mass production, 3nm process, target 1.6 PFLOPS BF16
  • 950 series (PR/DT) becomes inference-market mainstay, share over 30%
  • 960/970 launch, N+3 process, supports trillion-parameter models

Long term (2030+):

  • Huawei Ascend series becomes TOP 3 of the global AI chip market
  • Domestic AI chips exceed 20% of the global market
  • Transition from "following" to "running alongside" to "leading"

References​

  1. Huawei Ascend 910C — Baidu Baike
  2. Huawei Ascend series AI chip detailed parameter comparison (2025-2028) — EET-China
  3. Huawei Ascend 910C compute cluster powers domestic chip's successful trillion-scale AI large-model training — QQ News
  4. Huawei Ascend 910C completes DeepSeek V4 Pro training — Huxiu
  5. Huawei Ascend 910C measured efficiency surpasses H100, AI Infra software-hardware co-optimization shines at ten-thousand-card cluster — CNBlogs

Last updated: June 10, 2026

Computex 2026 Wrap-Up: AI PC Chip War Begins, NVIDIA RTX Spark Arrives Fall 2026

· 3 min read
Industry Research Team

June 6, 2026 — COMPUTEX 2026 concluded yesterday in Taipei. Under the theme "AI Together," this year's event set records with 1,500+ exhibitors and 6,000 booths. The head-to-head battle between NVIDIA, Intel, and AMD in the AI PC space was the defining story of the show.

1. NVIDIA RTX Spark: June Launch at $1,399​

Less than a week after its COMPUTEX debut, the NVIDIA-MediaTek RTX Spark Superchip confirmed its commercial timeline:

DetailInfo
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Starting PriceNot yet announced (analysts estimate $3,000-4,000)
Core SpecsArm CPU (up to 20 cores) + Blackwell GPU (6,144 CUDA cores)
Unified Memory128 GB LPDDR5X (300 GB/s)
Model CapacityRuns 120B parameter models, up to 1M token context

Market Reaction: AMD, Intel, and Qualcomm shares fell following the announcement. Analysts believe RTX Spark will reshape the market across three fronts — Windows AI PCs, creator workstations, and edge inference nodes.


2. Intel 18A in Full Production: Clearwater Forest + Crescent Island​

Intel CEO Lip-Bu Tan delivered his first COMPUTEX keynote with two key updates:

Clearwater Forest (Xeon 6+)​

  • 288 cores, Darkmont architecture
  • First Intel 18A process node data center CPU
  • Foveros Direct 3D packaging
  • Now in full production

Crescent Island AI GPU​

  • 480 GB LPDDR5x memory
  • 350 W air-cooled PCIe form factor
  • Native FP4 support, targeting agentic inference
  • Shipping H2 2026

"As AI moves into the agentic era, the CPU returns to the center of modern AI infrastructure." — Lip-Bu Tan


3. AMD Ryzen AI 400 Series Now Shipping​

AMD showcased the Ryzen AI 400 series (Zen 5 + Zen 5C hybrid + XDNA2 NPU) at COMPUTEX:

  • NPU performance: 60 TOPS, the highest in x86
  • 7 consumer SKUs + commercial PRO series
  • Multiple OEM models already available or launching soon
  • Advancing AI 2026 summit set for July in San Francisco

4. Chinese Domestic Chips Gaining Momentum​

VendorProductStatus
HuaweiAscend 950PR/950DTIn production, self-developed HBM
CambriconMLU6902 PFLOPS FP8, shipping
Moore ThreadsMTT S50001,000 TFLOPS, specs public

5. The AI PC Era: Three-Way Roadmap Comparison​

DimensionNVIDIA RTX SparkIntel Clearwater Forest + Crescent IslandAMD Ryzen AI 400
CPU Cores20-core Grace (Arm)288-core Darkmont (x86)Up to 12-core Zen5+5C
GPU/NPUBlackwell GPUCrescent Island (discrete GPU)XDNA2 NPU (60 TOPS)
AI Compute1 PFLOPSTBD60 TOPS NPU
TargetPersonal AI agentsDual-track: DC + AI PCCopilot+ PC
ProcessTSMC 4NPIntel 18ATSMC 4nm
AvailabilityJune 2026H2 2026Shipping now

This Week in AI Compute (6/1 – 6/6)​

DateEvent
Jun 1NVIDIA GTC Taipei: RTX Spark, Vera Rubin production, DGX Station for Windows
Jun 1Intel unveils Crescent Island, Clearwater Forest
Jun 2COMPUTEX 2026 opens: "AI Together"
Jun 5COMPUTEX closes: 1,500+ exhibitors, record scale
Jun 6RTX Spark confirmed June launch at $1,399

Sources: COMPUTEX Daily, Tencent News, Phoenix Technology, Xueqiu, The Silicon Review.

Computex 2026 AI Compute Card Major Events: DGX Station for Windows, Intel Crescent Island, and More Major Launches

· 4 min read
Industry Research Team

June 1-5, 2026, Taipei — Computex 2026 (Taipei International Information Technology Show) wrapped up successfully this week. With the theme "AI Together," industry giants including NVIDIA, Intel, AMD, and Qualcomm unveiled numerous AI compute products in rapid succession. Below, MirrorFrog brings you a roundup of the most noteworthy developments in the compute card space this week.

① NVIDIA DGX Station for Windows: A Desktop AI Supercomputer​

NVIDIA officially launched the DGX Station for Windows during its Computex 2026 keynote, calling it "the world's most powerful desktop AI supercomputer."

Core Specifications​

ItemSpecification
ChipGB300 Grace Blackwell Ultra Desktop Superchip
GPU Memory252 GB HBM3e (7.1 TB/s)
CPU Memory496 GB LPDDR5X (396 GB/s)
Unified Memory748 GB (NVLink-C2C interconnect)
FP4 Compute20 PFLOPS (sparse)
FP8 Compute10 PFLOPS (sparse)
NetworkConnectX-8 SuperNIC, up to 800 Gb/s
Model CapacityCan run 1 trillion parameter models
System Power1,600 W
Operating SystemMicrosoft Windows
ShippingQ4 2026

Significance: DGX Station compresses AI compute power (20 PFLOPS FP4) that previously required datacenter-class clusters into a single desktop workstation. 748GB of unified memory means developers can run models with hundreds of billions or even trillions of parameters locally, without cloud dependency.


② Intel Crescent Island: Inference-Specialized AI GPU​

At Computex, Intel disclosed detailed specifications for its next-generation datacenter AI inference GPU, Crescent Island.

ItemSpecification
MemoryUp to 480 GB LPDDR5x
Power350 W (PCIe form factor)
Precision SupportFP4/MXFP4 → FP64 (full precision coverage)
TargetAI inference workloads (Agentic Inference)
PositioningBetter price-performance than HBM solutions
ShippingH2 2026

Significance: Crescent Island represents Intel's key strategic move in the AI inference market. 480GB of massive LPDDR5x memory (non-HBM) means significantly lower cost compared to NVIDIA H200/B200 and other competing products, targeting enterprise inference deployment scenarios.


③ Intel Xeon 6+ (Clearwater Forest): First Intel 18A Datacenter CPU​

Intel also unveiled the new Xeon 6+ processor, codenamed Clearwater Forest, its first datacenter CPU built on the 18A process:

  • 288 Darkmont architecture cores
  • L2 288MB + L3 576MB cache
  • 12-channel DDR5-8000 memory
  • Foveros Direct 3D advanced packaging
  • AI Agent Era: CPU returns to the center of infrastructure

④ NVIDIA RTX Spark Ecosystem Takes Shape​

This week, the RTX Spark super chip developed in collaboration between NVIDIA and MediaTek continued to generate buzz. Multiple OEMs showcased RTX Spark-based laptop and compact desktop prototypes:

  • ASUS, Dell, HP, Lenovo, Microsoft Surface, MSI all confirmed as launch partners
  • Equipped with 20-core Grace CPU + Blackwell GPU (6144 CUDA cores)
  • AI compute 1 PFLOPS
  • Retail availability Fall 2026

⑤ Intel × Foxconn AI Infrastructure Partnership​

Intel and Foxconn announced a joint AI infrastructure initiative, covering the complete chain from chip → server → rack-scale system, targeting the datacenter market opportunity driven by surging AI inference demand.


⑥ Domestic AI Chip Developments​

According to the IDC 2025 annual report, total AI accelerator card shipments in China reached approximately 4 million units, with domestic vendors shipping approximately 1.65 million units, capturing a market share exceeding 41%. Huawei's Ascend 950 series has entered mass production and delivery, while Cambricon's MLU690 has begun shipping to internet customers.


This Week's Compute Roundup​

VendorProductHighlightTimeline
NVIDIADGX Station for Windows20 PFLOPS, 748GB unified memoryQ4 2026
NVIDIARTX Spark1 PFLOPS AI PC chipFall 2026
IntelCrescent Island GPU480GB LPDDR5x, 350WH2 2026
IntelXeon 6+ (Clearwater Forest)288 cores, Intel 18AH2 2026
Intel + FoxconnAI infrastructure partnershipChip→rack full chainStrategic partnership
HuaweiAscend 950PR/DT1 PFLOPS FP8, self-developed HBMIn mass production
CambriconMLU6902 PFLOPS FP8, 192GB HBM3EShipping

Sources: NVIDIA GTC Taipei 2026 / Computex 2026 official announcements, Intel press releases, ifeng Tech, IT Home.

Huawei Ascend 950 Mass Production and the Full Picture of China's AI Chip Ecosystem

· 4 min read
Industry Research Team

June 2026 — Huawei's Ascend 950 series (950PR / 950DT) has entered formal mass production and delivery, a landmark event for China's AI chip industry in 2026. Meanwhile, Cambricon's MLU690 has begun shipping and Moore Threads has announced MTT S5000 specifications, formally establishing China's tri-polar AI chip landscape.

Ascend 950 Series: A Historic Breakthrough with Self-Developed HBM​

Huawei HiSilicon's Ascend 950 series is the fourth-generation Ascend AI chip, first revealed at Huawei Connect 2025 in September and entering mass production in Q1 2026.

950PR (Prefill Inference Specialized)​

ItemSpecification
ArchitectureDa Vinci v5 (SIMD + SIMT dual-model)
ProcessN+2 (SMIC domestic)
HBMHiBL 1.0 (Huawei self-developed) , 128 GB
FP8 Compute1 PFLOPS (HiF8 format)
TDP~400 W
TargetInference Prefill (video recommendation, real-time interaction)

950DT (Decode + Training Specialized)​

ItemSpecification
ArchitectureDa Vinci v5 (SIMD + SIMT dual-model)
ProcessN+2 (SMIC domestic)
HBMHiZQ 2.0 (Huawei self-developed) , 144 GB, 4 TB/s
FP8 Compute1 PFLOPS (HiF8 format)
TDP~500 W
TargetInference Decode + Model Training

Historical Significance​

Self-developed HBM (HiBL 1.0 / HiZQ 2.0) represents the most important technical breakthrough of Huawei Ascend 950 — this is the first time a Chinese enterprise has achieved self-developed mass production of HBM memory, completely eliminating dependence on SK Hynix / Samsung HBM supply. Combined with the domestic N+2 process, Ascend 950 has achieved full-chain domestic production from HBM → Compute Die → Packaging → System.

Cambricon MLU690: China's Only Native FP8 Support​

Cambricon's seventh-generation AI chip MLU 690 (Siyuan 690) began volume production and shipping in H1 2026. This is the first domestic AI chip with native FP8 precision support.

ItemMLU 690
Process5nm (TSMC / SMIC)
FP8 dense2 PFLOPS
HBM192GB HBM3E, 5 TB/s
TDP~500 W
Unit Price (OAM)~$8,000-12,000

MLU 690's FP8 compute power (2 PFLOPS dense) is on paper comparable to NVIDIA Blackwell (B200 FP8 4.5 PFLOPS sparse). Leveraging its financing advantage as a STAR Market listed company, Cambricon targets 2026 revenue of ¥15-20B (2025: ¥7.2B).

Moore Threads MTT S5000: From Graphics to Training-Inference Unified​

Moore Threads publicly disclosed detailed specifications of the MTT S5000 in February 2026, featuring the fourth-generation MUSA "Pinghu" architecture, single-card AI compute of 1,000 TFLOPS, 80GB GDDR6X memory, 1.6 TB/s bandwidth.

Moore Threads pursues a full-function GPU path (graphics rendering + AI compute + general-purpose compute), closest to NVIDIA's strategy. The founding team comes from former NVIDIA China, and the MUSIFY toolchain helps auto-migrate CUDA code to the MUSA platform, lowering ecosystem migration costs.

China's Tri-Polar AI Chip Landscape​

DimensionHuawei AscendCambriconMoore Threads
Core ArchitectureDa Vinci v5MLUv07MUSA 4th Gen
ProcessN+2 domestic5nm6nm
FP8 Compute~1 PFLOPS2 PFLOPS0.5 PFLOPS (estimated)
HBM Self-Sufficiency✅ Self-developed HiBL/HiZQ❌ Purchased❌ Purchased
EcosystemCANN + MindSporeNeuWare + MindSporeMUSA + MUSIFY
AdvantageFull-chain domesticHighest FP8 computeFull-function + CUDA migration
2025 Revenue(Huawei internal)¥7.2B¥2.2B

Global Market Comparison (Q2 2026 Update)​

TierVendorFlagship ChipFP8/PFLOPSHBMMass Production
Tier 1NVIDIARubin R20025 PF (sparse)288GB HBM42026 H2
Tier 2AMDMI40020 PF (dense)432GB HBM42026
HuaweiAscend 950DT1 PF (dense)144GB self-developed HBM2026 Q1
CambriconMLU6902 PF (dense)192GB HBM3E2026 H1
AWSTrainium 35.7 PF (dense)144GB HBM2025 Q4 GA
Tier 3IntelGaudi 31.8 PF128GB HBM2eIn production
GoogleTPU v74.6 PF(TFLOPS)192GB HBM2025
Moore ThreadsMTT S50001 PF80GB GDDR6X2025 Q1

Note: NVIDIA uses sparse compute as standard, while AMD / Huawei / Cambricon use dense — not directly comparable.

Outlook for H2 2026​

  • NVIDIA Rubin R200: Official shipment in H2 2026, 288GB HBM4, 6-chip CoWoS-L packaging
  • Huawei Ascend 960: Roadmap H2 2027, expected FP8 compute doubled to 2 PFLOPS
  • Cambricon MLU790: Expected 2027, 3nm, 384GB HBM4, 2.5 PFLOPS
  • Moore Threads: Next-gen GPU expected with HBM3, 2× MTT S5000 compute

By 2026, China's AI chip industry has formed a complete product matrix from Training (Cambricon MLU690 / Ascend 950DT) → Inference (Ascend 950PR / Moore Threads S5000) → Systems (CloudMatrix / Distributed Clusters).


This article is based on public information from Huawei Connect 2025 (2025-09-18), industry analysis reports from April 2026, and the latest market data as of June 2026.

NVIDIA Launches RTX Spark: AI Compute Enters the Personal Computer Era

· 3 min read
Industry Research Team

June 1, 2026, Taipei — During the Computex 2026 opening keynote, NVIDIA CEO Jensen Huang officially unveiled the RTX Spark super chip, marking NVIDIA's formal entry into the personal computer processor market dominated by Intel, AMD, Qualcomm, and Apple.

RTX Spark: The "Heart" of the Personal AI Computer​

RTX Spark was developed in collaboration between NVIDIA and MediaTek, featuring a heterogeneous package with a 20-core Grace CPU + Blackwell RTX GPU, equipped with 6144 CUDA cores. AI compute reaches 1 PFLOPS (one quadrillion floating-point operations per second), meaning personal computers now possess computing power comparable to a datacenter-class H100 GPU for the first time.

SpecificationRTX Spark
CPU20-core Grace (MediaTek collaboration, Arm architecture)
GPUBlackwell RTX (6144 CUDA cores)
AI Compute1 PFLOPS
TargetPersonal AI Agent, local LLM inference
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Form FactorLaptop SoC + compact desktop workstation

Jensen Huang's "Full-Stack AI" Strategy​

The launch of RTX Spark is a key step in NVIDIA's "full-stack AI" strategy. Jensen Huang stated during the keynote: "AI should not only run in the cloud. Everyone's computer should have the ability to run AI agents."

RTX Spark transforms NVIDIA from a datacenter GPU monopolist into a full competitor in the personal computing market. Following the announcement, shares of AMD, Intel, and Qualcomm fell accordingly.

Market Impact​

  • Intel: Personal computer AI processor business faces direct threat
  • AMD: Ryzen AI series must compete at the same level
  • Qualcomm: Snapdragon X Elite's Copilot+ PC positioning challenged
  • Apple: M-series chips are no longer the only high-performance AI PC option

Vera Rubin Platform Enters Full Mass Production​

During the same keynote, Jensen Huang also announced that the NVIDIA Vera Rubin platform has entered full mass production. Rubin R200 features a 6-chip CoWoS-L package (1× Vera CPU + 2× Rubin GPU die + I/O/HBM die), equipped with 288GB HBM4, 22 TB/s bandwidth, and 50 PFLOPS FP4 compute (sparse).

The Rubin NVL72 rack (72 Rubin GPUs + 36 Vera CPUs) will begin shipping in H2 2026.

Other Highlights from Computex 2026​

  • AMD: Showcased the MI350 series (192GB HBM3e, 5 PFLOPS FP8 dense), officially launching in June
  • Intel: Jaguar Shores publicly unveiled for the first time
  • Qualcomm: AI 200 / 300 series inference card roadmap updated
  • Domestic AI Chip Zone: Huawei, Cambricon, Moore Threads, and others showcased their latest products

Industry Significance​

The launch of RTX Spark means AI compute is no longer confined to datacenters. Individual developers, designers, and researchers will be able to run large model tasks locally that previously required cloud GPUs, potentially redefining the market landscape for personal AI computing.

The mass production of Vera Rubin further consolidates NVIDIA's absolute leadership in datacenter AI training. Together, both product lines form NVIDIA's full-stack AI computing landscape of "cloud training + personal inference."


This report is based on official NVIDIA announcements from Computex 2026 / GTC Taipei on June 1, 2026.

China AI Chip Landscape 2025: Ascend, Cambricon, Hygon — Who Will Dominate?

· 5 min read
Industry Research Team

Escalating U.S. export controls are forcing China's AI chip industry to accelerate self-reliance. By 2025, the discussion around domestic Chinese AI chips has shifted from "are they usable?" to "which one should I choose?"

This article systematically reviews the major players, core products, and actual deployment status of domestic AI chips, helping developers and procurement decision-makers understand the competitive landscape.