Skip to main content

Huawei Ascend 920: China's Highest Bandwidth at 4 Tbps + 3× H20 Compute for Domestic Substitution

· 5 min read
Industry Research Team

Huawei Ascend 920 (昇腾 920) entered large-scale mass production in 2025 H2, representing a major breakthrough for Chinese domestic AI chips. This article analyzes its specifications, comparison with NVIDIA H20, the CloudMatrix 384 Ultra system, and its significance for China's AI industry.

Core Specifications​

ItemAscend 910CAscend 920Improvement
ArchitectureDa Vinci v3Da Vinci v4New generation
Process7nm6nm (SMIC domestic)More advanced
Chiplets2× (dual die)2×same
HBM capacity~128 GB~96 GBslight decrease
HBM bandwidth3.2 Tbps4 Tbps1.25×
BF16 compute780 TFLOPS900+ TFLOPS1.15×
FP16 compute1,560 TFLOPS1,800 TFLOPS1.15×
INT8 compute3,120 TOPS3,600 TOPS1.15×
TDP~310 W~400 W+29%
Release date2025-042025 H2—

4 Tbps bandwidth = China's highest domestic HBM bandwidth, a 25% improvement over Ascend 910C. The 900+ BF16 TFLOPS compute also surpasses 910C.

Ascend 920 vs NVIDIA H20 (Target Comparison)​

NVIDIA H20 is the "compliance" AI chip specifically designed for the Chinese market under U.S. export controls:

MetricAscend 920NVIDIA H20
PositioningDomestic substitutionChina-compliant AI chip
Process6nm (SMIC)TSMC 4N (partially domestic after restrictions)
Memory~96 GB96 GB HBM3
Memory bandwidth4 Tbps4.0 Tbps
BF16 compute900 TFLOPS296 TFLOPS
BF16 compute ratio3×1× (baseline)
InterconnectHCCS 1.2 TbpsNVLink 900 GB/s
SoftwareCANN + MindSporeCUDA (restricted)
Import compliance✅ Domestic⚠️ U.S. export controls

💡 Ascend 920 significantly leads H20 in BF16 compute (3×), with 4 Tbps bandwidth on par with H20. This is a key victory for domestic substitution.

CloudMatrix 384 Ultra System​

Ascend 920 will be used in the CloudMatrix 384 Ultra supernode system:

ItemConfiguration
Chip count384 Ascend 920 chips
Rack count16 (12 compute + 4 network)
Total HBM~36 TB (96GB × 384)
InterconnectFully optical mesh, 8,000+ LPO optical modules
BF16 compute (system)~345 PFLOPS (estimated 900 × 384)
TDP (system)~150 kW

CloudMatrix 384 Ultra system-level BF16 compute of ~345 PFLOPS ≈ 2.4× NVIDIA GB200 NVL72 cluster (~144 PF FP8 dense).

Why Ascend 920 Is the Key Victory for Domestic Substitution?​

1. First Time Surpassing H20 by 3× in Compute​

PeriodDomesticNVIDIA China EditionMultiple
2023910B = 320 TFLOPSH20 = 296 TFLOPS1.08×
2024910B = 320 TFLOPSH20 = 296 TFLOPS1.08×
2025 H1910C = 780 TFLOPSH20 = 296 TFLOPS2.6×
2025 H2920 = 900 TFLOPSH20 = 296 TFLOPS3.0×

Starting from 2025 H2, Chinese domestic AI chip compute stably surpasses H20 by three times.

2. SMIC 6nm Domestic Process​

Ascend 920 uses SMIC N+1 / N+2 6nm process:

  • ✅ Fully indigenous and controllable
  • ✅ Not subject to U.S. export controls
  • ⚠️ Yield and cost still lag behind TSMC 4N

3. 4 Tbps — China's Highest Domestic HBM​

Ascend 920's 4 Tbps HBM bandwidth:

  • First domestic chip to reach 4 Tbps level (previous max 3.2 Tbps)
  • On par with H20
  • Presumed to use CXMT (ChangXin Memory Technologies) HBM3 or indigenous HBM

4. CANN + MindSpore Software Stack​

  • CANN 8.x (Compute Architecture for Neural Networks): analog to CUDA
  • MindSpore 2.4+: Huawei's indigenous AI framework
  • PyTorch 2.3+ MindSpore backend: PyTorch compatible
  • vLLM 0.7+ Ascend backend: low-latency inference
  • ONNX-Runtime Ascend backend: cross-framework inference
  • Atlas 900/950 series servers: OEM complete systems

China Market Deployment Status​

Scaled-Up Customers​

CustomerApplication
China MobileLarge model training (990M customers)
China TelecomIntelligent customer service + business insights
China UnicomGovernment + industry AI
State GridPower grid scheduling + fault prediction
CNPCExploration + logistics optimization
Major banksRisk control + anti-fraud
Internet companies (Baidu, Alibaba, Tencent)LLM inference

Industry Layout​

  • Government: 100% domestic requirement
  • Finance: policy-driven domestic requirement
  • Telecom: fast HBM domestication progress
  • Energy: fast HBM domestication progress
  • Internet: sensitive workloads shifting to domestic
  • Education / Healthcare: gradual domestication

Limitations and Challenges​

LimitationImpact
FP8/FP4 supportAscend 920 still BF16/FP16-primary, FP8 optimization in progress
HBM capacity96 GB is below NVIDIA Rubin R200 288 GB / AMD MI400 432 GB
CUDA compatibilityCANN 8 still requires migration; direct CUDA app execution is limited
SMIC 6nm yield10-20% lower yield than TSMC 4N
HBM sourceCXMT HBM production capacity limited
Interconnect bandwidthHCCS 1.2 Tbps far below NVLink 6 (3.5 TB/s)

Comparison with Contemporaneous Domestic Chips​

VendorChipBF16 ComputeHBM BandwidthMass Production
HuaweiAscend 920900 TFLOPS4 Tbps2025 H2
HuaweiAscend 910C780 TFLOPS3.2 Tbps2025-04
CambriconSiyuan 590~480 TFLOPS2.4 Tbps2024
Moore ThreadsMTT S5000~250 TFLOPS1.6 Tbps2024
BirenBR104~300 TFLOPS1.6 Tbps2024
IluvatarCoreX Bi-150~200 TFLOPS1.2 Tbps2024

Huawei Ascend 920 maintains a clear lead among Chinese domestic AI chips.

Detailed Product Pages​

Summary​

Huawei Ascend 920 is a key victory for Chinese AI chips in 2025 H2:

  1. 900+ BF16 TFLOPS = 3× H20 — first time stably surpassing H20 by three times
  2. SMIC 6nm domestic — indigenous and controllable
  3. 4 Tbps — China's highest domestic HBM bandwidth — HBM domestication breakthrough
  4. CloudMatrix 384 Ultra system — single system surpasses GB200 NVL72
  5. CANN + MindSpore — maturing software ecosystem

Starting from 2025 H2, China's AI industry enters a new phase where "domestic chips can independently support large-scale AI applications."