Skip to main content

Cambricon MLU 590 (China AI Training/Inference)

Overview​

Cambricon Technologies is a leading Chinese AI chip company, founded in 2016 (spun out from the Institute of Computing Technology, Chinese Academy of Sciences), with its STAR Market IPO on 2020-07-20 (ticker 688256). The MLU 590 is its latest-generation dual-purpose training and inference AI accelerator: 7nm process, 256 TOPS INT8 compute, 96GB HBM2 memory, 600 GB/s bandwidth. Paired with the MindSpore full-stack AI framework (led by CAICT), key customers include government, state-owned enterprises, and Chinese internet companies.

Strategic position: Under NVIDIA H100/H200 export controls, Cambricon is one of China's national team mainstays for AI domestic replacement (alongside Huawei Ascend and Hygon DCU).

Core Specifications​

ItemSpec
ArchitectureCambricon MLU 5th Gen (MLUv05)
ProcessTSMC 7nm (with some SMIC localization)
HBM96 GB HBM2
Memory Bandwidth600 GB/s
INT8 Compute256 TOPS
BF16 Compute125 TFLOPS
FP32 Compute62.5 TFLOPS
TDP~250 W
PCIePCIe 4.0 x16
InterconnectMLU-Link (proprietary, NVLink-like)
Form FactorPCIe / OAM module
Mass Production2023-Q4
Unit Price (OAM)~$3,500-5,000

vs Previous MLU 370​

MetricMLU 590MLU 370Improvement
Process7nm7nmSame
HBM96GB HBM248GB HBM22x
Bandwidth600 GB/s307 GB/s1.95x
INT8256 TOPS128 TOPS2x
BF16125 TFLOPS64 TFLOPS1.95x
Interconnect BandwidthMLU-Link 600 GB/s200 GB/s3x
TDP250W150W+67%
Perf/W1.0 TOPS/W0.85 TOPS/W+18%

Siyuan 590 Training Cluster​

ItemConfig
Board8x Siyuan 590 OAM
Node2x Siyuan 590 servers
Cluster1024 nodes = 8192 cards
Total Compute1.05 EFLOPS BF16
Total HBM786 TB
InterconnectMLU-Link fully connected

Software Stack​

LayerFramework/ToolNotes
AI FrameworksMindSpore (Huawei/CAICT-led)PyTorch compatible
PyTorch (Cambricon backend)MLU device mapping
TensorFlow (Cambricon backend)Legacy ecosystem
CompilerBANG C/C++Cambricon proprietary language
Operator LibraryCNMLCUDA cuDNN-like
Model ZooModelZooCV/NLP/Multimodal

⚠️ Ecosystem limitations: Compared to NVIDIA CUDA + 10 years of software, Cambricon's ecosystem is only 3-4 years old. PyTorch models need conversion, BANG C has a steep learning curve, and model migration cost is relatively high.

Vendor Information​

ItemDetails
CompanyCambricon Technologies
FoundersChen Tianshi and Chen Yunji brothers (CAS ICT)
Founded2016-03
IPO2020-07-20 STAR Market (688256)
Market Cap (2026-05)~CNY 320B
2025 Revenue~CNY 7.2B (+340% YoY)
HeadquartersHaidian District, Beijing
Websitehttps://www.cambricon.com
Key CustomersChina Mobile, Inspur, Sugon, ByteDance, Zhipu AI
National Policy"East Data West Compute" recommended chip

Key Features​

  • High localization: HBM from Samsung/SK Hynix, domestic packaging (JCET)
  • Siyuan architecture evolution: MLU 100 (2018) -> 270 (2019) -> 290 (2020) -> 370 (2021) -> 590 (2023) -> 690 (2025 speculative)
  • Unified training + inference: Same hardware supports both
  • MindSpore ecosystem binding: Deep collaboration with Huawei (Ascend also uses MindSpore)
  • Multimodal support: CV / NLP / Speech / Multimodal LLM
  • Weakness: No FP8 support (NVIDIA Hopper/Blackwell 2-4x advantage), ecosystem weaker than CUDA

DeepSeek / Zhipu Performance Reference​

  • DeepSeek V3 training: Siyuan 590 cluster performance approximately 50-60% of H100 cluster
  • Zhipu GLM-4 inference: Siyuan 590 single card 256 GB/s x 4 = 1 TB/s total bandwidth, 50 tok/s inference speed (FP16 70B)
  • Stable Diffusion XL training: Siyuan 590 approx 80% A100 speed (BF16)

Use Cases​

  • ✅ China market LLM training and inference
  • ✅ Government, SOE AI projects (policy-mandated)
  • ✅ Large model inference deployment
  • ✅ Domestic replacement projects
  • ✅ Intelligent computing center construction ("East Data West Compute" hubs)
  • ❌ International market (CUDA ecosystem lock-in)
  • ❌ Cutting-edge frontier model training (FP8 missing)

Cambricon vs Huawei Ascend​

DimensionCambricon MLU 590Huawei Ascend 910C
Compute125 BF16 TFLOPS780 BF16 TFLOPS
Memory96GB HBM2128GB HBM2E
EcosystemMindSpore (PyTorch-compatible)MindSpore + CANN
National SupportSTAR Market listedHuawei in-house
Market PositionGeneral + intelligent computing centersData center + gov/enterprise cloud
2025 Revenue~CNY 7.2BIncluded in Huawei Cloud