Skip to main content

Cambricon MLU290 (Siyuan 290)

Product Overview​

The Cambricon MLU290 (Siyuan 290) is Cambricon's first training-class AI chip, released in 2020 and in mass production in early 2021. Built on TSMC's 7nm process, it integrates 46 billion transistors based on the extended MLUv02 architecture (MLUv02 Extended). It marked Cambricon's move from "cloud inference" to full-scenario coverage of "cloud training + inference + hybrid computing", and formed training cluster solutions together with the contemporary Xuansi 1000 intelligent accelerator (4 MLU290 chips integrated in 2U).

Compared with the MLU270, the MLU290 achieved 4× higher peak compute, 12× higher memory bandwidth, and 19× higher inter-chip communication bandwidth. It is the starting point of Cambricon's training product line, later evolving into the MLU370 (Chiplet training-inference integrated) and the MLU590 (third-generation flagship).

Core Specifications​

ParameterValue
ArchitectureCambricon extended MLUv02 architecture (MLUv02 Extended)
Process NodeTSMC 7nm
Transistor Count46 billion
INT8 Compute512 TOPS
INT16 Compute256 TOPS
CINT32 Compute64 TOPS
FP16 / BF16 Compute256 TFLOPS (estimated, from consolidated parameter tables; official release materials do not list floating-point compute)
FP32 Compute32 TFLOPS (estimated, same source)
Memory Capacity32 GB
Memory TypeHBM2
Memory Bus Width4096 bit
Memory Bandwidth1.23 TB/s (1228 GB/s)
MLU Cores64
TDP350 W
InterconnectMLU-Link™ (600 GB/s aggregated per card, multi-card cluster interconnect)
InterfaceOAM (Open Accelerator Module, 54V); system interface PCIe 4.0 ×16
Launch2020 (announced) / 2021-01 (mass production)
Mass Production/AvailabilityIn mass production

⚠️ Specification notes: INT8 512 TOPS / INT16 256 TOPS / CINT32 64 TOPS / 32GB HBM2 / 1228 GB/s / 350W are consistent Cambricon official figures (WAIC 2021, OpenI community compute leaderboard). FP16 256 TFLOPS and FP32 32 TFLOPS come from third-party consolidated parameter tables and are not listed in official releases; they are marked as estimates.

Key Features​

  • First training chip: fully supports AI training, inference, or hybrid compute acceleration
  • MLU-Link™ multi-chip interconnect: first introduction of Cambricon's proprietary inter-chip interconnect, supporting multi-card training clusters
  • High-bandwidth memory: 32GB HBM2 + 1.23 TB/s bandwidth, suited to medium/large model training
  • OAM form factor: Open Accelerator Module design for easy integration by OEMs
  • Transformer optimization: tuned for large model training

Vendor Information​

ItemDetails
CompanyCambricon Technologies Corporation Limited
HeadquartersBeijing
Founded2016
IPOSTAR Market 688256

Use Cases​

  • ✅ Small/medium model training (vision, NLP, recommendation)
  • ✅ Cloud inference (high throughput)
  • ✅ Research supercomputing platforms / enterprise AI training clusters
  • ✅ Mixed-precision training
  • ❌ Ultra-large model training (32GB memory constrained)
  • ❌ Strong CUDA ecosystem dependence (requires migration to Cambricon NeuWare)

References​