Skip to main content

Cambricon MLU270 (Siyuan 270)

Product Overview​

The Cambricon MLU270 (Siyuan 270) is Cambricon's second-generation cloud AI chip, officially released in 2019, built on the MLUv02 architecture and positioned for high energy-efficiency AI inference acceleration in the cloud and at the edge. Compared with the first-generation MLU100, theoretical peak performance on non-sparse models improved 4× to 128 TOPS (INT8), while remaining compatible with INT4 (256 TOPS) and INT16 (64 TOPS), plus FP16/FP32 mixed precision.

The MLU270 is a key part of Cambricon's "cloud-edge-terminal" product lineup: cloud inference is delivered in accelerator card form factors such as the MLU270-S4 (70W) and MLU270-F4 (150W), with ample hardware video/image codec units for vision scenarios, suited to data center video analytics, smart cities, and other inference workloads. It forms a complete generational sequence with the subsequent MLU290 (training), MLU370 (Chiplet training-inference integrated), and MLU590 (third-generation flagship).

Core Specifications​

ParameterValue
ArchitectureCambricon MLUv02
Process NodeTSMC 16nm
INT8 Compute128 TOPS
INT4 Compute256 TOPS
INT16 Compute64 TOPS
FP16 / BF16 ComputeNot disclosed (mixed precision supported; peak not published)
FP32 ComputeNot disclosed
Memory Capacity16 GB
Memory TypeDDR4 (ECC)
Memory Bus Width256 bit
Memory Bandwidth102 GB/s
TDP70 W (MLU270-S4) / 150 W (MLU270-F4)
InterconnectPCIe 3.0 ×16
InterfacePCIe ×16 (S4 half-height half-length / F4 full-height full-length dual-slot)
Video CodecHardware codec units (video/image)
Launch2019
Mass Production/AvailabilityIn mass production

⚠️ Specification notes: The memory type is DDR4 (ECC); some early sources loosely recorded it as "LPDDR/onboard memory". Follow the Cambricon official product pages (MLU270-S4/F4): 16GB DDR4 ECC / 102 GB/s. FP16/FP32 peak compute is not published on the official site; only mixed-precision support is confirmed.

Key Features​

  • MLUv02 architecture: based on a network-on-chip (NoC) that guarantees parallel efficiency across the chip's 16 tensor cores; hardware on-chip data compression improves effective cache capacity and bandwidth
  • High energy-efficiency inference: INT8 inference performance improved 4× over the first generation, offering roughly 40× the energy efficiency of a CPU
  • Rich precision support: INT4/INT8/INT16 + FP16/FP32 mixed precision
  • Video/vision optimization: integrated ample hardware video/image codec units, suited to video analytics and smart cities
  • Unified edge-cloud software: supports Cambricon NeuWare / MagicMind, compatible with mainstream frameworks such as TensorFlow, PyTorch, Caffe, and MXNet

Vendor Information​

ItemDetails
CompanyCambricon Technologies Corporation Limited
HeadquartersBeijing
Founded2016
IPOSTAR Market 688256

Use Cases​

  • ✅ Cloud AI inference (vision, speech, NLP, recommendation)
  • ✅ Video analytics / smart cities (hardware codec units)
  • ✅ Edge/non-data-center inference (F4 active cooling, deployable in workstations)
  • ✅ Traditional machine learning acceleration
  • ❌ Large-scale model training (positioned for inference; compute and memory constrained)
  • ❌ Strong CUDA ecosystem dependence (requires migration to Cambricon NeuWare)

References​