Skip to main content

Cambricon MLU370-X8 (Siyuan 370)

Product Overview​

The Cambricon MLU370-X8 is Cambricon's unified training and inference AI accelerator card based on MLUarch03 (third-generation MLU architecture), in a dual-chip Siyuan 370 configuration with a 7nm process. Official peak performance: FP16 96 TFLOPS / BF16 96 TFLOPS / INT8 256 TOPS / FP32 24 TFLOPS, equipped with 48GB LPDDR5 (614.4 GB/s), full-height full-length dual-slot 250W, and aggregated MLU-Link 200 GB/s (bidirectional). It is paired with the NeuWare software stack + MagicMind. In Cambricon's product line it sits above the MLU290 and below the MLU590, and was one of Cambricon's mainstream shipping models before the Siyuan 590 (official commercial availability in 2026-07).

📌 Data correction (2026-09 cross-validation): This page previously misrecorded "INT8 96 TOPS / BF16 48 TFLOPS / TDP 35W / 48GB HBM2". Cambricon's official product page confirms the MLU370-X8: INT8 256 TOPS, INT16 128 TOPS, FP16/BF16 96 TFLOPS each, FP32 24 TFLOPS, 48GB LPDDR5, 614.4 GB/s, 250W; the Siyuan 370 series uses LPDDR5 rather than HBM.

Key lineage:

  • MLU 270 (2019): 16nm — early training/inference
  • MLU 290 (2020): 7nm, MLU-Link — first 7nm training generation
  • MLU370-X8 (released 2021 / mass-produced 2022): 7nm dual-chip, 48GB LPDDR5, INT8 256 TOPS, 250W — this page
  • MLU590 (official commercial availability 2026-07): 7nm, 96GB HBM2e — current mainstream
  • MLU690 (mass production in early 2026): dual-Die Chiplet, 196GB HBM3 — flagship

Core Specifications​

ItemParameter
ArchitectureCambricon MLUarch03 (third generation, dual-chip Siyuan 370)
ProcessTSMC 7nm
FP3224 TFLOPS
FP1696 TFLOPS
BF1696 TFLOPS
INT8256 TOPS
INT16128 TOPS
Memory48GB LPDDR5
Memory Bandwidth614.4 GB/s
TDP250 W
Form FactorPCIe Gen4 ×16, full-height full-length dual-slot (passive cooling)
InterconnectAggregated MLU-Link 200 GB/s (bidirectional, 3.1x PCIe 4.0)
Video Codec264-channel HEVC full-HD decode / 48-channel encode, up to 8K
Release2021 (Siyuan 370 launch)
Mass Production2022
Unit Price~¥38,000-42,000 (e-commerce channel reference)

vs MLU 290 (2020)​

MetricMLU370-X8MLU 290 (2020)Change
Process7nm7nmSame generation
Memory48GB LPDDR532GB HBM2+50% capacity, switched to LPDDR
Bandwidth614.4 GB/s307 GB/s2x
INT8256 TOPS64 TOPS4x
FP16/BF1696 TFLOPS32 TFLOPS3x
TDP250 W50 WFull-size data center card
Interconnect200 GB/s100 GB/s2x
SoftwareNeuWareNeuWare 0.5New generation

vs NVIDIA T4​

MetricMLU370-X8NVIDIA T4Difference
Process7nm12nmMLU370 newer generation
INT8256 TOPS130 TOPSMLU370 doubles it
FP1696 TFLOPS65 TFLOPSMLU370 +48%
BF1696 TFLOPSNot supportedExclusive to MLU370
TDP250W70WT4 is more power-efficient
Memory48GB LPDDR516GB GDDR6MLU370 3x
Bandwidth614.4 GB/s320 GB/sMLU370 1.9x
SoftwareNeuWare + MagicMindCUDAT4 is mature

MLU370-X8 positioning: In tests published on its basic software platform SDK, Cambricon officially claims that on a single card the performance of four common AI models is on par with a mainstream 350W RTX GPU; in multi-card scenarios, MLU-Link (200GB/s) delivers better parallel speedup.

Use Cases​

  • ✅ Unified training and inference (official positioning, full FP32/FP16/BF16/INT8/INT4 coverage)
  • ✅ High-density inference + multimodal concurrency (264-channel HEVC full-HD decode)
  • ✅ Government/SOE AI projects (domestic substitution policy)
  • ✅ Single-server 8-card training / distributed inference (MLU-Link 200GB/s)
  • ❌ FP8 (not supported; requires MLU590/690)
  • ❌ Cutting-edge large-model training (move to Siyuan 590/690 recommended)
  • ❌ International market (no CUDA compatibility)

LLM Inference Performance (48GB Version)​

ModelQuantizationPerformance (tok/s)Notes
LLaMA 1 7BFP16~25 tok/sMainstream
LLaMA 1 13BFP16~12 tok/sFull FP16
LLaMA 1 30BQ4_K_M~5 tok/sQuantized
ChatGLM-6BFP16~30 tok/sChinese
Stable Diffusion 1.5FP162x vs MLU 290Image generation

48GB LPDDR5 advantage: Compared to the contemporaneous NVIDIA T4 with 16GB, it can fully load 13B-class LLM FP16 inference; combined with 264-channel video decoding, it was the mainstream domestic multimodal/vision + mid-scale LLM inference workhorse of 2022-2024.

Software Stack NeuWare​

LayerToolDescription
AI FrameworksNeuWareUnified programming platform
PyTorch (NeuWare backend)Automatic MLU mapping
TensorFlow (NeuWare backend)Compatible
MindSporeHuawei/CAICT-led, PyTorch compatible
CompilerBANG C/C++Cambricon proprietary language
Operator LibraryCNMLCUDA cuDNN-like
Inference EngineMagicMindInference acceleration engine
QuantizationNeuQuantINT8 automatic
Model ZooModelZooCV/NLP/LLM

Vendor Information​

ItemDetails
CompanyCambricon Technologies
FoundersChen Tianshi and Chen Yunji brothers (CAS Institute of Computing Technology)
Founded2016-03
IPO2020-07-20 STAR Market (688256)
Siyuan 370 Launch2021-Q4 (MLU370-X8 mass-produced in 2022)
Key CustomersChina Mobile, Inspur, Sugon, ByteDance, Zhipu AI
National ProjectsRecommended chip for the "East Data West Compute" project

Key Timeline​

DateEvent
2016-03Cambricon founded (CAS ICT spinout)
2018-05First chip MLU 100 released (16nm)
2020-07-20STAR Market IPO (688256)
2020MLU 290 (first 7nm generation)
2021-Q4Siyuan 370 released (this page's product)
2022MLU370-X8 mass production + customer deployment
First unveiled 2024MLU 590 (official commercial availability 2026-07)
Early 2026MLU 690 mass production (196GB HBM3)

Cambricon Product Line​

ProductLaunchProcessMemoryINT8TDPStatus
MLU370-X82021-Q4 / mass-produced 20227nm dual-chip48GB LPDDR5256 TOPS250WIn mass production and on sale
MLU590Unveiled 2024 / commercial 2026-077nm96GB HBM2e512 TOPS350WCurrent mainstream
MLU690Mass production in early 2026Dual-Die Chiplet196GB HBM32400 TOPS (media-reported figures)500WFlagship

Key Features​

  • Unified training and inference: official positioning, full FP32/FP16/BF16/INT16/INT8/INT4 precision coverage
  • 48GB LPDDR5: large memory among domestic cards of its 2022 generation (vs NVIDIA T4 16GB)
  • MLU-Link 200GB/s: 3.1x PCIe 4.0, 8-card full interconnect in a single server
  • 264-channel video decoding: a powerful tool for multimodal/vision scenarios
  • Weaknesses: no FP8, LPDDR5 bandwidth weaker than HBM, cutting-edge large-model training requires moving to 590/690

vs Contemporary Domestic AI Chips (2021-2022)​

MetricCambricon MLU370-X8Huawei Ascend 310Alibaba Hanguang 800 (2021)
Process7nm12nm12nm
INT8256 TOPS22 TOPS820 TOPS
TDP250W8W168W
Memory48GB LPDDR58GB LPDDR432GB HBM2
Bandwidth614.4 GB/s25 GB/s700 GB/s
TargetTraining + InferenceEdgeData center inference

2021-2022 domestic AI top three: Hanguang 800 has the strongest compute (820 TOPS), MLU370-X8 has the most complete ecosystem (unified training/inference + MLU-Link), Ascend 310 has the best efficiency (8W).