Skip to main content

Enflame YunSui T20 (2021)

Product Overview​

The YunSui T20 is the second-generation AI training accelerator card released by Enflame Technology on July 7, 2021 at the World Artificial Intelligence Conference (WAIC). Based on the in-house Suiyuan 2.0 (DTU 2.0) chip, it adopts 2.5D advanced packaging (57.5mm × 57.5mm, integrating 9 chips), delivers 160 TFLOPS of TF32 compute (the first domestic support for TF32) and 320 TOPS of INT8 compute, and is equipped with 32GB HBM2E memory (1.6 TB/s bandwidth), with a TDP of 300W. Its GCU-LARE interconnect technology supports scaling clusters to 8192 cards (1.3 EFLOPS).

Enflame is one of the "GPU Four Little Dragons", focusing on domestic cloud AI training and inference.

Product Evolution:

  • DTU 1.0 / YunSui T10 (2019): first generation, 12nm, FP32 20 TFLOPS
  • DTU 2.0 / YunSui T20/T21 (2021): 2.5D packaging, TF32 160 TFLOPS — this page
  • DTU 3.0 / YunSui T30 (planned): next generation

Core Specifications​

DTU 2.0 Chip​

ItemParameter
ArchitectureIn-house GCU-CARA full-domain compute architecture
ProcessNot disclosed (industry estimate: 12nm)
Packaging2.5D advanced packaging, integrating 9 chips
Package Size57.5mm × 57.5mm (China's largest compute chip at launch)
FP3240 TFLOPS
TF32160 TFLOPS (first domestic support)
FP16 / BF16Supported (specific figures not disclosed)
INT8320 TOPS
Memory32GB HBM2E (Samsung; first domestic card to support it)
Memory Bandwidth1.6 TB/s
InterconnectGCU-LARE® (Enflame intelligent interconnect), 300 GB/s bidirectional

📌 Data correction (2026-09 cross-validation): This page previously recorded "64GB HBM2E / 1.8 TB/s", an early misstatement; Enflame's IPO prospectus figures and official materials specify 32GB HBM2E, 1.6 TB/s, TDP 300W, which has now been corrected.

YunSui T20 Accelerator Card​

ItemParameter
Core ChipDTU 2.0
PositioningData center AI training accelerator card
Form FactorPCIe training accelerator card
Multi-Card InterconnectIn-server 4-card full interconnect / enhanced 8-card full interconnect
ClusterSupports scaling from single-server multi-card to thousand-card level
Software StackTopsRider 2.0
Development InterfacesC++ / Python, multi-level open APIs
TDP300 W
ReleaseJuly 7, 2021 (WAIC 2021)

GCU-LARE Interconnect and Clusters​

SpecificationParameter
Interconnect TechnologyGCU-LARE® full-domain interconnect
Inter-Chip Bandwidth300 GB/s bidirectional
In-Server Interconnect4-card full interconnect → enhanced 8-card full interconnect
Cluster SolutionEnflame intelligent computing cluster CloudBlazer Matrix 2.0
Maximum Cluster8192 YunSui training cards
Total Cluster ComputeUp to 1.3 EFLOPS (FP32)
CoolingLiquid cooling, PUE < 1.5
Rack SolutionHigh-density deployment in a single rack

Record at launch: Enflame's COO stated: "No one in the world has yet achieved over 1E of single-precision compute using 8,000 cards."

TopsRider 2.0 Software Stack​

LayerToolDescription
PlatformTopsRider 2.0Enflame's unified programming platform
AI FrameworksPyTorchNative support
TensorFlowSupported
PaddlePaddleBaidu PaddlePaddle
Development InterfacesC++ / PythonMulti-level APIs
Operator LibraryIn-house operator libraryCovers mainstream models
CompilerGCU-CARA toolchainAutomated optimization
PerformanceTF32 precision on average 2.5x that of competitors' sub-flagship cardsOn par with competitors' flagships across many model types

Vendor Information​

ItemDetails
CompanyShanghai Enflame Technology Co., Ltd.
FoundedMarch 2018
FoundersZhao Lidong (former AMD China executive), Zhang Yalin (COO)
T20 LaunchJuly 7, 2021 (WAIC 2021)
FundingSeveral billion RMB in total (Tencent, Sequoia, etc.)
PositioningDomestic cloud AI training/inference chips
EcosystemOne of the "GPU Four Little Dragons" (MetaX, Biren, Enflame, Moore Threads)
PartnershipsCollaborates with partners to build the Enflame intelligent computing cluster

Use Cases​

  • ✅ Domestic large-model AI training (8192-card cluster, 1.3 EFLOPS)
  • ✅ Data center training (efficient training at 160 TFLOPS TF32)
  • ✅ Thousand-card cluster deployment (mature GCU-LARE interconnect solution)
  • ✅ Broad model coverage (multi-precision and dynamic feature support)
  • ✅ Mandatory domestic compute (independent and controllable)
  • ❌ Single-card inference (not the primary positioning; the YunSui i20 inference card is a better fit)
  • ❌ CUDA ecosystem (in-house TopsRider; migration requires adaptation)
  • ❌ FP8 training (not supported; watch the L600/T30)

Comparison with Contemporary Domestic AI Training Cards (2021)​

MetricEnflame T20Cambricon MLU370-X8Huawei Ascend 910Difference
Release2021-072021-Q4 (X8 mass-produced in 2022)2019T20 launched mid-year
Packaging2.5D advanced packagingStandard packagingStandard packagingT20 is more advanced
TF32160 TFLOPSNot supportedNot supportedUnique to T20
FP3240 TFLOPS24 TFLOPS256 TFLOPSAscend 910 leads
INT8320 TOPS256 TOPS512 TOPSAscend 910 leads
Memory32GB HBM2E48GB LPDDR532GB HBM2MLU370 has the largest capacity
Bandwidth1.6 TB/s614 GB/s1.2 TB/sT20 is the largest
Interconnect300 GB/s200 GB/sHCCST20 leads
Cluster8192 cards, 1.3 EFLOPSThousand cards4096 cardsT20 is the largest

2021's leading domestic AI training card: the T20 has the largest bandwidth (1.6 TB/s), the strongest interconnect (300 GB/s), and the largest cluster (8192 cards, 1.3 EFLOPS). However, its FP32 compute (40 TFLOPS) and INT8 (320 TOPS) are lower than the Huawei Ascend 910's.

Key Timeline​

DateEvent
2018-03Enflame Technology founded
2019-12DTU 1.0 / YunSui T10 released (12nm)
2021-07-07DTU 2.0 / YunSui T20 released (WAIC)
2021-2023T20/T21 deployed at scale; Enflame intelligent computing cluster commercialized
PlannedDTU 3.0 / YunSui T30