Skip to main content

Enflame T21 (DTU 2.0)

Product Overview​

T21 is an OAM training module from Enflame Technology based on its second-generation cloud training chip DTU 2.0, released in July 2021. DTU 2.0 is the largest AI computing chip in China to date (3306mm²), built on GlobalFoundries' 12nm FinFET process, supporting TF32 precision (a first in China) and equipped with 64GB HBM2E memory (also a first in China).

Product positioning: a high-performance AI training accelerator card for large-scale model training scenarios.


Core Specifications​

ParameterValue
ArchitectureGCU-CARA (second-generation full-domain computing architecture)
Chip CodenameDTU 2.0
ProcessGlobalFoundries 12nm FinFET
Chip Area3306 mm² (57.5mm × 57.5mm)
Packaging2.5D CoWoS (ASE), 9-chip integration
FP3240 TFLOPS
TF32160 TFLOPS (first AI chip in China to support TF32)
FP16 / BF16134.4 TFLOPS
INT8320 TOPS
Memory64 GB HBM2E (4 Samsung HBM2E stacks)
Memory Bandwidth1.8 TB/s (maximum)
InterconnectGCU-LARE, 300 GB/s bidirectional
TDP~300W (estimated, based on the T20 accelerator card)
InterfaceOAM module (T21) / PCIe 4.0 (T20)
ReleaseJuly 2021
Mass Production2021 Q4

Data notes:

  • ✅ FP32, TF32, FP16, INT8, memory, and bandwidth figures are official data (confirmed by Zhihu columns)
  • ⚠️ TDP is an estimated value; not officially published

Product Highlights​

1. China's Largest AI Computing Chip​

  • 3306mm² chip area: at the limit of ASE's 2.5D packaging capability
  • 9-chip integration: 1 main chip + 4 HBM2E + 4 auxiliary chips
  • 12nm process: GlobalFoundries 12nm FinFET (not 7nm/5nm, but achieves high compute through a large die area)

2. First Chinese Chip to Support TF32 Precision​

  • TF32 (TensorFloat-32): a single-precision tensor format balancing FP32 numerical stability with FP16 compute efficiency
  • Full precision support: FP32, TF32, FP16, BF16, INT8
  • First in China: the first Chinese AI chip to support TF32 precision

3. Massive Memory Bandwidth​

  • 64GB HBM2E: first AI chip in China to support HBM2E
  • 1.8 TB/s bandwidth: massive throughput supporting large-scale model training
  • 4 Samsung HBM2E stacks: 4 HBM2E memories placed around the central main chip

4. High-Speed Full-Domain Interconnect​

  • GCU-LARE: full-domain interconnect technology developed specifically for AI training clusters
  • 300 GB/s bidirectional bandwidth: supports interconnecting thousands of accelerator cards
  • Excellent linear scaling: large cluster training performance approaches linear scaling

Software Stack: TopsRider​

TopsRider is Enflame's proprietary computing and programming platform.

ComponentFunctionCounterpart
Deep learning frameworksPyTorch, TensorFlow, PaddlePaddle adaptationMainstream frameworks
Distributed trainingSupports the Horovod distributed training frameworkHorovod
Operator generalizationBased on operator generalization technology and graph optimization strategies-
Programming modelOpen, upgradable programming modelCUDA
Operator interfacesExtensible operator interfacescuBLAS, cuDNN

Enflame Cluster: CloudBlazer Matrix 2.0​

Enflame and its partners (Inspur, etc.) jointly built the CloudBlazer Matrix 2.0 intelligent computing cluster:

  • 8192 Enflame training cards: forming a hyperscale intelligent computing cluster
  • 1.3 EFLOPS (1300 PFLOPS): single-precision AI compute
  • Exascale-class computing: quintillion-scale computing capability

Comparison with Competing Flagships (Official Benchmark)​

Enflame COO Zhang Yalin presented benchmark comparisons of the T20 against the NVIDIA V100 and A100 at the launch event:

MetricEnflame T20 (PCIe)Enflame T21 (OAM)NVIDIA A100NVIDIA V100
FP3233.6 TFLOPS40 TFLOPS19.5 TFLOPS?
TF32134.4 TFLOPS160 TFLOPS?-
FP16134.4 TFLOPS134.4 TFLOPS312 TFLOPS?
INT8268.8 TOPS320 TOPS624 TOPS?

Note: the Enflame T20/T21 beat the A100 at FP32 and TF32 precision, but still trail at FP16 and INT8.


Roadmap​

ChipProcessFP32TF32MemoryRelease
DTU 1.012nm20 TFLOPS-32GB HBM22019
DTU 2.012nm40 TFLOPS160 TFLOPS64GB HBM2E2021
DTU 3.0 (planned)7nm/5nm??HBM32024?


Enflame IPO (September 2026 Update)​

Enflame has officially listed on the Shanghai Stock Exchange STAR Market, becoming the last of China's GPU "Four Little Dragons" to complete capitalization — Enflame, Moore Threads, MetaX, and Biren have now all gone public.

ItemDetails
Listing DateSeptember 2026 (subscription 9/2, listed 9/11)
First-Day PerformanceSurged about 188%
Use of ProceedsDedicated to R&D and industrialization of the fifth-generation cloud AI chip
Fifth-Generation PlanOfficial release planned for 2027, supporting FP4 precision and cross-modal generation capability
Industry SignificanceThe capitalization phase for Chinese GPUs comes to a close; competition enters its "second half"

Key insight: the 188% first-day gain reflects the market pricing in the certainty of domestic substitution, not the strength of the current products. Enflame's disclosed capacity and customer contracts remain limited; the real test is whether the fifth-generation product lands on schedule.

References​

  • Zhihu: "Enflame Releases China's Largest AI Computing Chip, with INT8 Compute Reaching 320 TOPS" (2021-07-12)
  • EEWorld: "'China's Largest' Single AI Chip DTU 2.0 Released" (2021-07-08)
  • Enflame official launch event (WAIC 2021)

Last updated: July 3, 2026