Enflame T21 (DTU 2.0)
Product Overview
T21 is an OAM training module from Enflame Technology based on its second-generation cloud training chip DTU 2.0, released in July 2021. DTU 2.0 is the largest AI computing chip in China to date (3306mm²), built on GlobalFoundries' 12nm FinFET process, supporting TF32 precision (a first in China) and equipped with 64GB HBM2E memory (also a first in China).
Product positioning: a high-performance AI training accelerator card for large-scale model training scenarios.
Core Specifications
| Parameter | Value |
|---|---|
| Architecture | GCU-CARA (second-generation full-domain computing architecture) |
| Chip Codename | DTU 2.0 |
| Process | GlobalFoundries 12nm FinFET |
| Chip Area | 3306 mm² (57.5mm × 57.5mm) |
| Packaging | 2.5D CoWoS (ASE), 9-chip integration |
| FP32 | 40 TFLOPS |
| TF32 | 160 TFLOPS (first AI chip in China to support TF32) |
| FP16 / BF16 | 134.4 TFLOPS |
| INT8 | 320 TOPS |
| Memory | 64 GB HBM2E (4 Samsung HBM2E stacks) |
| Memory Bandwidth | 1.8 TB/s (maximum) |
| Interconnect | GCU-LARE, 300 GB/s bidirectional |
| TDP | ~300W (estimated, based on the T20 accelerator card) |
| Interface | OAM module (T21) / PCIe 4.0 (T20) |
| Release | July 2021 |
| Mass Production | 2021 Q4 |
Data notes:
- ✅ FP32, TF32, FP16, INT8, memory, and bandwidth figures are official data (confirmed by Zhihu columns)
- ⚠️ TDP is an estimated value; not officially published
Product Highlights
1. China's Largest AI Computing Chip
- 3306mm² chip area: at the limit of ASE's 2.5D packaging capability
- 9-chip integration: 1 main chip + 4 HBM2E + 4 auxiliary chips
- 12nm process: GlobalFoundries 12nm FinFET (not 7nm/5nm, but achieves high compute through a large die area)
2. First Chinese Chip to Support TF32 Precision
- TF32 (TensorFloat-32): a single-precision tensor format balancing FP32 numerical stability with FP16 compute efficiency
- Full precision support: FP32, TF32, FP16, BF16, INT8
- First in China: the first Chinese AI chip to support TF32 precision
3. Massive Memory Bandwidth
- 64GB HBM2E: first AI chip in China to support HBM2E
- 1.8 TB/s bandwidth: massive throughput supporting large-scale model training
- 4 Samsung HBM2E stacks: 4 HBM2E memories placed around the central main chip
4. High-Speed Full-Domain Interconnect
- GCU-LARE: full-domain interconnect technology developed specifically for AI training clusters
- 300 GB/s bidirectional bandwidth: supports interconnecting thousands of accelerator cards
- Excellent linear scaling: large cluster training performance approaches linear scaling
Software Stack: TopsRider
TopsRider is Enflame's proprietary computing and programming platform.
| Component | Function | Counterpart |
|---|---|---|
| Deep learning frameworks | PyTorch, TensorFlow, PaddlePaddle adaptation | Mainstream frameworks |
| Distributed training | Supports the Horovod distributed training framework | Horovod |
| Operator generalization | Based on operator generalization technology and graph optimization strategies | - |
| Programming model | Open, upgradable programming model | CUDA |
| Operator interfaces | Extensible operator interfaces | cuBLAS, cuDNN |
Enflame Cluster: CloudBlazer Matrix 2.0
Enflame and its partners (Inspur, etc.) jointly built the CloudBlazer Matrix 2.0 intelligent computing cluster:
- 8192 Enflame training cards: forming a hyperscale intelligent computing cluster
- 1.3 EFLOPS (1300 PFLOPS): single-precision AI compute
- Exascale-class computing: quintillion-scale computing capability
Comparison with Competing Flagships (Official Benchmark)
Enflame COO Zhang Yalin presented benchmark comparisons of the T20 against the NVIDIA V100 and A100 at the launch event:
| Metric | Enflame T20 (PCIe) | Enflame T21 (OAM) | NVIDIA A100 | NVIDIA V100 |
|---|---|---|---|---|
| FP32 | 33.6 TFLOPS | 40 TFLOPS | 19.5 TFLOPS | ? |
| TF32 | 134.4 TFLOPS | 160 TFLOPS | ? | - |
| FP16 | 134.4 TFLOPS | 134.4 TFLOPS | 312 TFLOPS | ? |
| INT8 | 268.8 TOPS | 320 TOPS | 624 TOPS | ? |
Note: the Enflame T20/T21 beat the A100 at FP32 and TF32 precision, but still trail at FP16 and INT8.
Roadmap
| Chip | Process | FP32 | TF32 | Memory | Release |
|---|---|---|---|---|---|
| DTU 1.0 | 12nm | 20 TFLOPS | - | 32GB HBM2 | 2019 |
| DTU 2.0 | 12nm | 40 TFLOPS | 160 TFLOPS | 64GB HBM2E | 2021 |
| DTU 3.0 (planned) | 7nm/5nm | ? | ? | HBM3 | 2024? |
Related Products
- Iluvatar CoreX BI-V150 - Chinese general-purpose GPU training card
- Huawei Ascend 910C - Strongest Chinese AI training chip
- Cambricon MLU690 - Chinese AI training chip
- Full comparison table
Enflame IPO (September 2026 Update)
Enflame has officially listed on the Shanghai Stock Exchange STAR Market, becoming the last of China's GPU "Four Little Dragons" to complete capitalization — Enflame, Moore Threads, MetaX, and Biren have now all gone public.
| Item | Details |
|---|---|
| Listing Date | September 2026 (subscription 9/2, listed 9/11) |
| First-Day Performance | Surged about 188% |
| Use of Proceeds | Dedicated to R&D and industrialization of the fifth-generation cloud AI chip |
| Fifth-Generation Plan | Official release planned for 2027, supporting FP4 precision and cross-modal generation capability |
| Industry Significance | The capitalization phase for Chinese GPUs comes to a close; competition enters its "second half" |
Key insight: the 188% first-day gain reflects the market pricing in the certainty of domestic substitution, not the strength of the current products. Enflame's disclosed capacity and customer contracts remain limited; the real test is whether the fifth-generation product lands on schedule.
References
- Zhihu: "Enflame Releases China's Largest AI Computing Chip, with INT8 Compute Reaching 320 TOPS" (2021-07-12)
- EEWorld: "'China's Largest' Single AI Chip DTU 2.0 Released" (2021-07-08)
- Enflame official launch event (WAIC 2021)
Last updated: July 3, 2026