Skip to main content

Hygon DCU K100 AI (2024)

Product Overview​

The Hygon DCU K100 AI (DCU Gen 3) is a high-performance GPGPU accelerator card launched by Hygon Information for AI data centers. Based on the in-house x86-compatible GPGPU architecture, it delivers 192 TFLOPS of FP16/BF16 compute and 392 TOPS of INT8 compute, equipped with 64GB HBM2e memory and 896 GB/s of memory bandwidth. Compatible with the ROCm/DTK software stack, it can greatly reduce CUDA migration costs and is purpose-built for domestic large-model training and inference. (Standard K100: FP64 24.5 TFLOPS, peak of about 100 TFLOPS, 64GB HBM2e, 896 GB/s, about 300W, retaining double-precision capability.)

Product Evolution:

  • DCU Gen 1 (2022): early GPGPU, DCU architecture validation
  • DCU Gen 2 (2023): double-precision K100 + AI-optimized edition
  • DCU K100 AI (2024): 192 TFLOPS FP16, x86 instruction set — this page
  • DCU Gen 3 (planned): next-generation GPGPU

Core Specifications​

ParameterValue
ArchitectureIn-house GPGPU, x86 instruction-set compatible
ProcessAdvanced node (estimated 7nm; not officially disclosed)
FP3249 TFLOPS
TF3296 TFLOPS
FP16 / BF16192 TFLOPS
INT8392 TOPS
Memory64 GB HBM2e
Memory Bandwidth896 GB/s (HBM2e, dual ring buses, measured utilization 92%+)
Bus TopologyDual ring HBM2e buses (read/write separation to avoid conflicts)
SchedulerUnified tensor scheduler that dynamically senses Attention QKV matrices
TDP350 W (K100-AI edition; standard K100 about 300W)
Form FactorPCIe full-height, full-length, dual-slot card
Release2024
Software EcosystemDTK (DCU Toolkit), based on ROCm, CUDA compatible

DTK Software Ecosystem​

LayerToolDescription
RuntimeROCmAMD's open-source GPGPU platform
Programming FrameworkDTK (DCU Toolkit)Hygon's in-house stack, HIP/CUDA compatible
AI FrameworksPyTorch (HIP backend)Automatically mapped through ROCm
TensorFlowSupported
PaddlePaddleBaidu PaddlePaddle
CompilerHIPIFYAutomatic CUDA code conversion tool
Operator LibraryMIOpencuDNN-like
QuantizationFP16/INT8 mixed precision supportedNative BF16 format

CUDA compatibility: through the DTK/HIP ecosystem, CUDA code can be automatically converted into DCU-executable code, with far lower migration cost than fully in-house architectures.

Vendor Information​

ParameterDetails
CompanyHygon Information Technology Co., Ltd.
Stock Code688041 (STAR Market)
Technology Originx86 licensing + in-house DCU architecture
K100 AI Launch2024
Key CustomersThe three major telecom operators, intelligent computing centers, state-owned enterprises in finance/energy
Benchmark ProductNVIDIA H20 (FP16 192 vs H20 148 TFLOPS)
Price AdvantageConsiderably cheaper than the H20

Key Technical Features​

  • Dual ring HBM2e buses: physically separated read/write paths, measured utilization steady at 92%+ (about 76% for same-generation competitor cards); excellent performance on training workloads such as ResNet-50
  • Unified tensor scheduler: dynamically senses QKV matrix size changes in Attention layers, eliminating scheduling jitter
  • x86-compatible ecosystem: underlying instruction set compatible with x86, lowering software development migration costs
  • Native BF16: hardware support for the Brain Floating Point format
  • Qwen-7B fine-tuning test: when batch size jumps from 4 to 8, the utilization curve shows almost no sharp rise (the A100, by contrast, exhibits clear scheduling jitter)

Use Cases​

  • ✅ Domestic intelligent computing centers (x86 ecosystem compatibility; preferred by SOEs/operators)
  • ✅ Large-model training (domestic models such as the Qwen series and Baichuan)
  • ✅ Large-model inference (192 TFLOPS FP16 inference services)
  • ✅ Computer vision training (ResNet-50, YOLOv8)
  • ✅ Scientific computing (x86 ecosystem + large-scale linear algebra, PDE solving)
  • ❌ Native CUDA ecosystem (requires HIP translation; some operators need manual optimization)
  • ❌ Extra-large model training (limited by 64GB memory; requires multi-card parallelism)

Comparison with NVIDIA H20​

MetricHygon DCU K100 AINVIDIA H20Difference
FP16192 TFLOPS148 TFLOPSDCU K100 +30%
INT8392 TOPS296 TOPSDCU K100 +32%
Memory64GB HBM2e96GB HBM3H20 1.5x
Software EcosystemDTK (ROCm) / HIPCUDAH20 more mature
PriceLowerHigherDCU K100 has the advantage
SupplyStable domestic supplyExport control riskDCU K100 is secure

DCU K100 strengths: compute surpasses the H20, lower price, secure supply; weaknesses: smaller memory, software ecosystem maturity behind CUDA.

Domestic GPU Ecosystem Comparison​

ProductArchitectureFP16 (TFLOPS)MemorySoftware EcosystemStrength
Hygon DCU K100GPGPU/x8619264GB HBM2eDTK (ROCm)x86 compatible
Cambricon MLU 590In-house MLUv0512896GB HBM2NeuWareMature domestic AI
Kunlunxin P800XPU-P345Not disclosedIn-houseStrongest compute
MetaX C600XCORE 1.5~300 (FP8:1000)144GB HBM3eMXMACALargest memory
Enflame T20GCU-CARA~80 (TF32:160)64GB HBM2ETopsRiderCluster solution

Key Timeline​

DateEvent
2016Hygon Information founded (AMD x86/Zen licensing)
2022DCU Gen 1 released
2023DCU Gen 2 double-precision K100 released
2024DCU K100 AI launched (DCU Gen 3 AI edition)
2025Large-scale deployment of the K100 AI