Skip to main content

Biren BR100

Product Overview​

Biren BR100 is Biren Technology's first flagship general-purpose GPU, officially unveiled at Hot Chips 34 in August 2022. Built on the proprietary "Biliren" architecture with TSMC 7nm process + 2.5D CoWoS packaging, it integrates about 77 billion transistors and 64GB HBM2e memory in a dual-chiplet design, making it the most powerful Chinese general-purpose GPU of its time. The BR100 delivers 1024 TFLOPS BF16, 2048 TOPS INT8, and 256 TFLOPS FP32 peak compute — on paper surpassing the NVIDIA A100 and approaching the H100 on some metrics.

The BR100 ships in OAM (OCP Accelerator Module) form factor, paired with Biren's proprietary BIRENSUPA software stack (CUDA-like) and BLink inter-chip interconnect, targeting large model training and inference. Due to subsequent geopolitical factors and TSMC foundry restrictions, mass production and commercial deployment of the BR100 were significantly affected.

Core Specifications​

ParameterValue
ArchitectureBiliren (Biren proprietary ISA), dual chiplet
Process NodeTSMC 7nm, 2.5D CoWoS packaging
Transistor Count~77 billion
BF16 Compute1024 TFLOPS
TF32+ Compute512 TFLOPS
INT8 Compute2048 TOPS
FP32 Compute256 TFLOPS
FP64 ComputeNot supported
Memory Capacity64 GB
Memory TypeHBM2e
Memory Bus Width4096 bit
Memory Bandwidth~2.3 TB/s (some sources cite 1.64 TB/s)
TDP550 W
InterconnectBLink™ (8 ports, 512 GB/s aggregated)
InterfaceOAM; PCIe 5.0 ×16, CXL 2.0 support
Video Codec64-stream HEVC/H.264 encode / 512-stream decode
Launch2022-08 (Hot Chips 34)
Mass Production/AvailabilityReleased; mass production affected by foundry restrictions

⚠️ Specification notes: BF16 1024 / INT8 2048 / FP32 256 TFLOPS, 64GB HBM2e, 550W, and BLink 512 GB/s are Biren's official 2022 disclosures (consistent across multiple industry research reports and WCCFtech/aiwiki). Memory bandwidth has two reported figures, 1.64 TB/s and 2.3 TB/s; the table uses ~2.3 TB/s with the discrepancy noted. FP64 is not supported. These are vendor-claimed values, not verified by large-scale independent third-party benchmarks.

Key Features​

  • Dual-chiplet design: two compute dies + 896 GB/s die-to-die interconnect, breaking past reticle size limits
  • Six Biliren features: TF32+, TDA tensor data access accelerator, C-Warp CUDA-like Warp scheduling, BLink interconnect, HBM unified addressing, and security virtualization
  • BIRENSUPA software stack: CUDA-like software ecosystem, lowering migration cost
  • BLink interconnect: 8-port high-bandwidth inter-chip interconnect, supporting multi-card training

Vendor Information​

ItemDetails
CompanyBiren Technology
Founded2019-09
IPO2025-01, HKEX
HeadquartersShanghai
SoftwareBIRENSUPA (CUDA-like software stack)

Use Cases​

  • ✅ Large model training (high BF16/INT8 compute, 10,000-card cluster exploration)
  • ✅ Large model inference (high throughput)
  • ✅ High-end Chinese GPU alternative (on-paper benchmark against A100/H100)
  • ❌ Double-precision scientific computing (no FP64 support)
  • ❌ Native CUDA ecosystem (requires migration to BIRENSUPA)
  • ❌ International markets (export controls and foundry restrictions)

References​