Skip to main content

Corerain CAISA (Dataflow AI Chip)

Product Overview​

Corerain was founded by a team with Imperial College London/Tsinghua backgrounds and focuses on high-performance AI acceleration, with its core asset being a custom dataflow computing architecture. On April 9, 2019, Corerain released CAISA2.0, the world's first general-purpose AI underlying architecture built on dataflow technology, and on June 23, 2019 launched CAISA (CAISA3.0 engine), the world's first dataflow AI chip, at a Shenzhen event; it has completed mass production.

Unlike mainstream "instruction set architecture" AI chips, CAISA adopts a dataflow architecture: the order of data movement controls the order of computation, with the compute stream and data stream running overlapped, eliminating idle compute units and raising chip utilization above 90% (measured at 95.4%). Third-party data show that accelerator cards based on CAISA achieve about 3× measured performance using only about 1/3 of the peak compute of comparable NVIDIA products — a typical Chinese "dataflow" innovative-architecture route.

Around CAISA, Corerain launched the Xingkong X3 (single chip, 10.9 TOPS, PCIe 3.0 ×8) and Xingkong X9 (4 CAISA chips, 43.6 TOPS, PCIe 3.0 ×16, 64 video streams) accelerator cards, plus the RainBuilder end-to-end automatic compilation toolchain supporting seamless deployment of TensorFlow/PyTorch/Caffe/ONNX frameworks, deployed in security, power, industrial, and other scenarios.

Core Specifications​

ParameterValue
ArchitectureCustom dataflow CAISA3.0 architecture (4 CAISA engines, 16,000+ MAC units)
Process Node28nm
FP16 / BF16 ComputeNot disclosed
INT8 Compute10.9 TOPS (single-chip peak; Xingkong X9 with four chips: 43.6 TOPS)
FP32 ComputeNot disclosed
Memory CapacityNot disclosed
Memory TypeDDR (dual channel)
Memory Bandwidth>340 Gbps per CAISA chip (dual DDR channels)
TDPNot disclosed (chip level); Xingkong X3 dynamic power about 20W
InterconnectNot disclosed
InterfacePCIe 3.0 ×4 (CAISA chip); X3 card PCIe 3.0 ×8, X9 card PCIe 3.0 ×16
Launch2019-06
Mass Production/Availability2019-06 (in mass production)

Note: The "about 40 TOPS class" mentioned in some materials corresponds to the Xingkong X9 accelerator card (4 CAISA chips, 43.6 TOPS peak), while a single CAISA chip peaks at 10.9 TOPS. FP16/BF16, FP32, and memory capacity are not listed separately by the vendor, hence "Not disclosed". Chip utilization of up to 95.4% is CAISA's biggest highlight.

Key Features​

  • Dataflow architecture paradigm: abandons the traditional instruction set architecture, using dataflow to eliminate idle compute units; chip utilization up to 95.4% (typically <30% for peers).
  • High measured compute cost-effectiveness: about 3× the measured performance of comparable NVIDIA products with only 1/3 the peak compute (ResNet-50, YOLO v3, etc.).
  • General CNN support: supports mainstream algorithms such as ResNet, YOLO, and DeepLab through operator configurations in the dataflow network.
  • RainBuilder toolchain: end-to-end automatic compilation; mainstream framework models deploy without rewriting, lowering migration barriers.
  • Low latency: Xingkong X9 latency within 3 milliseconds, supporting 64-stream video structuring.

Vendor Information​

ItemDetails
CompanyShenzhen Corerain Technologies Co., Ltd.
HeadquartersShenzhen, China
Founded2017 (Founder/CEO Niu Xinyu; chief architect from academician Wayne Luk's team)

Use Cases​

  • ✅ Vision inference acceleration: video structuring for security, work safety, transportation, and smart manufacturing
  • ✅ Cost-effective edge/data center compute: trading low peak compute for high measured performance
  • ✅ Multi-stream video analytics: X9 supports real-time parsing of 64 1080P streams
  • ❌ Large model training (positioned for AI inference; training capability not emphasized)
  • ❌ Extreme floating-point scientific computing (FP16/FP32 figures undisclosed; focused on INT8 inference)
  • TsingMicro TX81 — Chinese reconfigurable/innovative architecture comparison
  • Vastai VA10 — Chinese cloud inference accelerator card
  • Lightmatter Envise — International compute-in-memory/novel architecture comparison

References​