Iluvatar CoreX BI-V150
Product Overview
Iluvatar CoreX BI-V150 is a general-purpose GPU accelerator card from Iluvatar CoreX for the cloud training market, released in 2023 and in mass production in 2024. Based on Iluvatar CoreX's self-developed ivcore11 general-purpose GPU architecture, built on a 7nm process with 2.5D CoWoS packaging technology, it aims to provide Chinese compute solutions for AI training, high-performance computing, and other scenarios.
Product positioning: a Chinese general-purpose GPU training card, CUDA-ecosystem compatible, supporting FP32, FP16, and INT8 multi-precision computing.
Core Specifications
| Parameter | Value |
|---|---|
| Architecture | ivcore11 (second-generation general-purpose GPU architecture) |
| Process | TSMC 7nm |
| Packaging | 2.5D CoWoS |
| FP32 | 48 TFLOPS |
| FP16 | 192 TFLOPS (confirmed by official/procurement specifications) |
| INT8 | 384 TOPS (confirmed by official/procurement specifications) |
| Memory | 64 GB HBM2e |
| Memory Bandwidth | 1.2 TB/s (HBM2e, confirmed by official/procurement specifications) |
| TDP | 350 W |
| Interface | PCIe 4.0 x16 |
| Release | 2023 |
| Mass Production | 2024 |
| Form Factor | Air-cooled PCIe accelerator card / OAM module |
Data notes:
- ✅ FP32, memory, TDP, FP16, INT8, and memory bandwidth have all been confirmed by official/procurement specifications, including Nankai University's 2025 accelerator card procurement announcement
- Memory bandwidth 1.2 TB/s (HBM2e), FP16 192 TFLOPS, INT8 384 TOPS
Product Highlights
1. Fully Self-Developed Architecture
- ivcore11 architecture: Iluvatar CoreX's second-generation general-purpose GPU architecture, with a complete instruction set system
- CUDA-ecosystem compatible: supports CUDA C++ programming, low migration cost
- Multi-precision support: FP32, FP16, INT8, FP8 (requires the ixTE library)
2. Large Memory Capacity
- 64GB HBM2e: supports large-scale model training
- High bandwidth: 1.2 TB/s memory bandwidth (HBM2e, officially confirmed)
3. IXUCA Software Stack
- Mainstream framework compatibility: TensorFlow, PyTorch, PaddlePaddle
- Complete toolchain: compiler, math libraries, communication libraries, management tools
- Seamless migration: highly compatible with the CUDA ecosystem, migration time reduced by over 50%
Software Stack: IXUCA
IXUCA (Iluvatar Unified Computing Architecture) is the unified computing architecture software stack self-developed by Iluvatar CoreX.
| Component | Name | Function | Counterpart |
|---|---|---|---|
| Deep learning frameworks | PyTorch-Cambricon, TensorFlow-Cambricon | Adapted deep learning frameworks | PyTorch, TensorFlow |
| Inference framework | IGIE | High-performance inference framework | TensorRT |
| Inference engine | IxRT | Dedicated inference acceleration engine | TensorRT |
| LLM inference framework | IxFormer | Large-model inference and training optimization | vLLM |
| Compiler | IXUCA Compiler | Compiler | nvcc |
| Math libraries | ixDNN, ixBLAS | Fundamental deep learning operators | cuDNN, cuBLAS |
| Communication library | ixCCL | Multi-card communication library | NCCL |
| Management tool | ixsmi | GPU management tool | nvidia-smi |
Use Cases
- ✅ AI model training (CNN, RNN, Transformer, etc.)
- ✅ High-performance computing (HPC)
- ✅ Large-model pretraining (requires multi-card parallelism)
- ✅ Domestic substitution projects (government, state-owned enterprises, defense industry)
- ❌ Top-tier frontier model training (compute limitations)
- ❌ International markets (constrained by US export controls)
Performance Comparison
| Metric | BI-V150 | A100 80GB | Gap |
|---|---|---|---|
| FP32 | 48 TFLOPS | 19.5 TFLOPS | +146% |
| Memory | 64 GB | 80 GB | -20% |
| TDP | 350W | 400W | -12.5% |
Note: the BI-V150's FP32 compute is higher than the A100's (possibly due to different precision definitions or test conditions), but actual training performance also depends on the degree of software stack optimization.
Vendor Information
| Item | Details |
|---|---|
| Company | Shanghai Iluvatar CoreX Semiconductor Co., Ltd. |
| English Name | Iluvatar CoreX |
| Founded | 2015 |
| Founder | Diao Shijing |
| Headquarters | Shanghai |
| Positioning | Chinese general-purpose GPU chip design company |
| Official Website | https://www.iluvatar.com |
| Software Stack | https://support.iluvatar.com |
Market Progress: ByteDance Shipments Double (September 2026 Update)
According to a Reuters report on 2026-09-10, Iluvatar CoreX's GPU shipments to ByteDance this year have nearly doubled to about 100,000 units, and the company has reallocated GPUs originally reserved for internal use to prioritize this order.
In ByteDance's domestic chip supplier hierarchy, Huawei remains the largest supplier, followed by Cambricon, with Iluvatar CoreX third.
During the same period, Iluvatar CoreX, like other domestic vendors, raised its prices (+20~30%), mainly due to rising HBM costs and tight domestic HBM capacity.
Key insight: a shipment scale of 100,000 units shows Iluvatar CoreX has moved from "validation purchases" into the "volume supply" stage. However, note that these orders are mainly for inference; the company's software ecosystem maturity on the training side still lags the leading vendors.
Related Products
- Zhikai 100 (MR-V100) - Inference GPU (page to be created)
- BI-V100 - First-generation training GPU (page to be created)
- Huawei Ascend 910C - Strongest Chinese AI training chip
- Cambricon MLU690 - Chinese AI training chip
- Full comparison table
Information To Be Added
- Official FP16/INT8 compute figures (192 TFLOPS / 384 TOPS, confirmed by procurement specifications)
- Official memory bandwidth figure (1.2 TB/s, HBM2e)
- Actual training performance tests (ResNet, BERT, LLM, etc.)
- Multi-card scaling performance (ixCCL)
- Energy efficiency tests
Data sources:
- SMZDM unboxing review (2026-01-08)
- Gitee AI product documentation
- Iluvatar CoreX official materials
Last updated: 2026-06-28