Skip to main content

Kunlunxin P800 (2024)

Product Overview​

The Kunlunxin P800 is the third-generation AI accelerator card from Kunlunxin Technology (a Baidu company). Based on the in-house XPU-P architecture, it delivers 345 TFLOPS of peak FP16 compute (surpassing the NVIDIA H20's 148 TFLOPS) at a TDP of about 400W, in the OAM module form factor; it launched in March 2024. It supports running the full-strength DeepSeek-V3/R1 671B on 8 cards in a single server, and multiple 10,000-card clusters have been delivered.

Key Positioning:

  • Kunlunxin Gen 1 (2018): 14nm, deployed inside Baidu
  • Kunlunxin Gen 2 (2021): 7nm, in-house Kunlun Core II, 256 INT8 TOPS
  • Kunlunxin P800 (2024): XPU-P architecture, 345 TFLOPS FP16, OAM — this page
  • Kunlunxin M100 (early 2026): inference-dedicated — existing page
  • Kunlunxin M300 (early 2027): ultra-large-scale multimodal training

Core Specifications​

ParameterValue
ArchitectureIn-house XPU-P architecture
Process7nm
Memory96 GB HBM3
Memory Bandwidth2.4 TB/s
FP16345 TFLOPS (surpasses the H20's 148 TFLOPS)
Low-power mode128 TFLOPS @ 120 W
INT8820 TOPS (some reports cite 690–820 TOPS)
MoE SupportNative support for MoE architectures
TDP400 W
Form FactorOAM module
InterconnectXCCL (Kunlunxin interconnect), supports IB/ROCE
Release2024-03
Mass ProductionLaunched in March 2024, large-scale delivery since 2025
Cluster ScaleSupports 10,000-card clusters; an all-in-house 30,000-card cluster has been realized
SupernodeTianchi 256 / Tianchi 512
Supply StatusIn short supply, constrained by foundry capacity

Large-Model Adaptation​

ModelDeploymentNotes
DeepSeek-V3/R1 671B8 cards in a single server inferencePassed CAICT adaptation certification
DeepSeek MoE full-parameter training32 servers sufficeSupports MLA and multi-expert parallelism
ERNIE seriesNative Baidu Cloud supportMain deployment on Baidu AI Cloud
Llama / Qwen / ChatGLMSupportedIncludes MoE distilled versions
BaichuanSupportedDomestic model ecosystem

CUDA compatibility: models that run on CUDA migrate to the P800 at low cost; open-source inference frameworks such as vLLM are supported.

Vendor Information​

ParameterDetails
CompanyKunlunxin Technology (Beijing) Co., Ltd.
Parent CompanyBaidu (57.67% stake)
FoundedApril 2021 (spun off from Baidu)
P800 LaunchMarch 2024
IPO StatusStarted STAR Market IPO tutoring in May 2026
ValuationOver 10 billion RMB
Key CustomersBaidu AI Cloud, China Mobile (won the AI inference server centralized procurement)
CertificationCAICT five-star rating for "Stable Operation of Intelligent Computing Service Clusters"

Use Cases​

  • ✅ Domestic large-model training (full-parameter training of DeepSeek, ERNIE, etc.)
  • ✅ Large-model inference (671B on 8 cards in a single server)
  • ✅ Baidu AI Cloud (core compute foundation of the Baige platform)
  • ✅ Domestic intelligent computing centers (10,000-card clusters verified)
  • ✅ MoE model inference (native hardware optimization)
  • ❌ Deep CUDA ecosystem dependence (migration requires adaptation)
  • ❌ Low-power edge deployment (400W TDP is high)
  • ❌ International markets (restricted by export controls)

Key Timeline​

DateEvent
2018Kunlunxin Gen 1 released (14nm)
2021-04Kunlunxin Technology began independent operations
2021Kunlunxin Gen 2 mass-produced (7nm Kunlun Core II)
2024-03P800 officially launched (this page)
2025-02Passed DeepSeek 671B adaptation certification
2025Large-scale delivery of 10,000-card clusters
2026-05Started STAR Market IPO