Skip to main content

Zhonghao Xinying XuYu TPU (AI Training/Inference)

Product Overview​

Zhonghao Xinying is an emerging Chinese TPU-architecture AI chip startup. On June 30, 2026, it officially released the XuYu TPU (unified training and inference), becoming one of the few companies globally to master TPU architecture after Google. Per official launch figures: 896 TFLOPS mixed-precision floating point, 1792 TOPS INT8, TDP 600W, with up to 2048 chips connected via all-optical interconnect in a single supernode; the company claims about 50% lower power consumption than traditional chips at the same compute level. The Tianjin Mobile TPU AI Computing Center is already operational, marking the first benchmark case of domestic TPU commercialization.

Core design philosophy: Abandon GPU's graphics rendering modules; a pure ASIC design focused on AI computation (fully in-house IP, instruction set, and operator library), delivering significantly better energy efficiency than traditional GPUs at the same process node. The software stack is compatible with PyTorch / vLLM / SGLang / DeepSpeed / Megatron.

Core Specifications​

ItemParameter
Release2026-06-30 (XuYu official launch)
ArchitectureSelf-developed TPU (pure ASIC, no graphics rendering; in-house IP/instruction set/operator library)
ProcessNot disclosed
FP16/BF16 Compute896 TFLOPS (mixed-precision floating point, official launch figures)
INT8 Compute1792 TOPS
FP32 ComputeNot disclosed
TDP600 W
InterconnectAll-optical interconnect, up to 2048 chips per supernode
Software CompatibilityPyTorch / vLLM / SGLang / DeepSpeed / Megatron
Production StatusIn mass production and delivery
Unit PriceNot disclosed

📌 Data correction (2026-09 cross-validation): This page previously recorded "INT8 512 TOPS / FP16 256 TFLOPS (estimated) / TDP 400W / released 2026-05" based on early media reports; it has now been updated to the official XuYu launch figures of 2026-06-30: 896 TFLOPS mixed precision, 1792 TOPS INT8, 600W, up to 2048 chips all-optical interconnect per supernode.

Efficiency Comparison​

ChipPowerComputePositioning
Zhonghao Xinying XuYu TPU600 W896 TFLOPS mixed precision / 1792 TOPS INT8Official claim: ~50% lower power than traditional chips at the same compute level
NVIDIA H100700 W3959 TOPS INT8Baseline
Cambricon MLU590350 W512 TOPS INT8Domestic counterpart

ℹ️ Source of the efficiency advantage: Pure ASIC design with no graphics overhead; a dedicated Matrix Multiplication Unit (MXU) architecture analogous to Google's TPU delivers significantly lower power consumption and cooling costs than GPUs in inference scenarios.

Commercial Deployment​

ItemDetails
First CustomerTianjin Mobile
DeploymentTianjin Mobile TPU AI Computing Center
StatusOperational
Industry SignificanceAmong the first benchmarks of domestic TPU commercialization

Architecture Differences vs GPU​

DimensionZhonghao Xinying TPUTraditional GPU (e.g. H100)
Design PhilosophyPure AI ASICGeneral-purpose GPU (graphics+AI)
Energy EfficiencyHigh (no graphics overhead)Lower
Programming FlexibilityLower (fixed dataflow)High (CUDA general-purpose computing)
Ecosystem CompatibilitySelf-developed (no CUDA compatibility; compatible with mainstream framework interfaces)CUDA ecosystem
Use CasesAI inference + trainingGeneral-purpose computing

Use Cases​

  • ✅ AI inference (high-efficiency scenarios)
  • ✅ AI computing center construction (domestic compliance)
  • ✅ Training of hundred-billion-parameter large models (supernode-scale all-optical interconnect)
  • ✅ Low power / low cooling cost scenarios
  • ❌ Complex dataflow models (less flexible than GPU)
  • ❌ Graphics rendering / general-purpose computing

Manufacturer Info​

ItemContent
CompanyZhonghao Xinying (Hangzhou) Technology Co., Ltd.
PositioningEmerging domestic TPU-architecture AI chip player
Core ProductXuYu TPU (released 2026-06-30)
First CustomerTianjin Mobile
Launch DateJune 30, 2026
FundingMultiple rounds