Skip to main content

NVIDIA A800

Product Overview​

NVIDIA A800 is the A100 China-specific slowed-down variant launched by NVIDIA in 2022 in response to US export controls on China. The two share virtually identical core compute; the only difference is that NVLink interconnect bandwidth is reduced from 600 GB/s to 400 GB/s. For single-card training/inference, performance is exactly the same as the A100; efficiency only drops slightly in interconnect-intensive large-scale cluster (16+ card) training.

Strategic position: After the A100 could no longer enter China through official channels, the A800 became one of the main workhorses for domestic LLM training with long-term stable supply, and was the mainstream choice in China's intelligent computing centers during 2022-2023 (later gradually replaced by the H800, Ascend 910B, and others).

Core Specifications​

ParameterValue
ArchitectureNVIDIA Ampere (GA100)
ProcessTSMC 7nm
CUDA Cores6,912
Tensor Cores432 (3rd Gen)
FP16 Compute312 TFLOPS
FP3219.5 TFLOPS
INT8 Compute624 TOPS
Memory80 GB HBM2e
Memory TypeHBM2e
Memory Bandwidth2.0 TB/s
NVLink Bandwidth400 GB/s (A100: 600 GB/s)
PCIeGen 4.0 x16
TDP400 W
Release2022
Priceapprox. ¥104,000 - ¥140,000 (China market)

Comparison with A100​

MetricA800 80GBA100 80GBDifference
ArchitectureAmpereAmpereIdentical
Memory80GB HBM2e80GB HBM2eIdentical
Bandwidth2.0 TB/s2.0 TB/sIdentical
FP16312 TFLOPS312 TFLOPSIdentical
NVLink400 GB/s600 GB/sSlowed 33%
PCIeGen 4.0Gen 4.0Identical
TDP400 W400 WIdentical

Key Features​

  • Compute identical to the A100: zero difference for single-card training/inference
  • Slowed NVLink: only affects interconnect-intensive large-scale cluster scenarios (8-card efficiency drops 10-15%, 16+ cards drop 20-30%)
  • Compliant supply: normally procurable in China, long-term stable supply
  • Complete ecosystem: full CUDA / cuDNN / NCCL stack compatibility

Vendor Information​

ParameterValue
CompanyNVIDIA Corporation
Websitehttps://www.nvidia.com
Release2022
ArchitectureAmpere

Use Cases​

  • ✅ LLM training (single card / small clusters)
  • ✅ AI inference
  • ✅ HPC scientific computing
  • ⚠️ Very large clusters (slowed NVLink hurts interconnect efficiency)