Product Overview
NVIDIA A800 is the A100 China-specific slowed-down variant launched by NVIDIA in 2022 in response to US export controls on China. The two share virtually identical core compute; the only difference is that NVLink interconnect bandwidth is reduced from 600 GB/s to 400 GB/s. For single-card training/inference, performance is exactly the same as the A100; efficiency only drops slightly in interconnect-intensive large-scale cluster (16+ card) training.
Strategic position: After the A100 could no longer enter China through official channels, the A800 became one of the main workhorses for domestic LLM training with long-term stable supply, and was the mainstream choice in China's intelligent computing centers during 2022-2023 (later gradually replaced by the H800, Ascend 910B, and others).
Core Specifications
| Parameter | Value |
|---|
| Architecture | NVIDIA Ampere (GA100) |
| Process | TSMC 7nm |
| CUDA Cores | 6,912 |
| Tensor Cores | 432 (3rd Gen) |
| FP16 Compute | 312 TFLOPS |
| FP32 | 19.5 TFLOPS |
| INT8 Compute | 624 TOPS |
| Memory | 80 GB HBM2e |
| Memory Type | HBM2e |
| Memory Bandwidth | 2.0 TB/s |
| NVLink Bandwidth | 400 GB/s (A100: 600 GB/s) |
| PCIe | Gen 4.0 x16 |
| TDP | 400 W |
| Release | 2022 |
| Price | approx. ¥104,000 - ¥140,000 (China market) |
Comparison with A100
| Metric | A800 80GB | A100 80GB | Difference |
|---|
| Architecture | Ampere | Ampere | Identical |
| Memory | 80GB HBM2e | 80GB HBM2e | Identical |
| Bandwidth | 2.0 TB/s | 2.0 TB/s | Identical |
| FP16 | 312 TFLOPS | 312 TFLOPS | Identical |
| NVLink | 400 GB/s | 600 GB/s | Slowed 33% |
| PCIe | Gen 4.0 | Gen 4.0 | Identical |
| TDP | 400 W | 400 W | Identical |
Key Features
- Compute identical to the A100: zero difference for single-card training/inference
- Slowed NVLink: only affects interconnect-intensive large-scale cluster scenarios (8-card efficiency drops 10-15%, 16+ cards drop 20-30%)
- Compliant supply: normally procurable in China, long-term stable supply
- Complete ecosystem: full CUDA / cuDNN / NCCL stack compatibility
Use Cases
- ✅ LLM training (single card / small clusters)
- ✅ AI inference
- ✅ HPC scientific computing
- ⚠️ Very large clusters (slowed NVLink hurts interconnect efficiency)
External Links