Skip to main content

AI Compute Card Full Comparison Table (100+ Models)

Quick Filter​

ScenarioRecommended Models
Trillion-parameter training (GPT-4 class)Rubin R200, B300 Ultra, MI400, TPU Ironwood
10B–100B parameter trainingH100, H200, B200, MI300X, MI325X
China market (domestic alternatives)Ascend 950DT, Ascend 910C, Ascend 920, MLU690
High-throughput inferenceL40S, L4, H200 (inference mode), Crescent Island
Edge AIJetson Orin, Edge TPU, Hailo-8L

Datacenter Training GPU​

Compute unified with precision and sparsity/density labels. ✅ Available · 🔄 Upcoming · 🔮 Forward-looking

ModelFP8/FP4 ComputeFP16 ComputeMemoryMemory BandwidthTDPReleaseStatus
NVIDIA Rubin R20035 PFLOPS (FP4 training)~25 PFLOPS288GB HBM422 TB/s1,800-2,300W2026 H2🔮
NVIDIA B300 Ultra14 PFLOPS (sparse)~7 PFLOPS288GB HBM3e8 TB/s1,400W2026 Q1✅
NVIDIA GB300——288GB HBM3e8 TB/s1,600W2025 H2✅
NVIDIA B2009 PFLOPS (sparse)4.5 PFLOPS (sparse)192GB HBM3e8 TB/s1,000W2025 Q2✅
NVIDIA GB200——192GB HBM3e8 TB/s1,000W2024 Q4✅
NVIDIA B1007 PFLOPS (sparse)3.5 PFLOPS (sparse)192GB HBM3e8 TB/s700W2024 Q4✅
NVIDIA H2003,958 TFLOPS (sparse)1,979 TFLOPS141GB HBM3e4.8 TB/s700W2024 Q2✅
NVIDIA H100 SXM3,958 TFLOPS (sparse)1,979 TFLOPS80GB HBM33.35 TB/s700W2022 Q3✅
NVIDIA H20296 TFLOPS148 TFLOPS96GB HBM34.0 TB/s400W2024 Q1✅
NVIDIA A100—312 TFLOPS80GB HBM2e2.0 TB/s300-400W2020 Q3✅
AMD MI400 Series40 PFLOPS (FP4)~10 PFLOPS432GB HBM419.6 TB/s1,200-1,500W2026 H2🔮
AMD MI455X20 PFLOPS~10 PFLOPS432GB HBM419.6 TB/s1,200-1,500W2026 H2🔮
AMD MI350X10.1 PFLOPS (MXFP6)~5 PFLOPS288GB HBM3e8 TB/s1,400W2025 H2🔄
AMD MI325X2,614 TFLOPS1,307 TFLOPS256GB HBM3e6.48 TB/s750W2024 Q4✅
AMD MI300X2,614 TFLOPS1,307 TFLOPS192GB HBM35.3 TB/s750W2023 Q4✅
Huawei Ascend 950PR1 PFLOPS (FP8)~780 TFLOPS128GB HiBL~3 TB/s600W2026 H1🔄
Huawei Ascend 950DT1 PFLOPS (FP8)~500 TFLOPS144GB HiZQ4 TB/s500W2026 H1🔄
Huawei Ascend 920—1,800 TFLOPS~96GB HBM3~4 TB/s400W2025 H2✅
Huawei Ascend 910C780 TFLOPS (BF16)~390 TFLOPS128GB HBM2e (dual-die)~1.2 TB/s310W2025 H1✅
Huawei Ascend 910B—256 TFLOPS64GB HBM2e1.2 TB/s310W2023✅
Cambricon MLU690—700+ TFLOPS196GB HBM33.35 TB/s~500W2025✅
Moore Threads MTT S5000—~63 TFLOPS (FP32)80GB GDDR6X1.6 TB/s300W2025 Q1✅
MetaX C6001,000 TFLOPS——3.6 TB/s400W2025 Q4🔄
Kunlun P800—345 TFLOPS96GB HBM3—400W2024 Q1✅
Iluvatar BI-V150—~192 TFLOPS64GB HBM2e—350W2023✅

Domestic chip note: Huawei Ascend, Cambricon MLU, Moore Threads MTT, MetaX, Kunlun, and Iluvatar are Chinese domestic AI chip representatives, primarily targeting the China market due to US export controls. MTT S5000 is priced in CNY (¥55,000).

Datacenter Inference GPU​

ModelFP8 ComputeINT8 ComputeMemoryTDPUse CaseStatus
NVIDIA L40S733 TFLOPS (sparse)1,466 TOPS48GB GDDR6350WDatacenter inference✅
NVIDIA RTX 6000 Ada1,458 TFLOPS (sparse)2,905 TOPS48GB GDDR6300WWorkstation inference✅
NVIDIA RTX Pro 6000 Blackwell——96GB GDDR7 ECC600WWorkstation inference✅
NVIDIA L4485 TFLOPS970 TOPS24GB GDDR672WEdge inference✅
NVIDIA L296 TFLOPS (sparse)193 TOPS16GB GDDR650WLow-power inference✅
NVIDIA T465 TFLOPS130 TOPS16GB GDDR670WEntry-level inference✅
Intel Arc Pro B60——24GB GDDR6200WMid-range inference✅
Intel Arc Pro B50——16GB GDDR670WEntry-level inference✅
Qualcomm AI 200800 TFLOPS——280WDatacenter inference🔄

AI Training ASIC (TPU / Gaudi / Trainium)​

ModelVendorCompute (BF16/FP8)MemoryInterconnectReleaseStatus
Google TPU Ironwood (v7)Google~2,000 TFLOPS192GB HBM~5 Tb/s2026 H1🔄
Google TPU v6pGoogle—96GB HBM2—2024 Q4✅
Google TPU v6e (Trillium)Google918 TFLOPS32GB HBM1.6 Tb/s2024 Q4✅
Google TPU v5pGoogle———2023 Q4✅
Google TPU v5eGoogle—16GB HBM2—2023 Q3✅
Google TPU v4Google—32GB HBM2—2020 Q3✅
Google TPU 8t (Training)Google———2026 Q2🔮
Google TPU 8i (Inference)Google~1,500 TOPS——2026 Q2🔮
Intel Gaudi 3Intel1,600 TFLOPS128GB SRAM2.4 Tb/s2024 Q2✅
Intel Gaudi 2Intel865 TFLOPS (FP8)96GB HBM2e2.4 Tb/s2022 Q2✅
Intel Gaudi 4Intel—192GB HBM3e—2026 Q2🔮
Intel Crescent IslandIntelTBD480GB LPDDR5xTBD2026 H2🔄
AWS Trainium 3AWS~5.7 PFLOPS~144GB~4.5 Tb/s2025 Q4🔄
AWS Trainium 2AWS1,299 TFLOPS (dense)64GB~1.6 Tb/s2024 Q4✅
AWS Trainium 1AWS191 TFLOPS (FP8)32GB HBM—2020 Q4✅
AWS Inferentia 2AWS190 TFLOPS (FP16)32GB HBM2e—2022 Q4✅
AWS Inferentia 1AWS———2019 Q4✅
Microsoft Maia 200Microsoft5+ PFLOPS——2026 Q1🔄
Meta MTIA v3Meta———2026 Q3🔮

Wafer-Scale Training​

ModelVendorTransistorsOn-Chip MemoryFP8 ComputeReleaseStatus
Cerebras WSE-4Cerebras~5-6 trillion44GB SRAM~400 PFLOPS2027🔮
Cerebras WSE-3Cerebras4 trillion40GB SRAM125 PFLOPS2024 Q1✅
Cerebras WSE-2Cerebras2.6 trillion40GB SRAM85 PFLOPS2021 Q3✅

Edge AI & On-Device NPU​

ModelVendorCompute (TOPS)PowerUse CaseStatus
NVIDIA Jetson ThorNVIDIA2,070 TOPS130WRobotics / autonomous driving✅
NVIDIA Jetson Orin AGXNVIDIA275 TOPS60WEdge inference✅
Qualcomm AI 100Qualcomm70 TOPS15WDatacenter edge inference✅
Huawei Ascend 310Huawei22 TOPS8WOn-device inference✅
Hailo-8LHailo13 TOPS1.5WOn-device vision AI✅
Google Edge TPUGoogle4 TOPS2WIoT on-device inference✅

Innovative Architectures​

ModelArchitecture TypeKey FeatureVendorStatus
Groq LPU v2LPU (Language Processing Unit)Ultra-low latency inference (~500 tok/s)Groq✅
Graphcore IPU (Bow)IPU (Intelligence Processing Unit)Native graph computing, 1,400 IPU coresGraphcore✅
Tesla Dojo (D1)Distributed training waferIntegrated auto-labeling + model trainingTesla✅
Apple M5 UltraSoC + NPUOn-device 50 TOPS, unified memoryApple🔮
BrainChip Akida 2Spiking Neural Network (SNN)Ultra-low-power neuromorphicBrainChip✅

Pricing Reference​

Prices fluctuate with market supply and demand. Purchase prices are affected by export controls. CNY denotes Chinese domestic chip pricing in RMB (reference rate: 1 USD ≈ 7.2 CNY). Data for reference only.

NVIDIA​

ModelMSRP (USD)Market Price (USD)Notes
Rubin R200$85,000—Estimated
GB300$75,000$72,000
GB200$65,000$62,000
B300 Ultra$55,000$52,000
B200$45,000$42,000
B100$38,000$35,000
H200$40,000$38,000
H100$30,000$28,000
H100 NVL$40,000$38,000
H20$14,000$13,000China-specific
A100$15,000$28,000Secondary market
L40S$7,000$6,500
RTX Pro 6000 Blackwell$6,800$6,500
RTX 6000 Ada$4,500$4,200
RTX 5090$2,000$1,900
L4 / L2$2,500$2,300
T4$2,500$1,800Secondary market
Jetson Thor$800—Module
Jetson Orin$400$380Module
H800¥280,000¥250,000CNY, China market

AMD​

ModelMSRP (USD)Market Price (USD)Notes
MI400$55,000—Estimated
MI350X$40,000$37,000
MI355X$22,000$20,500
MI325X$18,000$16,500
MI300X$15,000$13,500
MI250$12,000$10,000Secondary market
MI210$9,000$8,500

Intel​

ModelMSRP (USD)Market Price (USD)Notes
Gaudi 4$25,000—Estimated
Gaudi 3$18,000$16,500
Gaudi 2$12,000$11,000
Gaudi 1$8,000$7,000Discontinued
Max Series$4,000$3,700
Flex Series$1,000$900
Arc Pro B60$500$480
Arc Pro B50$350$330

Huawei Ascend​

ModelMSRP (USD)Market Price (USD)Notes
Ascend 950DT$22,000—Estimated
Ascend 950PR$18,000—Estimated
Ascend 910D$18,000$16,000
Ascend 920$25,000$23,000
Ascend 910C$16,000$14,500
Ascend 910B$12,000$10,500

Google TPU / AWS / Cloud​

ModelMSRP (USD)Market Price (USD)Notes
TPU Ironwood$45,000—Estimated
TPU v6p$40,000—
TPU v5p$35,000—
TPU v6e$22,000—
TPU v5e$18,000—
TPU v4$25,000—
Trainium 3$30,000—
Trainium 2$22,000—
Inferentia 2$12,000—
Trainium 1$15,000—
Inferentia 1$8,000—

Chinese Domestic Chips (CNY Pricing)​

ModelMSRP (CNY)Market Price (CNY)USD Equivalent
Cambricon MLU690¥150,000¥140,000~$19,444
Moore Threads MTT S5000¥55,000¥50,000~$6,944
Iluvatar TG150¥85,000¥78,000~$10,833
Enflame T21¥70,000¥65,000~$9,028
MetaX C500¥55,000¥52,000~$7,222
Moore Threads S4000¥60,000¥55,000~$7,639

Other Vendors​

ModelMSRP (USD)Market Price (USD)Notes
Cerebras WSE-4$8,000,000—Rack system
Cerebras WSE-3$5,000,000—Rack system
Cerebras WSE-2$3,000,000—Rack system
SambaNova SN40L$200,000—System
Groq LPU v2$35,000—
Rubin Ultra$150,000—Estimated
Graphcore IPU$15,000—
Tenstorrent Blackhole$20,000—
Qualcomm AI 100$8,000$7,000
Hailo-8L$300$280Module

Purchasing Guide​

By Model Scale​

By Region​

  • North America / Europe: NVIDIA + AMD, freely available
  • China: Ascend 950 / 910C / 920 / Cambricon MLU690 (domestic alternatives)
  • Cloud (no hardware preference): Any vendor, choose by price

← Back to Home | Roadmap → | TCO Calculator → | Industry News →