NVIDIA GH200 Grace Hopper
Product Overview
The NVIDIA GH200 Grace Hopper superchip fuses the 72-core Arm Grace CPU and one Hopper H200 GPU into a single package via the 900 GB/s NVLink-C2C coherence bus, providing a unified CPU-GPU memory address space. Whole-chip TDP is 1000 W (CPU + GPU + HBM), roughly 5x lower power than traditional x86+H100 solutions with 7x the bandwidth.
Strategic position: The GH200 is NVIDIA's key product in the move from a "GPU accelerator card" to a "CPU+GPU fused superchip", designed for memory-intensive workloads (trillion-parameter LLMs, large-table recommendation systems, vector databases). The DGX GH200 combines 256 GH200s into a single giant system with 1 EFLOPS and 144 TB of unified memory.
Core Specifications
| Parameter | Value |
|---|---|
| Architecture | Grace (Arm Neoverse V2) + Hopper |
| Process | TSMC 4nm |
| CPU | 72-core NVIDIA Grace (Arm), 480 GB LPDDR5X |
| GPU | 1× Hopper H200, 96 GB HBM3 (expandable to 144 GB HBM3e) |
| FP8 Compute | 3,958 TFLOPS |
| FP16 Compute | 1,979 TFLOPS |
| FP32 | 134 TFLOPS |
| INT8 Compute | 3,958 TOPS |
| Memory | 96 GB HBM3 (high-end: 144 GB HBM3e) |
| Memory Type | HBM3 |
| Memory Bandwidth | 4.8 TB/s (96GB version) / 6.35 TB/s (144GB version) |
| CPU Memory Bandwidth | Up to 500 GB/s (LPDDR5X) |
| Interconnect | NVLink-C2C 900 GB/s (CPU↔GPU) |
| TDP | 1000 W (whole chip, incl. CPU+GPU+HBM) |
| Release | Announced 2023, mass production 2024 |
Key Features
- Unified memory: The Grace CPU and Hopper GPU share a 96-144 GB HBM3 address space; CUDA Unified Memory needs no manual copies
- NVLink-C2C 900 GB/s: 14x PCIe 5.0 x16
- Native liquid cooling: the 1000W single node fits a 1U liquid-cooled chassis
- Mature ecosystem: full CUDA / cuDNN / NCCL / TensorRT / OpenACC stack; PyTorch / JAX / TensorFlow ship native Grace binaries
Vendor Information
| Parameter | Value |
|---|---|
| Company | NVIDIA Corporation |
| Website | https://www.nvidia.com |
| Release | 2023 |
| Architecture | Grace Hopper |
Use Cases
- ✅ Cloud trillion-parameter training (DGX GH200 clusters)
- ✅ Generative AI inference farms (a single card loads a 175B FP16 model)
- ✅ HPC unified-memory acceleration (OpenFOAM / WRF / GROMACS)
- ✅ Edge AI hyper-converged servers
Related Comparisons
- NVIDIA H200 - Standalone Hopper GPU
- NVIDIA GB200 - Blackwell-generation rack-level superchip
- NVIDIA H100 - Hopper flagship accelerator card