Skip to main content

NVIDIA GH200 Grace Hopper

Product Overview​

The NVIDIA GH200 Grace Hopper superchip fuses the 72-core Arm Grace CPU and one Hopper H200 GPU into a single package via the 900 GB/s NVLink-C2C coherence bus, providing a unified CPU-GPU memory address space. Whole-chip TDP is 1000 W (CPU + GPU + HBM), roughly 5x lower power than traditional x86+H100 solutions with 7x the bandwidth.

Strategic position: The GH200 is NVIDIA's key product in the move from a "GPU accelerator card" to a "CPU+GPU fused superchip", designed for memory-intensive workloads (trillion-parameter LLMs, large-table recommendation systems, vector databases). The DGX GH200 combines 256 GH200s into a single giant system with 1 EFLOPS and 144 TB of unified memory.

Core Specifications​

ParameterValue
ArchitectureGrace (Arm Neoverse V2) + Hopper
ProcessTSMC 4nm
CPU72-core NVIDIA Grace (Arm), 480 GB LPDDR5X
GPU1× Hopper H200, 96 GB HBM3 (expandable to 144 GB HBM3e)
FP8 Compute3,958 TFLOPS
FP16 Compute1,979 TFLOPS
FP32134 TFLOPS
INT8 Compute3,958 TOPS
Memory96 GB HBM3 (high-end: 144 GB HBM3e)
Memory TypeHBM3
Memory Bandwidth4.8 TB/s (96GB version) / 6.35 TB/s (144GB version)
CPU Memory BandwidthUp to 500 GB/s (LPDDR5X)
InterconnectNVLink-C2C 900 GB/s (CPU↔GPU)
TDP1000 W (whole chip, incl. CPU+GPU+HBM)
ReleaseAnnounced 2023, mass production 2024

Key Features​

  • Unified memory: The Grace CPU and Hopper GPU share a 96-144 GB HBM3 address space; CUDA Unified Memory needs no manual copies
  • NVLink-C2C 900 GB/s: 14x PCIe 5.0 x16
  • Native liquid cooling: the 1000W single node fits a 1U liquid-cooled chassis
  • Mature ecosystem: full CUDA / cuDNN / NCCL / TensorRT / OpenACC stack; PyTorch / JAX / TensorFlow ship native Grace binaries

Vendor Information​

ParameterValue
CompanyNVIDIA Corporation
Websitehttps://www.nvidia.com
Release2023
ArchitectureGrace Hopper

Use Cases​

  • ✅ Cloud trillion-parameter training (DGX GH200 clusters)
  • ✅ Generative AI inference farms (a single card loads a 175B FP16 model)
  • ✅ HPC unified-memory acceleration (OpenFOAM / WRF / GROMACS)
  • ✅ Edge AI hyper-converged servers