Skip to main content

NVIDIA Rubin

The NVIDIA Rubin GPU was officially unveiled at CES 2026 (January 2026) as NVIDIA's next-generation AI accelerator chip following Blackwell. It adopts an all-new six-chip architecture (Vera CPU + Rubin GPU + NVLink 6 + BlueField-4 DPU + ConnectX-9 SuperNIC + Spectrum-6 Ethernet switch), forming a complete, unified AI system.

Core Specifications​

ParameterSpecification
ArchitectureRubin architecture (Blackwell successor)
ProcessTSMC 3nm (N3/N3P)
Transistor Count336 billion
Chip DesignDual compute chiplet + dual I/O chiplet MCM design (NV-HBI inter-die interconnect)
Memory Capacity288 GB HBM4
Memory Bandwidth22 TB/s
FP4 Compute50 PFLOPS (inference figure)
Compute Units224 SM / 896 Tensor Cores (officially disclosed at Hot Chips 2026)
AnnouncementJanuary 2026 (CES 2026 preview) / full production at GTC 2026
TDP2300 W
Mass Production2026 H2 full production ramp (officially confirmed on 8/21 that shipments to hyperscale customers have started, with Microsoft the first to deploy)

Memory Specifications​

ParameterSpecification
Memory TypeHBM4
Memory Capacity288GB
Memory Bandwidth22 TB/s
Per-Stack Bandwidth>3.0 TB/s
Per-Pin Data Rate>11 Gbps
HBM4 Stack Count8 stacks
Interface Width2048-bit/stack (doubled vs HBM3e)

Compute Performance​

PrecisionPerformance
FP4 (inference)50 PFLOPS
FP4 (training)35 PFLOPS
FP8~25 PFLOPS (dense, inference; NVFP4 is 2x that)
FP16TBD
FP32TBD
INT8TBD

Power and Cooling​

ParameterSpecification
Single GPU Power1800W - 2300W
Cooling Solution100% liquid cooling (no air-cooled configuration)
NVL72 Rack Power120-130 kW
NVL144 CPX Rack Power~260 kW
NVL576 Rack Power~600 kW

Interconnect Technology​

ParameterSpecification
NVLink VersionNVLink 6
Single GPU NVLink Bandwidth3.6 TB/s (2x Blackwell's NVLink 5)
CPU-GPU InterconnectNVLink-C2C, 1.8 TB/s bandwidth
NetworkConnectX-9 SuperNIC + Spectrum-6 Ethernet switch

Architecture Features​

1. Third-Generation Transformer Engine​

  • Supports NVFP4 adaptive compression
  • Automatically optimizes precision formats without manual tuning
  • Balances performance and accuracy

2. Simultaneous Multithreading (SMT)​

  • A single GPU supports 176 threads
  • Improves parallel processing capability

3. Full-Stack AI Factory Solution​

The Rubin platform consists of 6 co-designed chips:

  • Vera CPU: control plane
  • Rubin GPU: compute core
  • NVLink 6: high-speed GPU-to-GPU interconnect
  • BlueField-4 DPU: data center infrastructure processing
  • ConnectX-9 SuperNIC: network connectivity
  • Spectrum-6 Ethernet switch: network switching

All components are co-designed, eliminating the bottlenecks of multi-vendor deployments.

4. Dynamo Inference Scheduling Framework​

  • Supports splitting inference tasks
  • Prefill stage: assigned to Vera CPU / Rubin NVL144 CPX racks
  • Decode stage: assigned to Rubin GPUs
  • Greatly improves inference energy efficiency

5. Heterogeneous Collaboration with the Groq 3 LPU​

  • Interconnected via Spectrum X networking
  • Offloads decode tasks of trillion-parameter models to LPUs without any CUDA code changes
  • Inference task scheduling is handled automatically by the Dynamo layer

System Configurations​

Rubin NVL72​

  • GPU Count: 72 Rubin GPUs
  • Rack Power: 120-130 kW
  • Use Case: large-scale training and inference

Rubin NVL144 CPX​

  • GPU Count: 144 Rubin GPUs
  • Rack Power: ~260 kW
  • Use Case: ultra-large-scale inference

Rubin NVL576​

  • GPU Count: 576 Rubin GPUs
  • Rack Power: ~600 kW
  • Use Case: ultra-large-scale AI factories

Comparison with Blackwell​

ParameterBlackwell (B200)Rubin (Rubin GPU)
ProcessTSMC 4nmTSMC 3nm
Transistors208 billion336 billion
HBMHBM3e 192GBHBM4 288GB
HBM Bandwidth8 TB/s22 TB/s
FP4 Compute? PFLOPS50 PFLOPS
TDP1000W1800W-2300W
NVLinkNVLink 5 (1.8 TB/s)NVLink 6 (3.6 TB/s)

Release Timeline and Availability​

StageTimeline
AnnouncementJanuary 2026 (CES 2026 preview) / full production at GTC 2026
Sample Delivery2026 Q2 (shipments to hyperscale customers already started)
Volume Production2026 H2 production ramp (officially confirmed on 8/21)
Early Customer DeliverySecond half of 2026 (Microsoft the first to deploy)

Production and Shipment Cadence (2026-09 Update)​

StageTimeline / Status
ES engineering samples2026-05 (provided to top-tier customers)
PS production samplesFrom mid-September 2026, new compute boards / switch boards fully transition to PS versions
Rack-level mass productionFrom mid-September 2026, officially launched with more than ten major ODMs / OEMs worldwide
2026 shipmentsJust over 1 million units (new-product ramp cadence)
2027 shipmentsExpected to ramp sharply; the real volume year

📌 Figure note: the B300 / GB300 are expected to ship over 5 million units in 2026 and remain the year's shipment mainstay; Rubin's volume ramp only arrives in 2027.

Confirmed Deployment Partners​

AWS, Google Cloud, Microsoft Azure, Oracle Cloud, CoreWeave, Lambda, Nebius, and Nscale have all confirmed first-batch deployments. India's Yotta announced it will deploy 40,000 Vera Rubin units at Noida D4 (120MW) and 40,000 GB300 units at Navi Mumbai NM2 (80MW), a combined investment of about $12 billion — the first order of this scale from a non-US hyperscaler without sovereign-fund backing.

Measured Performance Figures (Hot Chips 2026)​

At Hot Chips 2026, NVIDIA shifted its narrative from specs to measured results: running the DeepSeek-V4-Pro AgentX benchmark, Vera Rubin NVL72 delivers up to 30x throughput per megawatt and up to 35x lower token cost versus the GB300 NVL72 (per NVIDIA's figures). The advantage is greatest in high-interactivity / low-latency scenarios.

⚠️ The 30x / 35x figures are NVIDIA marketing numbers; independent verification (InferenceX) is not yet complete — discount them before relying on them.

Architecture Roadmap Change: Kyber NVL144 Removed from the 2027 Roadmap​

The Kyber NVL144, originally slated for 2027 production, has been removed from the Rubin Ultra product roadmap, mainly because manufacturing difficulty with the orthogonal midplane caused severe schedule slippage. The 2027 mainstream shifts to NVL72 and NVL576 expanded from NVL72; single-node form factors such as NVL8 remain. The Kyber form factor itself has not been fully terminated; it may have to wait until after 2028, in combination with the Feynman architecture.

Vera CPU (Key Companion on the Same Platform)​

  • 88-core in-house cores (codename Olympus, ARM v9.2, not a stock Neoverse design)
  • 176 threads (SMT-X spatial multithreading), 164MB L3, single die
  • 1.2 TB/s SOCAMM2 memory bandwidth
  • Official figures: agentic tasks complete 1.8x faster than traditional x86, with 2x energy efficiency
  • Already in full mass production with Vera Rubin systems, marking NVIDIA's entry into the server CPU market

References​


Tags: GPU Training Inference NVIDIA Rubin 2026 HBM4 NVLink 6