NVIDIA Rubin
The NVIDIA Rubin GPU was officially unveiled at CES 2026 (January 2026) as NVIDIA's next-generation AI accelerator chip following Blackwell. It adopts an all-new six-chip architecture (Vera CPU + Rubin GPU + NVLink 6 + BlueField-4 DPU + ConnectX-9 SuperNIC + Spectrum-6 Ethernet switch), forming a complete, unified AI system.
Core Specifications
| Parameter | Specification |
|---|---|
| Architecture | Rubin architecture (Blackwell successor) |
| Process | TSMC 3nm (N3/N3P) |
| Transistor Count | 336 billion |
| Chip Design | Dual compute chiplet + dual I/O chiplet MCM design (NV-HBI inter-die interconnect) |
| Memory Capacity | 288 GB HBM4 |
| Memory Bandwidth | 22 TB/s |
| FP4 Compute | 50 PFLOPS (inference figure) |
| Compute Units | 224 SM / 896 Tensor Cores (officially disclosed at Hot Chips 2026) |
| Announcement | January 2026 (CES 2026 preview) / full production at GTC 2026 |
| TDP | 2300 W |
| Mass Production | 2026 H2 full production ramp (officially confirmed on 8/21 that shipments to hyperscale customers have started, with Microsoft the first to deploy) |
Memory Specifications
| Parameter | Specification |
|---|---|
| Memory Type | HBM4 |
| Memory Capacity | 288GB |
| Memory Bandwidth | 22 TB/s |
| Per-Stack Bandwidth | >3.0 TB/s |
| Per-Pin Data Rate | >11 Gbps |
| HBM4 Stack Count | 8 stacks |
| Interface Width | 2048-bit/stack (doubled vs HBM3e) |
Compute Performance
| Precision | Performance |
|---|---|
| FP4 (inference) | 50 PFLOPS |
| FP4 (training) | 35 PFLOPS |
| FP8 | ~25 PFLOPS (dense, inference; NVFP4 is 2x that) |
| FP16 | TBD |
| FP32 | TBD |
| INT8 | TBD |
Power and Cooling
| Parameter | Specification |
|---|---|
| Single GPU Power | 1800W - 2300W |
| Cooling Solution | 100% liquid cooling (no air-cooled configuration) |
| NVL72 Rack Power | 120-130 kW |
| NVL144 CPX Rack Power | ~260 kW |
| NVL576 Rack Power | ~600 kW |
Interconnect Technology
| Parameter | Specification |
|---|---|
| NVLink Version | NVLink 6 |
| Single GPU NVLink Bandwidth | 3.6 TB/s (2x Blackwell's NVLink 5) |
| CPU-GPU Interconnect | NVLink-C2C, 1.8 TB/s bandwidth |
| Network | ConnectX-9 SuperNIC + Spectrum-6 Ethernet switch |
Architecture Features
1. Third-Generation Transformer Engine
- Supports NVFP4 adaptive compression
- Automatically optimizes precision formats without manual tuning
- Balances performance and accuracy
2. Simultaneous Multithreading (SMT)
- A single GPU supports 176 threads
- Improves parallel processing capability
3. Full-Stack AI Factory Solution
The Rubin platform consists of 6 co-designed chips:
- Vera CPU: control plane
- Rubin GPU: compute core
- NVLink 6: high-speed GPU-to-GPU interconnect
- BlueField-4 DPU: data center infrastructure processing
- ConnectX-9 SuperNIC: network connectivity
- Spectrum-6 Ethernet switch: network switching
All components are co-designed, eliminating the bottlenecks of multi-vendor deployments.
4. Dynamo Inference Scheduling Framework
- Supports splitting inference tasks
- Prefill stage: assigned to Vera CPU / Rubin NVL144 CPX racks
- Decode stage: assigned to Rubin GPUs
- Greatly improves inference energy efficiency
5. Heterogeneous Collaboration with the Groq 3 LPU
- Interconnected via Spectrum X networking
- Offloads decode tasks of trillion-parameter models to LPUs without any CUDA code changes
- Inference task scheduling is handled automatically by the Dynamo layer
System Configurations
Rubin NVL72
- GPU Count: 72 Rubin GPUs
- Rack Power: 120-130 kW
- Use Case: large-scale training and inference
Rubin NVL144 CPX
- GPU Count: 144 Rubin GPUs
- Rack Power: ~260 kW
- Use Case: ultra-large-scale inference
Rubin NVL576
- GPU Count: 576 Rubin GPUs
- Rack Power: ~600 kW
- Use Case: ultra-large-scale AI factories
Comparison with Blackwell
| Parameter | Blackwell (B200) | Rubin (Rubin GPU) |
|---|---|---|
| Process | TSMC 4nm | TSMC 3nm |
| Transistors | 208 billion | 336 billion |
| HBM | HBM3e 192GB | HBM4 288GB |
| HBM Bandwidth | 8 TB/s | 22 TB/s |
| FP4 Compute | ? PFLOPS | 50 PFLOPS |
| TDP | 1000W | 1800W-2300W |
| NVLink | NVLink 5 (1.8 TB/s) | NVLink 6 (3.6 TB/s) |
Release Timeline and Availability
| Stage | Timeline |
|---|---|
| Announcement | January 2026 (CES 2026 preview) / full production at GTC 2026 |
| Sample Delivery | 2026 Q2 (shipments to hyperscale customers already started) |
| Volume Production | 2026 H2 production ramp (officially confirmed on 8/21) |
| Early Customer Delivery | Second half of 2026 (Microsoft the first to deploy) |
Production and Shipment Cadence (2026-09 Update)
| Stage | Timeline / Status |
|---|---|
| ES engineering samples | 2026-05 (provided to top-tier customers) |
| PS production samples | From mid-September 2026, new compute boards / switch boards fully transition to PS versions |
| Rack-level mass production | From mid-September 2026, officially launched with more than ten major ODMs / OEMs worldwide |
| 2026 shipments | Just over 1 million units (new-product ramp cadence) |
| 2027 shipments | Expected to ramp sharply; the real volume year |
📌 Figure note: the B300 / GB300 are expected to ship over 5 million units in 2026 and remain the year's shipment mainstay; Rubin's volume ramp only arrives in 2027.
Confirmed Deployment Partners
AWS, Google Cloud, Microsoft Azure, Oracle Cloud, CoreWeave, Lambda, Nebius, and Nscale have all confirmed first-batch deployments. India's Yotta announced it will deploy 40,000 Vera Rubin units at Noida D4 (120MW) and 40,000 GB300 units at Navi Mumbai NM2 (80MW), a combined investment of about $12 billion — the first order of this scale from a non-US hyperscaler without sovereign-fund backing.
Measured Performance Figures (Hot Chips 2026)
At Hot Chips 2026, NVIDIA shifted its narrative from specs to measured results: running the DeepSeek-V4-Pro AgentX benchmark, Vera Rubin NVL72 delivers up to 30x throughput per megawatt and up to 35x lower token cost versus the GB300 NVL72 (per NVIDIA's figures). The advantage is greatest in high-interactivity / low-latency scenarios.
⚠️ The 30x / 35x figures are NVIDIA marketing numbers; independent verification (InferenceX) is not yet complete — discount them before relying on them.
Architecture Roadmap Change: Kyber NVL144 Removed from the 2027 Roadmap
The Kyber NVL144, originally slated for 2027 production, has been removed from the Rubin Ultra product roadmap, mainly because manufacturing difficulty with the orthogonal midplane caused severe schedule slippage. The 2027 mainstream shifts to NVL72 and NVL576 expanded from NVL72; single-node form factors such as NVL8 remain. The Kyber form factor itself has not been fully terminated; it may have to wait until after 2028, in combination with the Feynman architecture.
Vera CPU (Key Companion on the Same Platform)
- 88-core in-house cores (codename Olympus, ARM v9.2, not a stock Neoverse design)
- 176 threads (SMT-X spatial multithreading), 164MB L3, single die
- 1.2 TB/s SOCAMM2 memory bandwidth
- Official figures: agentic tasks complete 1.8x faster than traditional x86, with 2x energy efficiency
- Already in full mass production with Vera Rubin systems, marking NVIDIA's entry into the server CPU market
References
- NVIDIA Official Blog: Inside the NVIDIA Rubin Platform
- TechInsider: NVIDIA Rubin GPU Analysis
- Baidu Baike: Rubin GPU
Tags: GPU Training Inference NVIDIA Rubin 2026 HBM4 NVLink 6