Skip to main content

2 posts tagged with "AI Factory"

AI factory and datacenter infrastructure

View all tags

AWS Adds Another 2 Million NVIDIA GPUs: Vera CPU Debuts on AWS, 100,000 GPUs Reserved for the U.S. Government AI Factory

· 5 min read
Industry Research Team

This article is based on official announcements from AWS and NVIDIA (September 5, 2026) and public statements by executives of both companies.

On September 5, 2026, AWS and NVIDIA announced an expanded strategic partnership: on top of the "1 million additional GPUs starting in 2026" plan announced at GTC 2026, they will deploy 2 million more NVIDIA GPUs, covering three architecture generations — Blackwell Ultra, Rubin, and Rubin Ultra — with a deployment window of 2027-2028. Demand growth exceeding all previous forecasts was the direct reason both companies cited.

This is no longer a simple "chip purchase" — it is a full-stack partnership spanning GPUs, CPUs, interconnect, memory, open-source models, and software. We break the key information into six points.

1. Composition and Timeline of the 2 Million GPUs​

  • Scale: 2 million GPUs (added on top of the original 1 million GPU plan, tripling the total committed scale)
  • Architectures: Blackwell Ultra, Rubin, Rubin Ultra
  • Timeline: 2027-2028, deployed across AWS global infrastructure (including newly built AI factories)
  • Context: Amazon's 2026 capital expenditure guidance has been raised from $200 billion to $220 billion, and CEO Andy Jassy has explicitly said it is still not enough to meet AI compute demand

For reference, the figures Jensen Huang gave at GTC 2026 put cumulative orders and demand for the Blackwell and Rubin platforms through 2027 on track to reach $1 trillion (at GTC 2025, the estimate for 2026 was roughly $500 billion). AWS's add-on order is one of the heaviest puzzle pieces in that big picture.

2. Vera CPU Comes to AWS for the First Time​

This is the most structurally significant change in the partnership: NVIDIA Vera CPU infrastructure will enter AWS.

The Vera CPU is the general-purpose processor in the Vera Rubin platform designed for agentic AI, positioned to efficiently turn AI resources into "completed agent tasks." As agentic workloads rise, CPU-side pressure on task orchestration, memory management, and data scheduling rises in step — pure GPU expansion is no longer enough, and AWS needs a matching high-performance CPU layer. This also aligns with AWS's strategy of "offering the broadest compute choices, from in-house chips (Graviton/Trainium) to partner chips": Vera does not replace Trainium, but fills in the CPU compute layer alongside the accelerated infrastructure.

At re:Invent 2025, AWS announced that its next-generation Trainium chip would support NVIDIA NVLink Fusion high-speed interconnect. This time, both companies took the partnership one step further:

  • Amazon Annapurna Labs will support NVIDIA's new custom high-bandwidth memory NVHBM (developed in collaboration with memory vendors)
  • Trainium thereby gains a faster, more power-efficient memory option
  • Trainium and GPUs can work together within the same rack-scale architecture, sharing scale-up interconnect

For the chip industry, this is a signal worth watching closely: NVLink Fusion + NVHBM means NVIDIA's interconnect and memory technologies have started "supplying" competing ASICs. The boundary between the in-house ASIC camp (Trainium, TPU, MTIA) and the NVIDIA GPU camp is shifting from "either/or" to "hybrid deployment."

4. 100,000 GPUs: The U.S. Government Sovereign AI Factory​

A dedicated public-sector business is carved out of the partnership: AWS and NVIDIA will build AI factories for the U.S. government, deploying 100,000 GPUs on secure AWS infrastructure to host federal and national security workloads (Impact Level 6 and above), supporting the development of advanced AI models within strict regulatory frameworks.

Sovereign AI turning from a slogan into concrete numbers is a defining feature of this cycle — government customers are becoming first-class buyers of AI compute.

5. Software and Physical AI Deepen in Parallel​

Beyond hardware, the software layer of the partnership is also strengthening:

  • NVIDIA Nemotron open-source models continue to arrive on Amazon Bedrock and SageMaker
  • Amazon EMR data processing and OpenSearch vector indexing are accelerated by cuDF / cuVS
  • Amazon Robotics officially adopts the NVIDIA physical AI platform (Jetson, Omniverse, Isaac) for warehouse automation and next-generation robotics

Physical AI (robotics, embodied intelligence) is becoming a new growth pole in cloud providers' compute narratives — the same trend as JD.com, Tesla, and others writing embodied intelligence into their compute procurement logic.

6. Implications for Compute Buyers​

ObservationImplication
Demand "beat all forecasts"Supply tightness is not a short-term phenomenon; the 2027-2028 compute window must be locked in now
Mixed procurement across three architecturesDuring the Rubin/Rubin Ultra production ramp-up, Blackwell Ultra remains the delivery mainstay; procurement needs cross-generation planning
CPU layer revaluedagentic AI pushes the bottleneck from GPU to CPU orchestration and memory bandwidth; do not focus only on accelerator cards when selecting
Sovereign AI landsGovernment-scale orders enter the market, further tightening the allocatable supply of high-end GPUs

For decision-makers weighing build versus rent, every massive add-on order from cloud giants reprices the future rental curve. If you are evaluating GPU purchase or rental options, we recommend running a quantitative calculation with the TCO Calculator: enter chip price, power draw, utilization, and rental rates to compare 3-year total cost of ownership.

Summary​

2 million GPUs, Vera CPU in the cloud, NVHBM opening up, 100,000 sovereign compute GPUs — this round of expansion between AWS and NVIDIA pushes the "AI factory" race into the full-stack era. Compute scarcity will most likely only tighten before 2027; whether you are buying, renting, or betting on domestic alternatives, locking in supply and cost curves early is the surest move right now.

(The data in this article comes from official AWS/NVIDIA announcements and public statements by executives of both companies; architecture performance figures are as released by the vendors.)

NVIDIA Vera Rubin Enters Full Production: The Agentic AI Factory Era Begins

· 5 min read
Industry Research Team

On June 1, 2026, NVIDIA founder and CEO Jensen Huang officially announced at COMPUTEX 2026 (Taipei) that: the Vera Rubin platform has entered full production. This marks a fundamental paradigm shift for AI hardware from "discrete accelerators" to "integrated AI factories."

Key Highlights​

  • Rubin GPU: Next-gen AI compute chip, FP4 compute is 3.6× that of Blackwell
  • Vera CPU: 88 custom Arm cores (176 threads), replacing the Grace CPU
  • NVLink 6: GPU-to-GPU interconnect bandwidth reaches 260 TB/s (double Blackwell)
  • CX8 SuperNIC: 800Gb/s network, ConnectX-9 link reaching 28.8 TB/s
  • HBM4 memory: 288GB per chip, 13 TB/s bandwidth
  • Agentic throughput: 10× over Grace Blackwell

Complete Vera Rubin Platform Specs​

Vera Rubin is not a single GPU but a complete AI factory platform comprising 7 chips:

ChipTypePurpose
Rubin GPUMain AI compute chipTraining + inference
Rubin Ultra GPUFlagship versionUltra-scale inference
Vera CPUCPU paired with RubinHost CPU + data preprocessing
NVLink 6Interconnect chipHigh-speed GPU interconnect (260 TB/s)
CX8 SuperNICNIC800Gb/s network
XDR 800G switchDatacenter networkCross-rack communication
Rubin Platform PODWhole cabinetPre-configured AI factory (144 GPUs)

Rubin GPU Detailed Specs (estimated)​

ParameterRubin GPURubin UltraBlackwell (B200)
ArchitectureRubinRubin UltraBlackwell
ProcessTSMC 3nm (est.)TSMC 3nmTSMC 4NP
Memory288GB HBM4288GB HBM4E (est.)192GB HBM3e
Memory bandwidth13 TB/s13+ TB/s8 TB/s
FP4 compute~3,600 TFLOPS (est.)~5,000 TFLOPS (est.)2,250 TFLOPS
TDP1,000W (est.)1,200W (est.)700-1000W
InterconnectNVLink 6 (260 TB/s)NVLink 6NVLink 5 (1800 GB/s)
Mass production2026 Q3H2 20272024 Q4

📌 Note: Rubin's exact specs are not fully public yet; some values above are estimates.

Vera CPU: The New Host CPU Replacing Grace​

Vera CPU is NVIDIA's self-designed Arm-architecture CPU, replacing the previous Grace CPU:

ParameterVera CPUGrace CPU
Cores88 cores (176 threads)72 cores (144 threads)
ArchitectureCustom Armv9 (est.)Arm Neoverse V2
InterfaceNVLink 5.0 (1.8 TB/s)NVLink 4.0 (900 GB/s)
TDP~500W (est.)350-500W
PurposeAI factory Host CPUHPC / AI Host

Key upgrade: Vera's co-design with the Rubin GPU achieves end-to-end optimization in compute, data loading, and preprocessing, comparable to Google TPU 8t's Arm Axion integration.

Performance vs Blackwell​

NVIDIA officially claims that under the same POD configuration (144 GPU chips):

MetricGrace Blackwell (GB200 NVL72)Vera Rubin NVL144Improvement
FP4 compute1.1 PFLOPS3.6 PFLOPS3.3×
Memory capacity288GB×72 = 20.7TB288GB×144 = 41.4TB2×
Memory bandwidth8 TB/s×7213 TB/s×144~3.3×
NVLink bandwidth1800 GB/s×72260 TB/s (full POD)~2×
Agentic throughputBaseline10×10×
Performance per wattBaseline25× (vs CPU alone)25×

💡 Why "10× agentic throughput"? Agentic AI workloads differ from training/inference: one prompt may trigger multiple stages including reasoning, retrieval, tool calls, and response generation, involving thousands of steps. The Rubin platform is optimized for this long-chain, high-concurrency workload.

MGX Third-Gen Rack-Scale System​

Vera Rubin adopts the MGX third-gen open rack-scale system design:

  • Five-rack synergy: Vera Rubin NVL72 system + Vera CPU + Groq 3 LPX + Vera BlueField-4 STX storage + Spectrum-6 SPX Ethernet
  • Global supply chain: 30 countries, 350+ factories, hundreds of partners (Dell, HPE, Lenovo, Supermicro, Asus, Foxconn, etc.)
  • Spectrum-X Ethernet silicon photonics: World's first switch based on CPO (co-packaged optics) supporting 200Gb/s SerDes, now in mass production

Mass Production Timeline​

TimeEvent
Jan 2026CES 2026 first unveils Rubin platform
June 1, 2026COMPUTEX 2026 announces full production
Fall 2026Vera Rubin officially starts mass production and shipment
H2 2027Rubin Ultra launch (HBM4E upgrade)
2028Feynman architecture (next gen)

AI Factory: From Selling Chips to Selling "Smart Production Lines"​

Huang said something at the launch that shook the industry:

"Rubin's Agentic AI throughput is 10× that of Blackwell. Rubin is a complete AI factory platform."

This marks a fundamental shift in NVIDIA's business model:

  • Past: Sold GPUs (H100/B200), customers built systems themselves
  • Now: Sells "complete AI factory solutions" (Vera Rubin POD), including GPU, CPU, network, storage, software stack
  • Future: Becomes the "TSMC" of global AI infrastructure (providing smart production capacity)

vs Competitors​

VendorProductPositioningAdvantageDisadvantage
NVIDIAVera RubinComplete AI factory solutionMost complete ecosystem, most mature softwareExpensive, extremely high power
AMDMI455X (MI400 series)Training competitorPrice/performance, open ecosystemSoftware ecosystem gap
GoogleTPU 8i/8tCloud training/inferenceDeep Gemini integrationGoogle Cloud only
HuaweiAscend 910C/950Domestic substitutionChina localization, AscendMind frameworkAffected by export controls

Industry Impact​

  1. AI labs: Frontier model training time shrinks from "months" to "weeks"
  2. Cloud providers: Must decide whether to procure Vera Rubin POD (conflicts with self-developed chip strategy)
  3. Hyperscale datacenters: AI factory becomes a new competitive dimension (whoever has the strongest compute can train the strongest model)
  4. Domestic chips: Ascend 910C/950, Cambricon MLU590, etc. must catch up to Blackwell in 2026-2027, or the gap will widen to the Rubin era

References​


This article is compiled from NVIDIA official announcements and public materials. Some specs are estimates, subject to final official release.