Skip to main content

14 posts tagged with "NVIDIA"

NVIDIA AI chips and GPUs

View all tags

June 2026 AI Chip Major Events Roundup: Ascend 910C Trains Trillion-Parameter Model, OpenAI Custom Chip, RTX Spark Launch

· 6 min read
Industry Research Team

June 2026 saw multiple milestone events in the AI chip field, marking acceleration of two major trends: "domestic substitution" and "de-NVIDIA-ization."

1. Huawei Ascend 910C Completes 1.6-Trillion-Parameter DeepSeek V4 Pro Training (2026-06-05)​

Event Overview​

June 5, 2026, Shenzhen Hetao College, together with Harbin Institute of Technology (Shenzhen), Shenzhen Big Data Research Institute, Huawei, and other teams, relied on an Ascend 910C domestic AI compute cluster to successfully complete full-parameter post-training of the 1.6-trillion-parameter DeepSeek V4 Pro large model.

Technical Significance​

MetricValue
Model parameters1.6 trillion
Training chipAscend 910C cluster
Training typeFull Parameter Post-Training
SignificanceFirst time domestic AI chips complete trillion-parameter-level model training

Industry Impact​

  1. Breaks technology blockade: Proves domestic AI chips can train trillion-parameter models
  2. Accelerates "farewell to NVIDIA": DeepSeek fully switches to Huawei Ascend, reducing dependence on H100
  3. Domestic substitution inflection point: From "inference substitution" to "training substitution"

2. OpenAI Launches First Custom AI Inference Chip Jalapeño (2026-06-24)​

Event Overview​

June 24, 2026, OpenAI and Broadcom jointly launched the first custom AI inference chip Jalapeño, with a design cycle of only 9 months (industry average 18 months), using TSMC 3nm process.

Key Metrics​

MetricJalapeñoComparison (Blackwell)
ProcessTSMC 3nmTSMC 4nm
ArchitectureSystolic ArrayBlackwell GPU
Design cycle9 months~18 months
Inference cost-50%Baseline
AI-assisted design✅ First❌ No
DeploymentEnd of 2026Shipped

Strategic Significance​

  1. First AI chip with AI-assisted design: OpenAI used models like GPT-5.3-Codex-Spark to assist architecture exploration
  2. Accelerates "de-NVIDIA-ization": Tech giants (Google, Amazon, Microsoft, Meta, OpenAI) collectively develop custom chips
  3. Inference cost revolution: For OpenAI processing hundreds of millions of API calls daily, a 50% cost reduction is significant

3. NVIDIA Launches RTX Spark AI PC Superchip at Computex 2026 (2026-06-01)​

Event Overview​

June 1, 2026, NVIDIA CEO Jensen Huang launched the RTX Spark AI PC superchip at Computex 2026 / GTC Taipei, in collaboration with MediaTek, using an Arm CPU + Blackwell GPU unified-memory architecture.

Key Metrics​

MetricRTX Spark
CPUUp to 20-core Arm (with MediaTek)
GPU6,144 CUDA cores (Blackwell)
Unified memory128GB LPDDR5X (shared CPU+GPU)
Memory bandwidth300 GB/s
AI compute~1 PFLOPS (est.)
Model capacityCan run 120B-parameter models
ContextUp to 1 million tokens
TDP~100W (est.)
AvailabilityFall 2026

Industry Impact​

  1. NVIDIA enters PC chip market: Challenges Intel's dominance in personal computers
  2. New AI PC standard: Run 120B-parameter models locally, 1M-token context
  3. Windows transforms into AI Agent platform: Deep collaboration with Microsoft OpenShell framework

4. MIIT Publishes "2026 AI Chip Industry Development White Paper" (2026-06-09)​

Event Overview​

June 9, 2026, China's Ministry of Industry and Information Technology published the "2026 AI Chip Industry Development White Paper," predicting the domestic AI chip market will exceed 200 billion RMB in 2026.

Key Predictions​

Metric2026 Prediction
Market sizeExceed 200 billion RMB
Domestic chip share>50% (41% in 2025)
Edge inference chipsSignificant progress
Shipment growthMore than double (vs 2025)

Industry Significance​

  1. Domestic AI chip capitalization accelerates: Cambricon, Enflame, Moore Threads, etc. accelerate IPOs
  2. Edge inference becomes the breakthrough: Easier to achieve domestic substitution than training chips
  3. Policy dividend continues: Domestic substitution upgraded from "market behavior" to "national strategy"

5. ByteDance in Talks to Procure 50K Iluvatar Inference Chips (2026-06-17)​

Event Overview​

June 17, 2026, Reuters reported that ByteDance is in talks with Shanghai AI chip firm Iluvatar to procure at least 50,000 AI chips, mainly for inference tasks.

Deal Details​

ItemContent
BuyerByteDance
SupplierIluvatar
Chip modelZhiKai series (inference GPU)
QuantityAt least 50,000
UseInference workloads
Training chipTianTai series

Industry Significance​

  1. Domestic GPU top player "adds a member": Iluvatar enters a top internet company's supply chain for the first time
  2. ByteDance 2026 capex raised over 200B RMB: Mainly for AI compute and datacenters
  3. "Domestic substitution" extends from government/SOEs to private tech giants

Trend 1: "Domestic Substitution" Moves from Inference to Training​

  • Ascend 910C completes 1.6-trillion-parameter model training → Proves domestic chips have training capability
  • DeepSeek fully switches to Ascend → Leading AI companies first to "farewell to NVIDIA"
  • ByteDance procures Iluvatar → Private tech giants follow

Trend 2: "De-NVIDIA-ization" from Slogan to Action​

  • OpenAI Jalapeño → First custom chip, inference cost -50%
  • Google TPU, Amazon Trainium, Microsoft Maia → Continuous iteration
  • Meta MTIA, Apple M5 Ultra → Increased investment

Trend 3: AI PC and Edge Inference Become New Battlefield​

  • NVIDIA RTX Spark → New AI PC standard, launches Fall 2026
  • Edge inference chip localization accelerates → Key mention in MIIT white paper
  • "Local trillion-parameter model execution" → New consumer market selling point

Looking Ahead (2026 H2)​

  1. Ascend 950DT full scale-up (2026 Q4) → Huawei's latest-gen training chip
  2. NVIDIA Rubin R200 shipment (2026 H2) → Next-gen flagship
  3. AMD MI400 Helios rack (2026 H2) → Targets NVIDIA GB200
  4. OpenAI Jalapeño deployment (end of 2026) → Gigawatt-scale datacenters
  5. Domestic AI chip shipments more than double → CITIC Securities prediction

References​


This article is continuously updated. Please provide the latest developments.

Cambricon MLU690 vs NVIDIA H100: In-Depth Comparison — Can a Domestic AI Chip Replace the H100?

· 6 min read
AI Hardware Analyst

In 2026, against the backdrop of U.S. export controls on AI chips to China, Cambricon's MLU690 has drawn intense attention as a "China-made H100." This article compares the two in depth across compute, memory, power, software ecosystem, measured performance, and price to help you make a selection decision.

Core Verdict (Read This First)​

DimensionMLU690H100WinnerGap
BF16 compute600 TFLOPS989 TFLOPSH100+65%
Memory capacity64GB HBM380GB HBM3H100+25%
Memory bandwidth2 TB/s3.35 TB/sH100+68%
TDP280W700WMLU690-60%
Energy efficiency2.14 TFLOPS/W1.41 TFLOPS/WMLU690+52%
Software ecosystemNeuWare (~75% coverage)CUDA (100% coverage)H100large gap
Price~¥140,000~¥200,000MLU690-30%
Availabilitydomestic spot stockexport-controlledMLU690✅

One-line summary: MLU690 delivers roughly 60% of H100's compute, but at only 40% of the power and 70% of the price — a strong fit for AI training and inference in the Chinese market.


1. Detailed Spec Comparison​

1.1 Compute​

PrecisionMLU690H100 SXM5H200 SXM5Note
FP8~300 TFLOPS (est.)3,958 TFLOPS3,958 TFLOPSH100 supports FP8; MLU690 likely does not
BF16/FP16600 TFLOPS989 TFLOPS989 TFLOPSH100 leads by 65%
FP32~150 TFLOPS (est.)60 TFLOPS60 TFLOPSMLU690 estimate; H100 actually higher
INT81,200 TOPS1,979 TOPS1,979 TOPSH100 leads by 65%

Key findings:

  • ✅ MLU690 reaches 60% of H100's BF16 compute
  • ⚠️ H100 supports FP8 (4-bit); MLU690 likely does not (needs confirmation)
  • ⚠️ H100's higher INT8 compute favors inference scenarios

1.2 Memory​

ItemMLU690H100H200Note
Capacity64GB HBM380GB HBM3141GB HBM3eH200 largest
Bandwidth2 TB/s3.35 TB/s4.8 TB/sH200 highest
TypeHBM3HBM3HBM3eH200 uses latest HBM3e

Key findings:

  • ⚠️ MLU690 has 20% less memory than H100 (64GB vs 80GB)
  • ⚠️ MLU690 bandwidth is 40% lower than H100 (2 TB/s vs 3.35 TB/s)
  • ❌ When running 70B+ parameter models, MLU690 may run out of memory (model parallelism required)

1.3 Power​

ItemMLU690H100H200
TDP280W700W700W
Efficiency (FP16/W)2.14 TFLOPS/W1.41 TFLOPS/W1.41 TFLOPS/W
8-card server power~3.5kW~6kW~6kW
Annual electricity (¥0.6/kWh)~¥18,400~¥36,800~¥36,800

Key findings:

  • ✅ MLU690 draws only 40% of H100's power, sharply cutting data-center electricity cost
  • ✅ MLU690 leads efficiency by 52%, better suited to large-scale deployment
  • ✅ For power-sensitive inference, MLU690 has a clear edge

2. Software Ecosystem​

2.1 Framework Support​

FrameworkMLU690 (NeuWare)H100 (CUDA)Note
PyTorch✅ (PyTorch-Cambricon)✅ nativeMLU690 needs an extra plugin
TensorFlow✅ (TensorFlow-Cambricon)✅ nativesame
JAX⚠️ partial✅ nativeMLU690 limited
ONNX⚠️ partial✅ nativesame
vLLM⚠️ in progress✅ nativeMLU690 awaits community port

2.2 Operator Coverage​

CategoryMLU690H100Note
Basic operators✅ 95%✅ 100%conv, matmul, etc.
Transformer operators✅ 85%✅ 100%Attention, LayerNorm, etc.
Custom operators⚠️ hand-written✅ CUDA C++MLU690 harder to develop
LLM inference opt.⚠️ basic✅ mature (FlashAttention, PagedAttention)H100 leads

Key findings:

  • ⚠️ NeuWare is only 5–6 years old, with ~75–85% operator coverage
  • ❌ Complex LLMs (e.g., GPT-4, Claude) may need manual optimization
  • ✅ Common models (Llama, Qwen, GLM) are essentially already supported

3. Measured Performance​

3.1 Training​

ModelMLU690 (time)H100 (time)Speedup
Llama 7B~48 h (est.)~30 h1.6x
Llama 70B~7 days (est.)~4.5 days1.6x
Qwen 72B~8 days (est.)~5 days1.6x

Note: above figures are estimates; real performance depends on software optimization.

3.2 Inference​

ModelMLU690 (tok/s)H100 (tok/s)Note
Llama 7B~80 tok/s (est.)~120 tok/sH100 +50%
Llama 70B~20 tok/s (est.)~35 tok/sH100 +75%
Qwen 72B~18 tok/s (est.)~30 tok/sH100 +67%

Key findings:

  • ⚠️ H100 leads inference by 50–75%
  • ✅ But MLU690 draws only 40% the power, with better efficiency
  • ✅ For cost-sensitive inference, MLU690 is more economical

4. Price​

4.1 Hardware Procurement​

ItemMLU690H100H200
Per-card (domestic)~¥140,000~¥200,000~¥300,000
8-card server (turnkey)~¥1,200,000~¥1,800,000~¥2,600,000
Cost gap-+50%+117%

4.2 TCO (3 years)​

ItemMLU690H100Note
Hardware¥1,200,000¥1,800,000MLU690 33% cheaper
Electricity (3y)¥55,200¥110,400MLU690 50% cheaper
Facility¥150,000¥250,000MLU690 40% cheaper
TCO (3y)¥1,405,200¥2,160,400MLU690 35% cheaper

Key findings:

  • ✅ MLU690's TCO is 35% lower than H100's
  • ✅ For large-scale deployment (100+ cards), the cost advantage is pronounced

5. Selection Advice​

5.1 Choose MLU690 if...​

  • ✅ Your business is primarily in the Chinese market
  • ✅ You are affected by U.S. export controls and cannot buy H100/H200
  • ✅ You are power-sensitive (edge data centers, high electricity-cost regions)
  • ✅ Your models use common architectures (Llama, Qwen, GLM)
  • ✅ You have domestic-substitution requirements (government, SOEs, military)

5.2 Choose H100/H200 if...​

  • ✅ Your business is global
  • ✅ You need to train frontier models (GPT-4 class)
  • ✅ Your models use complex operators (need the CUDA ecosystem)
  • ✅ You demand extreme performance (low-latency inference)
  • ✅ You can legally procure H100/H200
ScenarioRecommended
TrainingH100 (high perf) + MLU690 (low-cost scale-out)
InferenceMLU690 (cost-sensitive) + H100 (low-latency)
Domestic projectall MLU690
International marketall H100/H200

6. Outlook​

6.1 MLU690's weaknesses​

  • ⚠️ Immature software ecosystem: 75–85% operator coverage; complex models need manual tuning
  • ⚠️ Small memory: 64GB limits support for 70B+ parameter models
  • ⚠️ Weak interconnect: Cambricon Link bandwidth below NVLink
  • ⚠️ Limited international market: affected by U.S. export controls

6.2 MLU690's improvement path​

  • 📅 MLU790 (2027): expected 5nm process, ~2x compute
  • 📅 Memory upgrade: next gen may adopt HBM3e, capacity up to 128GB
  • 📅 Software: NeuWare ecosystem improving, operator coverage target 95%

7. Summary​

DimensionMLU690H100Recommended scenario
Compute⭐⭐⭐⭐⭐⭐⭐⭐⭐H100 for top-tier training
Memory⭐⭐⭐⭐⭐⭐⭐H100 for large models
Power⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for inference
Ecosystem⭐⭐⭐⭐⭐⭐⭐⭐H100 for complex models
Price⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for large-scale deployment
Domestic⭐⭐⭐⭐⭐❌MLU690 for Chinese market

Final recommendation:

  • 🇨🇳 Chinese market: prefer MLU690 (domestic + low cost)
  • 🌍 International market: prefer H100/H200 (performance + ecosystem)
  • 💡 Hybrid: train on H100, infer on MLU690

References​


Disclaimer: Data in this article is based on public sources and reasonable estimates; actual performance is subject to vendor official testing. MLU690's software ecosystem is evolving rapidly — watch NeuWare updates.

Last updated: 2026-06-23

2026 H2 AI Chip Roadmap Major Update: Qualcomm Enters, AMD MI400 Three Models Unveiled, Huawei Three-Generation Roadmap

· 7 min read
AI Hardware Analyst

June 2026 update — the AI compute card market is undergoing its most dramatic reshuffling in years. This article walks through the latest roadmap developments.


Key Takeaways​

  • Qualcomm AI 200/250 officially enters the datacenter AI inference market, targeting NVIDIA H200
  • AMD MI400 series unveils three models: MI430X (HPC), MI440X (enterprise), MI455X (flagship)
  • Huawei publishes a three-generation roadmap: 950 (2026) → 960 (2027-Q4) → 970 (2028-Q4)
  • Intel Jaguar Shores timeline uncertain, possibly delayed to 2027 or later
  • NVIDIA Rubin R200 is in full mass production; the Vera CPU + Rubin GPU combination is now shipping

1. Qualcomm: Mobile Giant Moves Into Datacenter AI​

AI 100 → AI 200 → AI 250​

Qualcomm officially launched the AI 200 datacenter inference chip in October 2025, marking the mobile giant's formal entry into the datacenter AI market.

ModelLaunchAvailabilityKey Features
AI 1002025-102026 H2Rack-scale AI inference, 768GB LPDDR per card
AI 2502025-102027 H1Near-memory computing architecture, 10x effective memory bandwidth

Why Qualcomm Can Succeed​

  1. Low TCO: LPDDR memory is far cheaper than HBM
  2. Energy efficiency: Mobile chip design heritage, excellent power control
  3. Inference-focused: Not chasing training performance, focused on inference scenarios
  4. Rack form factor: Direct liquid cooling, 160kW rack-level power, Ethernet interconnect

Market Impact​

  • Takes on NVIDIA H200: AI 200 inference performance approaches H200 but with 30-40% lower TCO
  • Pressures NVIDIA: May push NVIDIA to launch inference-specific chips (e.g., Rubin CPX)
  • Diversifies choice: Breaks NVIDIA's monopoly in the inference market

2. AMD MI400 Series: Three Models, Precise Positioning​

At CES 2026 (January 2026), AMD officially unveiled the three models of the MI400 series, precisely covering different markets:

MI430X (HPC + Sovereign AI)​

FeatureSpec
PositioningHPC + sovereign AI
FP32/FP64Supported (key differentiator)
Use casesScientific computing, climate simulation, national AI infrastructure
CompetitorNVIDIA does not make FP64 AI cards

MI440X (Enterprise Servers)​

FeatureSpec
PositioningEnterprise 8-GPU servers
CompatibilityWorks with existing datacenter infrastructure
Use casesEnterprise AI, private cloud, edge inference
AdvantageCheaper and easier to deploy than MI455X

MI455X (Flagship AI Training)​

FeatureSpec
PositioningFlagship AI training + inference
Optimized precisionFP4/FP8/BF16
Helios rackCore component
CompetitorNVIDIA Rubin R200

Helios Rack-Scale Solution​

AMD also launched the Helios rack-scale AI solution at CES 2026:

  • 18 Zen 6 CPUs (2nm process)
  • 72 MI455X GPUs
  • Direct liquid cooling
  • Shipment expected in 2026 H2

3. Huawei Three-Generation Roadmap: 950 → 960 → 970​

Huawei unveiled its three-generation chip roadmap at HC 2025 (September 2025) with a very clear timeline:

Ascend 950 Series (2026)​

ModelLaunchKey Features
950PR2026-Q1PR (inference-optimized), already in mass production
950DT2026-Q4DT (Decode + training), expected to scale up

Technical highlights:

  • Added FP8/MXFP8/MXFP4 support
  • Interconnect bandwidth 2TB/s (2.5x over 910C)

Ascend 960 (2027-Q4)​

  • Doubled compute: All specs double versus the 950 series
  • FP8: ~2 PFLOPS expected
  • Process: N+3 (equivalent to 5nm)
  • Positioning: Targets NVIDIA B200

Ascend 970 (2028-Q4)​

  • Third-generation flagship: Only timeline announced, specs TBD
  • Significance: Huawei's first complete generation-spanning roadmap
  • Signal: China's domestic AI chips have entered a "roadmap-driven" phase

4. Intel Jaguar Shores: Timeline Uncertain​

Original Plan​

  • Launch: 2026
  • Architecture: Xe-HPC + Gaudi fusion
  • Process: 18A (Intel's most advanced)
  • Memory: Possibly HBM4E (instead of originally planned HBM4)

Latest Developments​

  • Possible delay: Some sources suggest a slip to 2027
  • Competitors: AMD MI400 already unveiled, NVIDIA Rubin in mass production
  • Market pressure: Intel is losing ground in the AI chip market; Jaguar Shores is its last chance

Impact on Roadmap​

If Jaguar Shores slips to 2027, Intel will essentially be out of the AI chip market.


5. NVIDIA Rubin Platform: Full Mass Production​

Rubin R200 (2026-Q2 full mass production)​

FeatureSpec
HBM288GB HBM4
Compute50 PFLOPS FP4
NVLinkNVLink 6 (1800 GB/s)
ProcessTSMC 4NP

Rubin NVL72 Cabinet (2026 H2 shipment)​

  • 72 Rubin GPUs
  • 36 Vera CPUs
  • 1.8 EFLOPS FP4
  • Direct liquid cooling

Vera CPU (Debut)​

  • Architecture: Custom CPU replacing Grace
  • Positioning: Deep co-design with Rubin GPU
  • Significance: NVIDIA's transformation from a GPU company into a computing platform company

6. Google TPU v8: Training/Inference Officially Split​

TPU 8t (training) + TPU 8i (inference)​

At Cloud Next 2026, Google announced TPU v8 would officially split into training and inference versions:

FeatureTPU 8t (training)TPU 8i (inference)
OptimizationHigh compute, high bandwidthLow latency, low cost
InterconnectOptical interconnectEthernet
Launch20272027

Significance​

  • Industry trend: Specialization of training/inference chips
  • Followers: Qualcomm AI 200 is also inference-only
  • NVIDIA pressure: Does it need an inference-specific chip?

7. Cerebras WSE-4: Wafer-Scale Engine Evolves​

Core Specs​

FeatureSpec
Transistors1.4 trillion
Compute125 PFLOPS FP8
Launch2026 H2
ProcessTSMC 5nm

Competitive Advantages​

  • Massive model training: A single WSE-4 can train 10T+ parameter models
  • Low-latency inference: Entire model on one chip, no communication overhead
  • Mature software stack: Cerebras stack already supports PyTorch, TensorFlow

8. Market Landscape Analysis​

Training Market​

RankVendorProductMarket Share (est.)
1NVIDIARubin R20070%
2AMDMI455X15%
3GoogleTPU v8t10%
4HuaweiAscend 9605% (mostly China)

Inference Market (New Battlefield)​

RankVendorProductAdvantage
1NVIDIAH200 / Rubin CPXMature ecosystem
2QualcommAI 200Low TCO
3AMDMI440XGood compatibility
4IntelGaudi 4Low price

Trend 1: Rise of Inference-Specific Chips​

  • Qualcomm AI 200: Mobile giant enters the market
  • NVIDIA Rubin CPX: NVIDIA's first inference-specific chip
  • Google TPU 8i: Training/inference officially split

Trend 2: Rack-Scale Solutions Become Standard​

  • NVIDIA NVL72: 72 GPU + 36 CPU
  • AMD Helios: 18 CPU + 72 GPU
  • Qualcomm rack: 160kW liquid-cooled rack

Trend 3: China's Domestic Chips Enter "Roadmap-Driven" Phase​

  • Huawei three-generation roadmap: 950 → 960 → 970
  • Clear timeline: 2026-Q1 → 2027-Q4 → 2028-Q4
  • Significance: From "catch-up" to "planning"

Trend 4: HBM Capacity Becomes the Bottleneck​

  • SK hynix: HBM4 capacity already booked by NVIDIA
  • Samsung: HBM4E samples delivered to AMD
  • Impact: MI400 and Rubin R200 shipments constrained by HBM capacity

10. Procurement Recommendations​

If Procuring in 2026 H2​

  1. Training scenarios:

    • First choice: NVIDIA Rubin R200 (best performance)
    • Alternative: AMD MI455X (better price/performance)
    • Domestic: Huawei Ascend 950DT (China-based customers)
  2. Inference scenarios:

    • First choice: NVIDIA H200 (mature ecosystem)
    • Best value: Qualcomm AI 200 (if available)
    • Cost-sensitive: AMD MI440X
  3. HPC scenarios:

    • Only choice: AMD MI430X (FP64 support)

If Procuring in 2027​

  • Wait for Rubin Ultra: Performance possibly 2x R200
  • Watch MI500: AMD's next-generation product
  • Evaluate TPU v8: If already on Google Cloud

Conclusion​

2026 H2 will be the most fiercely contested half-year in AI chip market history:

  • NVIDIA continues to lead, but its advantage is narrowing
  • AMD precisely positions three models; market share will keep rising
  • Qualcomm enters the inference market; its low-TCO strategy may disrupt the market
  • Huawei has a clear three-generation roadmap; domestic substitution accelerates
  • Intel's Jaguar Shores is make-or-break

For procurement decision-makers, this is the hardest time to decide — every option has clear pros and cons.

For engineers, this is the best of times — chip performance doubles yearly, architectural innovation is endless.


References​

  • AI Compute Card Future Roadmap - MirrorFrog real-time updates
  • NVIDIA Rubin R200 deep dive (see related articles on this site)
  • AMD MI400 series CES 2026 launch (see related articles on this site)
  • Qualcomm AI 100 launch analysis (coming soon)

Last updated: 2026-06-20
Author: Charles Qing
Tags: #roadmap #market-analysis #procurement

NVIDIA Vera Rubin Enters Full Production: The Agentic AI Factory Era Begins

· 5 min read
Industry Research Team

On June 1, 2026, NVIDIA founder and CEO Jensen Huang officially announced at COMPUTEX 2026 (Taipei) that: the Vera Rubin platform has entered full production. This marks a fundamental paradigm shift for AI hardware from "discrete accelerators" to "integrated AI factories."

Key Highlights​

  • Rubin GPU: Next-gen AI compute chip, FP4 compute is 3.6× that of Blackwell
  • Vera CPU: 88 custom Arm cores (176 threads), replacing the Grace CPU
  • NVLink 6: GPU-to-GPU interconnect bandwidth reaches 260 TB/s (double Blackwell)
  • CX8 SuperNIC: 800Gb/s network, ConnectX-9 link reaching 28.8 TB/s
  • HBM4 memory: 288GB per chip, 13 TB/s bandwidth
  • Agentic throughput: 10× over Grace Blackwell

Complete Vera Rubin Platform Specs​

Vera Rubin is not a single GPU but a complete AI factory platform comprising 7 chips:

ChipTypePurpose
Rubin GPUMain AI compute chipTraining + inference
Rubin Ultra GPUFlagship versionUltra-scale inference
Vera CPUCPU paired with RubinHost CPU + data preprocessing
NVLink 6Interconnect chipHigh-speed GPU interconnect (260 TB/s)
CX8 SuperNICNIC800Gb/s network
XDR 800G switchDatacenter networkCross-rack communication
Rubin Platform PODWhole cabinetPre-configured AI factory (144 GPUs)

Rubin GPU Detailed Specs (estimated)​

ParameterRubin GPURubin UltraBlackwell (B200)
ArchitectureRubinRubin UltraBlackwell
ProcessTSMC 3nm (est.)TSMC 3nmTSMC 4NP
Memory288GB HBM4288GB HBM4E (est.)192GB HBM3e
Memory bandwidth13 TB/s13+ TB/s8 TB/s
FP4 compute~3,600 TFLOPS (est.)~5,000 TFLOPS (est.)2,250 TFLOPS
TDP1,000W (est.)1,200W (est.)700-1000W
InterconnectNVLink 6 (260 TB/s)NVLink 6NVLink 5 (1800 GB/s)
Mass production2026 Q3H2 20272024 Q4

📌 Note: Rubin's exact specs are not fully public yet; some values above are estimates.

Vera CPU: The New Host CPU Replacing Grace​

Vera CPU is NVIDIA's self-designed Arm-architecture CPU, replacing the previous Grace CPU:

ParameterVera CPUGrace CPU
Cores88 cores (176 threads)72 cores (144 threads)
ArchitectureCustom Armv9 (est.)Arm Neoverse V2
InterfaceNVLink 5.0 (1.8 TB/s)NVLink 4.0 (900 GB/s)
TDP~500W (est.)350-500W
PurposeAI factory Host CPUHPC / AI Host

Key upgrade: Vera's co-design with the Rubin GPU achieves end-to-end optimization in compute, data loading, and preprocessing, comparable to Google TPU 8t's Arm Axion integration.

Performance vs Blackwell​

NVIDIA officially claims that under the same POD configuration (144 GPU chips):

MetricGrace Blackwell (GB200 NVL72)Vera Rubin NVL144Improvement
FP4 compute1.1 PFLOPS3.6 PFLOPS3.3×
Memory capacity288GB×72 = 20.7TB288GB×144 = 41.4TB2×
Memory bandwidth8 TB/s×7213 TB/s×144~3.3×
NVLink bandwidth1800 GB/s×72260 TB/s (full POD)~2×
Agentic throughputBaseline10×10×
Performance per wattBaseline25× (vs CPU alone)25×

💡 Why "10× agentic throughput"? Agentic AI workloads differ from training/inference: one prompt may trigger multiple stages including reasoning, retrieval, tool calls, and response generation, involving thousands of steps. The Rubin platform is optimized for this long-chain, high-concurrency workload.

MGX Third-Gen Rack-Scale System​

Vera Rubin adopts the MGX third-gen open rack-scale system design:

  • Five-rack synergy: Vera Rubin NVL72 system + Vera CPU + Groq 3 LPX + Vera BlueField-4 STX storage + Spectrum-6 SPX Ethernet
  • Global supply chain: 30 countries, 350+ factories, hundreds of partners (Dell, HPE, Lenovo, Supermicro, Asus, Foxconn, etc.)
  • Spectrum-X Ethernet silicon photonics: World's first switch based on CPO (co-packaged optics) supporting 200Gb/s SerDes, now in mass production

Mass Production Timeline​

TimeEvent
Jan 2026CES 2026 first unveils Rubin platform
June 1, 2026COMPUTEX 2026 announces full production
Fall 2026Vera Rubin officially starts mass production and shipment
H2 2027Rubin Ultra launch (HBM4E upgrade)
2028Feynman architecture (next gen)

AI Factory: From Selling Chips to Selling "Smart Production Lines"​

Huang said something at the launch that shook the industry:

"Rubin's Agentic AI throughput is 10× that of Blackwell. Rubin is a complete AI factory platform."

This marks a fundamental shift in NVIDIA's business model:

  • Past: Sold GPUs (H100/B200), customers built systems themselves
  • Now: Sells "complete AI factory solutions" (Vera Rubin POD), including GPU, CPU, network, storage, software stack
  • Future: Becomes the "TSMC" of global AI infrastructure (providing smart production capacity)

vs Competitors​

VendorProductPositioningAdvantageDisadvantage
NVIDIAVera RubinComplete AI factory solutionMost complete ecosystem, most mature softwareExpensive, extremely high power
AMDMI455X (MI400 series)Training competitorPrice/performance, open ecosystemSoftware ecosystem gap
GoogleTPU 8i/8tCloud training/inferenceDeep Gemini integrationGoogle Cloud only
HuaweiAscend 910C/950Domestic substitutionChina localization, AscendMind frameworkAffected by export controls

Industry Impact​

  1. AI labs: Frontier model training time shrinks from "months" to "weeks"
  2. Cloud providers: Must decide whether to procure Vera Rubin POD (conflicts with self-developed chip strategy)
  3. Hyperscale datacenters: AI factory becomes a new competitive dimension (whoever has the strongest compute can train the strongest model)
  4. Domestic chips: Ascend 910C/950, Cambricon MLU590, etc. must catch up to Blackwell in 2026-2027, or the gap will widen to the Rubin era

References​


This article is compiled from NVIDIA official announcements and public materials. Some specs are estimates, subject to final official release.