Skip to main content

11 posts tagged with "Product Launch"

Official releases and first reviews of new AI chips / GPUs / ASICs

View all tags

Groq 3 LPX Enters Full Mass Production: Samsung 4nm Foundry, 315 PFLOPS FP8 per Rack — NVIDIA Turns the LPU into the Seventh Chip of Vera Rubin

· 5 min read
Industry Research Team

This article is based on NVIDIA official product pages, the March 2026 GTC architecture blog, and third-party benchmark data; performance figures are vendor architecture comparisons, and actual gains should be verified with your own tests.

NVIDIA has confirmed that Groq 3 LPX — the "seventh chip" of the Vera Rubin platform — has entered full mass production, manufactured at Samsung's Pyeongtaek campus. This is the first time the deal has landed in the form of production silicon since December 2025, when NVIDIA spent roughly $20 billion to obtain a non-exclusive IP license to Groq plus its core engineering team.

For the inference hardware landscape, this is a paradigm-level event: for the first time, NVIDIA has incorporated a "specialized inference architecture defined by someone else" into its own flagship platform — and not as an add-on sale, but as deep co-design.

1. The Groq 3 LPU Chip and the LPX Rack: Specs at a Glance​

Single Groq 3 LPU (4nm, Samsung foundry):

ItemValue
Compiler-managed on-chip SRAM500 MB
SRAM bandwidth150 TB/s
Chip-to-chip scale-up bandwidth2.5 TB/s (96 chip-to-chip links, 112 Gbps each)

A single LPX rack (32 1U liquid-cooled trays, 8 LPUs per tray):

ItemValue
Total LPUs256
FP8 compute315 PFLOPS
Total on-chip SRAM128 GB
Aggregate SRAM bandwidth40 PB/s
In-rack scale-up bandwidth640 TB/s
DDR5 memory12 TB (hosting large model weights)

The 500 MB of SRAM per LPU may not look like much, but multiplied by 256 LPUs and stacked with 40 PB/s of aggregate bandwidth, it forms the physical foundation of "deterministic low-latency decoding" — the LPU's design philosophy is precisely to use compiler static scheduling of on-chip SRAM to completely eliminate memory-fetch stalls during the inference decoding phase, which is exactly the most typical bottleneck of GPU inference.

2. AFD: Attention and FFN Split Up, GPU and LPU Each Do Their Own Job​

LPX does not replace the GPU; instead, it forms a heterogeneous system with Vera Rubin NVL72, centered on AFD (Attention-FFN Disaggregation):

  1. Rubin GPUs handle prefill and attention — building the KV cache over large contexts and executing attention layers, consuming HBM capacity and high throughput;
  2. Groq 3 LPUs handle FFN / MoE expert-layer decoding — latency-sensitive and pattern-predictable, exactly the home turf of deterministic SRAM scheduling;
  3. Intermediate activations are exchanged between the two engines token by token, orchestrated and routed by NVIDIA Dynamo: the GPU computes every attention layer, the LPU computes every feed-forward layer, jointly producing each output token.

The logic of this division of labor is clear: agentic AI applications can consume 15x the tokens of traditional AI applications, and the bottleneck shifts from "can it finish computing" to "can it keep emitting tokens at stable low latency." Let the big HBM container run attention, and let the deterministic SRAM engine run decode — each plays to its strengths.

3. Measured Results and Official Figures​

  • Third-party benchmarks (Artificial Analysis): with Gemma 4 31B at a 100K token context, LPX output reaches roughly 3400 tokens/s — inter-token intervals below 1 millisecond, a qualitative leap in interactivity for long-context agentic scenarios
  • Official architecture comparisons: with LPX added, Vera Rubin NVL72 achieves up to 35x higher throughput per megawatt on trillion-parameter models; up to 10x more revenue opportunity per watt in "high-value token" scenarios
  • Launch customers: Nebius is the first to deploy LPX racks to expand its token capacity; CoreWeave already connects Vera Rubin racks in production with Spectrum-X Multiplane

It must be emphasized: the 35x/10x figures are architecture comparisons at specific high-interaction operating points, not universal conclusions. For training, high-throughput batch inference, and workloads that need CUDA ecosystem flexibility, the GPU remains the right answer; LPX's home turf is scenarios where "single-user interactive latency is the product" — agent loops, real-time coding assistants, long-context conversations.

4. Three Industry Signals​

1. Specialized inference chips get absorbed, not opposed. The old narrative of the LPU as a "GPU challenger" has become part of Vera Rubin. The endgame for inference hardware may not be one architecture winning, but the heterogeneous combination of "GPU + specialized decoding engines" becoming the standard. For other inference chip startups (Cerebras, Etched, etc.), this is both proof that the ceiling has risen and a warning: get integrated or find differentiation.

2. Samsung foundry lands a high-end AI order. Against the backdrop of TSMC's near-monopoly on AI main chips, Samsung 4nm taking on LPX mass production is highly significant — combined with Tesla's earlier AI6 2nm order, Samsung foundry has a shot at returning to full-year profitability in 2027.

3. Inference economics enters the "priced per MW" era. When vendors start telling their story with tokens/MW and revenue per watt, the core KPI of compute selection has completely shifted from peak compute (TFLOPS) to token output per unit of energy. This is consistent with the power-and-electricity cost model built into our TCO Calculator: the chip with the best-looking peak specs is not necessarily the chip with the lowest cost per token.

Summary​

Groq 3 LPX mass production marks the arrival of a heterogeneous era for inference hardware: "GPUs manage throughput, LPUs manage latency." For teams currently selecting inference clusters, the recommendation is to evaluate workloads separately: keep batch offline inference on GPUs, and separately calculate the unit cost of LPX-class solutions for interactive long-context agents. If in doubt, run the 3-year total cost of ownership of both architectures through the TCO Calculator before deciding.

(Performance data in this article comes from NVIDIA official architecture comparisons and Artificial Analysis third-party benchmarks; for actual deployments, please rely on your own testing.)

Alibaba Zhenwu V900 Unveiled at Apsara Conference: 3x M890 Performance, 216GB Memory, 500,000-Card Cluster, Mass Production in 2027 Q1

· 5 min read
Industry Research Team

Just five days after Huawei officially announced the Ascend 960 at HUAWEI CONNECT on September 17, Alibaba unveiled its new unified training-and-inference AI chip, Zhenwu V900, at the Apsara Conference in Hangzhou on September 22 — which Alibaba calls "the most powerful Chinese self-developed AI chip in terms of compute performance to date." This article is compiled from Alibaba's official announcements and reports from Sina Tech, C114, Huanqiu.com, and other sources.


1. Single Chip: 3x M890, 216GB + 1200GB/s​

MetricZhenwu V900 (this launch)Zhenwu M890 (previous generation)
Performance3x Zhenwu M890 (official figure; absolute value undisclosed)FP16 600 TFLOPS (as catalogued on this site)
Memory216 GB144 GB (HBM3)
Chip-to-chip interconnect1200 GB/s—
PrecisionNative FP8 / FP4 (including high-precision training)—
PositioningUnified training and inference (trillion-parameter-scale training + low-precision inference)Unified training and inference
Mass production2027 Q1, scaled deployment in Alibaba Cloud data centersAlready deployed at scale

Three takeaways:

  • Memory crosses into the 200GB+ tier: 216GB puts it in the same capacity class as the Ascend 960DT (288GB self-developed HBM) and NVIDIA Rubin (288GB HBM4), leveling the capacity threshold for long-context and very large MoE models;
  • 1200GB/s chip-to-chip interconnect: laying the foundation for the supernode's "memory-semantics interconnect" — the chip-level prerequisite that pairs with ICN Switch;
  • Precision coverage across all scenarios: officially described as "high-precision training, low-precision and ultra-low-precision inference across all scenarios," with native FP8/FP4 support aligned with the common spec of 2026's new cards.

The absolute single-chip compute figure has not been officially disclosed; this site records it using the official relative figure of "3x M890." See the Zhenwu M890 spec page for M890 details.

2. System Level: ICN Switch + Panjiu Supernode, a 500,000-Card Single Cluster​

V900's real selling point is not the single chip but system-level collaboration:

  • ICN Switch self-developed interconnect chip: once connected, the supernode gains native memory semantics and unified memory addressing, with a thousand cards interconnected at full bandwidth — over a thousand V900s can "work together like a single super chip";
  • Panjiu supernode server: integrates V900 (compute) + ICN Switch (interconnect) + Panmai intelligent NIC (network) + Zhenyue SSD controller (storage), a fully self-developed compute-storage-network stack;
  • 500,000-card single cluster: combined with Alibaba Cloud's next-generation intelligent computing center network architecture, a single cluster can scale up to 500,000 cards.

This mirrors Huawei's "11 key chips" approach: the unit of competition has shifted from the single chip to whole-system delivery capability — Zhenwu handles compute, Yitian handles general-purpose computing, Panmai handles networking, Zhenyue handles storage, and ICN Switch handles interconnect.

3. Business and Roadmap​

  • The Zhenwu family has served over 650+ enterprise customers (as of June 2026), spanning autonomous driving, finance, large models, embodied AI, energy, and manufacturing;
  • Supernodes based on the M890 are already deployed at scale, running models with over 2 trillion parameters such as Qwen3.8 and Kimi K3; Alibaba Cloud will add new serving nodes in Q4 to expand supernode supply;
  • Roadmap: V900 enters mass production and sales in 2027 Q1; Zhenwu J900 is planned for release in 2027 Q3;
  • CPU synergy: Yitian 720 / 730 arrive in 2027 (the 730 is the first to adopt T-Head's fully self-developed CPU microarchitecture, with single-core SPECint2017/GHz up to 1.4x that of Yitian 710); the 2029 Yitian 750 will interconnect directly with Zhenwu AI chips via the ICN bus;
  • Alibaba Group CEO Eddie Wu said T-Head's AI chip annual shipment volume will increase substantially.

4. Competitive Coordinates: Two Swords in One Week​

Viewing the two mid-September launches side by side, the landscape of "system-level competition" among domestic AI chips is now clear:

DimensionHuawei Ascend 960 (9-17)Alibaba Zhenwu V900 (9-22)
Launch eventHC2026Apsara Conference 2026
Per-card memory288GB (self-developed HBM)216GB
Memory bandwidth9.6 TB/sUndisclosed
InterconnectLinJu UnifiedBus / NPO optical interconnectICN Switch memory-semantics interconnect
SupernodeAscend 960 supernode (4,096 cards, 8 EFLOPS FP8)Panjiu supernode (500,000-card single cluster)
Availability2027 Q22027 Q1

Huawei is taking the "supernode + open-source CANN ecosystem" route, while Alibaba is taking the "full cloud stack + open-source model ecosystem" route; both still trail in per-card specs, but what they deliver are procurable, operable ultra-large-scale clusters. The substitution logic against NVIDIA's CUDA ecosystem is shifting from "performance parity" to "supply certainty + system efficiency."

5. Summary​

  • V900: 3x M890, 216GB, 1200GB/s chip-to-chip interconnect, native FP8/FP4, mass production in 2027 Q1;
  • System: ICN Switch with thousand-card full bandwidth, Panjiu supernode with a fully self-developed compute-storage-network stack, 500,000-card single cluster;
  • Cadence: J900 in 2027 Q3, Yitian 730 in 2027, Yitian 750 in 2029 — T-Head's first three-year generational roadmap;
  • What to watch: actual mass production and Alibaba Cloud deployment scale in 2027 Q1, disclosure of V900's absolute compute, and the possibility of external supply of ICN Switch.

Further Reading​

References​

  • Huanqiu.com: T-Head launches new unified training-and-inference AI chip Zhenwu V900 (2026-09-22)
  • Sina Tech / TechWeb: Alibaba T-Head launches new AI chip Zhenwu V900, mass production in the first quarter of 2027
  • C114: Alibaba launches new AI chip Zhenwu V900, mass production in 2027 Q1

This article is compiled from Alibaba's official announcements and public reports. The absolute single-chip compute of V900 has not been officially disclosed; supernode and cluster scale figures are official claims, subject to actual future delivery.

昇腾 960 官宣发布:FP4 算力翻倍、全球首个 NPO 超节点,华为确认一年一代

· 6 min read
Industry Research Team

2026 年 9 月 17 日,华为全联接大会(HC2026),华为常务董事、ICT 基础设施业务总裁汪涛正式发布昇腾 960。与 8 月数字中国峰会上的"路线图预告"不同,这次是完整的规格官宣——而且提前三个季度就绪。本文基于华为官网新闻稿及多方报道整理,逐项拆解 960 的规格、超节点与路线图含义。


1. 昇腾 960:双版本,训练先行​

昇腾 960 分为两个版本,节奏错开一个季度:

指标昇腾 960DT(训练版)昇腾 960PR(推理版)
FP8 算力2 PFLOPS(较 950 翻倍)待披露
FP4 算力4 PFLOPS待披露
显存288GB(自研 HBM)待披露
显存带宽9.6 TB/s待披露
配套超节点Atlas 860(风冷)Atlas 960(液冷)
就绪 / 上市2027 Q1 / 2027 Q22027 Q3

三个读数:

  • 提前三个季度就绪:960DT 原计划 2027 年 Q4,现在 2027 Q1 就绪、Q2 上市——与 950DT"提前上线华为云"一脉相承,路线图可信度在持续兑现;
  • 精度口径与英伟达对齐:FP8 / FP4 主口径(辅以自研 HiF8 / HiFP8),FP4 已是 2026 新卡的通用口径;
  • 288GB 自研 HBM + 9.6TB/s:显存容量与带宽同步翻倍,对训练长上下文大模型是实打实的容量红利。

单卡规格详见昇腾 960 规格页;上一代昇腾 950DT 规格页。

2. 昇腾 960 超节点:全球首个 NPO 超节点​

这是本场发布真正的"重器"——昇腾 960 超节点是全球第一个采用 NPO(近封装光学)的超节点:

指标昇腾 960 超节点
单节点规模4096 卡
系统算力8 EFLOPS FP8 / 16 EFLOPS FP4
HBM 总容量1 PB
互联 RTT低至 2μs
可用度99.8%

NPO 意味着什么:传统光模块挂在交换机面板上,信号要先出封装再转电—光;NPO 把光引擎挪到交换芯片封装近旁,大幅缩短 SerDes 距离、降低功耗与时延。华为的具体数字:

  • 5500 个自研 Hi-ONE 光引擎(业界首个量产 NPO,单引擎 7.2T);
  • 替代 4.8 万颗 800G 光模块;
  • 降低功耗 550kW 以上,同时支撑 2μs 级互联时延。

对比 8 月数字中国峰会的口径(Atlas 960 单超节点 15488 卡),官方最新口径为单超节点 4096 卡(1 个超集群可由多超节点组成)——以华为官网最新新闻稿为准。

⚠️ 华为还给出了"MFU 提升 2.75 倍""推理时延降低 70%"等收益数据,均出自华为马尔科夫实验室仿真,尚无独立第三方实测,阅读时注意口径。

3. 一年一代:970(2028)→ 980(2029)​

华为首次把"一年一代"从愿景变成官宣承诺:

年份产品状态
2026昇腾 950 系列950PR 已量产,950DT Q4 放量
2027昇腾 960本次官宣,提前就绪
2028昇腾 970HC2026 确认
2029昇腾 980HC2026 确认

支撑这一节奏的是华为提出的"韬定律":算力规格每代翻倍,访存带宽、访存容量、互联带宽同步大幅提升。

生态侧的同场数据也值得记录:CANN 外部开发者占比首次超过 61%、月活开发者 5200 人;910C 超节点部署已超 1000 套。

4. 竞争坐标:单卡有代差,系统级对打​

把 960DT 放到 2027 年的棋盘上看:

  • 对 NVIDIA Rubin(R200:288GB HBM4 / 50 PFLOPS FP4):单卡 FP4 算力约为 R200 的 8%(4 vs 50 PFLOPS),单卡代差客观存在;但 4096 卡超节点 + NPO 互联是系统级竞争——用"可交付的超大集群"对打单卡性能;
  • 对采购方:960 的价值主张是"在国产量产约束内拿到最大可用集群"。DeepSeek 拟采购 16 万颗 950DT 跑推理的订单已经证明:当供给与生态到位,头部实验室愿意主动选国产;
  • 对国产链条:950→960 的显存翻倍直接拉动国产 HBM 迭代,这是整条链的胜负手。

5. 小结​

  • 960DT:FP8 2P / FP4 4P、288GB HBM、9.6TB/s,2027 Q2 上市,提前三个季度;
  • 960 超节点:全球首个 NPO 超节点,4096 卡 / 8 EFLOPS FP8 / 1PB HBM / RTT 2μs;
  • 路线图:一年一代官宣至 2029,950、960 均提前兑现,规划可信度显著上升;
  • 看什么:2027 Q2 实际交付节奏、国产 HBM 产能、CANN 生态的第三方模型覆盖度。

相关阅读​

参考资料​

  • 华为官网:发布全球首个采用 NPO 的超节点(昇腾 960 超节点)(2026-09-17)
  • 电子工程专辑:国产 AI 芯片新突破!华为昇腾 960 芯片将提前发布
  • 环球网:华为全联接大会 2026 主题演讲报道

本文基于华为官网新闻稿与公开报道整理。MFU、时延等收益数据为华为实验室仿真口径;2028 / 2029 产品仅为路线图官宣,规格以未来发布为准。

Computex 2026 Wrap-Up: AI PC Chip War Begins, NVIDIA RTX Spark Arrives Fall 2026

· 3 min read
Industry Research Team

June 6, 2026 — COMPUTEX 2026 concluded yesterday in Taipei. Under the theme "AI Together," this year's event set records with 1,500+ exhibitors and 6,000 booths. The head-to-head battle between NVIDIA, Intel, and AMD in the AI PC space was the defining story of the show.

1. NVIDIA RTX Spark: June Launch at $1,399​

Less than a week after its COMPUTEX debut, the NVIDIA-MediaTek RTX Spark Superchip confirmed its commercial timeline:

DetailInfo
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Starting PriceNot yet announced (analysts estimate $3,000-4,000)
Core SpecsArm CPU (up to 20 cores) + Blackwell GPU (6,144 CUDA cores)
Unified Memory128 GB LPDDR5X (300 GB/s)
Model CapacityRuns 120B parameter models, up to 1M token context

Market Reaction: AMD, Intel, and Qualcomm shares fell following the announcement. Analysts believe RTX Spark will reshape the market across three fronts — Windows AI PCs, creator workstations, and edge inference nodes.


2. Intel 18A in Full Production: Clearwater Forest + Crescent Island​

Intel CEO Lip-Bu Tan delivered his first COMPUTEX keynote with two key updates:

Clearwater Forest (Xeon 6+)​

  • 288 cores, Darkmont architecture
  • First Intel 18A process node data center CPU
  • Foveros Direct 3D packaging
  • Now in full production

Crescent Island AI GPU​

  • 480 GB LPDDR5x memory
  • 350 W air-cooled PCIe form factor
  • Native FP4 support, targeting agentic inference
  • Shipping H2 2026

"As AI moves into the agentic era, the CPU returns to the center of modern AI infrastructure." — Lip-Bu Tan


3. AMD Ryzen AI 400 Series Now Shipping​

AMD showcased the Ryzen AI 400 series (Zen 5 + Zen 5C hybrid + XDNA2 NPU) at COMPUTEX:

  • NPU performance: 60 TOPS, the highest in x86
  • 7 consumer SKUs + commercial PRO series
  • Multiple OEM models already available or launching soon
  • Advancing AI 2026 summit set for July in San Francisco

4. Chinese Domestic Chips Gaining Momentum​

VendorProductStatus
HuaweiAscend 950PR/950DTIn production, self-developed HBM
CambriconMLU6902 PFLOPS FP8, shipping
Moore ThreadsMTT S50001,000 TFLOPS, specs public

5. The AI PC Era: Three-Way Roadmap Comparison​

DimensionNVIDIA RTX SparkIntel Clearwater Forest + Crescent IslandAMD Ryzen AI 400
CPU Cores20-core Grace (Arm)288-core Darkmont (x86)Up to 12-core Zen5+5C
GPU/NPUBlackwell GPUCrescent Island (discrete GPU)XDNA2 NPU (60 TOPS)
AI Compute1 PFLOPSTBD60 TOPS NPU
TargetPersonal AI agentsDual-track: DC + AI PCCopilot+ PC
ProcessTSMC 4NPIntel 18ATSMC 4nm
AvailabilityJune 2026H2 2026Shipping now

This Week in AI Compute (6/1 – 6/6)​

DateEvent
Jun 1NVIDIA GTC Taipei: RTX Spark, Vera Rubin production, DGX Station for Windows
Jun 1Intel unveils Crescent Island, Clearwater Forest
Jun 2COMPUTEX 2026 opens: "AI Together"
Jun 5COMPUTEX closes: 1,500+ exhibitors, record scale
Jun 6RTX Spark confirmed June launch at $1,399

Sources: COMPUTEX Daily, Tencent News, Phoenix Technology, Xueqiu, The Silicon Review.

Computex 2026 AI Compute Card Major Events: DGX Station for Windows, Intel Crescent Island, and More Major Launches

· 4 min read
Industry Research Team

June 1-5, 2026, Taipei — Computex 2026 (Taipei International Information Technology Show) wrapped up successfully this week. With the theme "AI Together," industry giants including NVIDIA, Intel, AMD, and Qualcomm unveiled numerous AI compute products in rapid succession. Below, MirrorFrog brings you a roundup of the most noteworthy developments in the compute card space this week.

① NVIDIA DGX Station for Windows: A Desktop AI Supercomputer​

NVIDIA officially launched the DGX Station for Windows during its Computex 2026 keynote, calling it "the world's most powerful desktop AI supercomputer."

Core Specifications​

ItemSpecification
ChipGB300 Grace Blackwell Ultra Desktop Superchip
GPU Memory252 GB HBM3e (7.1 TB/s)
CPU Memory496 GB LPDDR5X (396 GB/s)
Unified Memory748 GB (NVLink-C2C interconnect)
FP4 Compute20 PFLOPS (sparse)
FP8 Compute10 PFLOPS (sparse)
NetworkConnectX-8 SuperNIC, up to 800 Gb/s
Model CapacityCan run 1 trillion parameter models
System Power1,600 W
Operating SystemMicrosoft Windows
ShippingQ4 2026

Significance: DGX Station compresses AI compute power (20 PFLOPS FP4) that previously required datacenter-class clusters into a single desktop workstation. 748GB of unified memory means developers can run models with hundreds of billions or even trillions of parameters locally, without cloud dependency.


② Intel Crescent Island: Inference-Specialized AI GPU​

At Computex, Intel disclosed detailed specifications for its next-generation datacenter AI inference GPU, Crescent Island.

ItemSpecification
MemoryUp to 480 GB LPDDR5x
Power350 W (PCIe form factor)
Precision SupportFP4/MXFP4 → FP64 (full precision coverage)
TargetAI inference workloads (Agentic Inference)
PositioningBetter price-performance than HBM solutions
ShippingH2 2026

Significance: Crescent Island represents Intel's key strategic move in the AI inference market. 480GB of massive LPDDR5x memory (non-HBM) means significantly lower cost compared to NVIDIA H200/B200 and other competing products, targeting enterprise inference deployment scenarios.


③ Intel Xeon 6+ (Clearwater Forest): First Intel 18A Datacenter CPU​

Intel also unveiled the new Xeon 6+ processor, codenamed Clearwater Forest, its first datacenter CPU built on the 18A process:

  • 288 Darkmont architecture cores
  • L2 288MB + L3 576MB cache
  • 12-channel DDR5-8000 memory
  • Foveros Direct 3D advanced packaging
  • AI Agent Era: CPU returns to the center of infrastructure

④ NVIDIA RTX Spark Ecosystem Takes Shape​

This week, the RTX Spark super chip developed in collaboration between NVIDIA and MediaTek continued to generate buzz. Multiple OEMs showcased RTX Spark-based laptop and compact desktop prototypes:

  • ASUS, Dell, HP, Lenovo, Microsoft Surface, MSI all confirmed as launch partners
  • Equipped with 20-core Grace CPU + Blackwell GPU (6144 CUDA cores)
  • AI compute 1 PFLOPS
  • Retail availability Fall 2026

⑤ Intel × Foxconn AI Infrastructure Partnership​

Intel and Foxconn announced a joint AI infrastructure initiative, covering the complete chain from chip → server → rack-scale system, targeting the datacenter market opportunity driven by surging AI inference demand.


⑥ Domestic AI Chip Developments​

According to the IDC 2025 annual report, total AI accelerator card shipments in China reached approximately 4 million units, with domestic vendors shipping approximately 1.65 million units, capturing a market share exceeding 41%. Huawei's Ascend 950 series has entered mass production and delivery, while Cambricon's MLU690 has begun shipping to internet customers.


This Week's Compute Roundup​

VendorProductHighlightTimeline
NVIDIADGX Station for Windows20 PFLOPS, 748GB unified memoryQ4 2026
NVIDIARTX Spark1 PFLOPS AI PC chipFall 2026
IntelCrescent Island GPU480GB LPDDR5x, 350WH2 2026
IntelXeon 6+ (Clearwater Forest)288 cores, Intel 18AH2 2026
Intel + FoxconnAI infrastructure partnershipChip→rack full chainStrategic partnership
HuaweiAscend 950PR/DT1 PFLOPS FP8, self-developed HBMIn mass production
CambriconMLU6902 PFLOPS FP8, 192GB HBM3EShipping

Sources: NVIDIA GTC Taipei 2026 / Computex 2026 official announcements, Intel press releases, ifeng Tech, IT Home.

NVIDIA Launches RTX Spark: AI Compute Enters the Personal Computer Era

· 3 min read
Industry Research Team

June 1, 2026, Taipei — During the Computex 2026 opening keynote, NVIDIA CEO Jensen Huang officially unveiled the RTX Spark super chip, marking NVIDIA's formal entry into the personal computer processor market dominated by Intel, AMD, Qualcomm, and Apple.

RTX Spark: The "Heart" of the Personal AI Computer​

RTX Spark was developed in collaboration between NVIDIA and MediaTek, featuring a heterogeneous package with a 20-core Grace CPU + Blackwell RTX GPU, equipped with 6144 CUDA cores. AI compute reaches 1 PFLOPS (one quadrillion floating-point operations per second), meaning personal computers now possess computing power comparable to a datacenter-class H100 GPU for the first time.

SpecificationRTX Spark
CPU20-core Grace (MediaTek collaboration, Arm architecture)
GPUBlackwell RTX (6144 CUDA cores)
AI Compute1 PFLOPS
TargetPersonal AI Agent, local LLM inference
Launch OEMsASUS, Dell, HP, Lenovo, Microsoft Surface, MSI
AvailabilityFall 2026
Form FactorLaptop SoC + compact desktop workstation

Jensen Huang's "Full-Stack AI" Strategy​

The launch of RTX Spark is a key step in NVIDIA's "full-stack AI" strategy. Jensen Huang stated during the keynote: "AI should not only run in the cloud. Everyone's computer should have the ability to run AI agents."

RTX Spark transforms NVIDIA from a datacenter GPU monopolist into a full competitor in the personal computing market. Following the announcement, shares of AMD, Intel, and Qualcomm fell accordingly.

Market Impact​

  • Intel: Personal computer AI processor business faces direct threat
  • AMD: Ryzen AI series must compete at the same level
  • Qualcomm: Snapdragon X Elite's Copilot+ PC positioning challenged
  • Apple: M-series chips are no longer the only high-performance AI PC option

Vera Rubin Platform Enters Full Mass Production​

During the same keynote, Jensen Huang also announced that the NVIDIA Vera Rubin platform has entered full mass production. Rubin R200 features a 6-chip CoWoS-L package (1× Vera CPU + 2× Rubin GPU die + I/O/HBM die), equipped with 288GB HBM4, 22 TB/s bandwidth, and 50 PFLOPS FP4 compute (sparse).

The Rubin NVL72 rack (72 Rubin GPUs + 36 Vera CPUs) will begin shipping in H2 2026.

Other Highlights from Computex 2026​

  • AMD: Showcased the MI350 series (192GB HBM3e, 5 PFLOPS FP8 dense), officially launching in June
  • Intel: Jaguar Shores publicly unveiled for the first time
  • Qualcomm: AI 200 / 300 series inference card roadmap updated
  • Domestic AI Chip Zone: Huawei, Cambricon, Moore Threads, and others showcased their latest products

Industry Significance​

The launch of RTX Spark means AI compute is no longer confined to datacenters. Individual developers, designers, and researchers will be able to run large model tasks locally that previously required cloud GPUs, potentially redefining the market landscape for personal AI computing.

The mass production of Vera Rubin further consolidates NVIDIA's absolute leadership in datacenter AI training. Together, both product lines form NVIDIA's full-stack AI computing landscape of "cloud training + personal inference."


This report is based on official NVIDIA announcements from Computex 2026 / GTC Taipei on June 1, 2026.

AWS Trainium 3 GA: 3nm Process + 4.4× Compute + 4× Efficiency + 144-Chip UltraServer

· 4 min read
Industry Research Team

On December 2, 2025, at the re:Invent 2025 conference, AWS formally GA'd its third-generation custom AI training chip Trainium 3. This is a critical upgrade to the AWS compute landscape: 3nm process, 4.4× compute improvement, 4× efficiency improvement, Trn3 UltraServer with 144 chips. This article provides a detailed analysis.