Skip to main content

3 posts tagged with "Power & Cooling"

Power efficiency, thermal design and liquid cooling for AI systems

View all tags

Intelligent Compute to Hit 9,800 EFLOPS by 2030: Five Hard Targets in the ICT Industry "15th Five-Year" Plan

· 6 min read
Industry Research Team

This article is based on the MIIT "15th Five-Year" Plan for ICT Industry Development (issued in September 2026) and official interpretations by the People's Post and Telegraph News, The Paper, Huaxia Times, and other outlets.

In early September, the MIIT issued the "15th Five-Year" Plan for ICT Industry Development, drawing a roadmap for the ICT industry over the next five years with 13 major indicators and 26 key tasks. For the AI compute industry, this is the most substantive policy document — we have picked out five hard targets and break down their industry implications one by one.

Target 1: Intelligent Compute from 1,590 to 9,800 EFLOPS, More Than 5x Growth in Five Years​

This is the standout number in the entire plan. For reference:

  • As of the end of June 2026, national intelligent compute capacity had reached 2,185 EFLOPS, up 177% year over year — the first half alone overshot the year-end 2025 target (1,590 EFLOPS)
  • The country has built 42 10,000-card-class intelligent computing clusters; in July 2026, the first fully domestic 100,000-card-class AI compute cluster, the Sugon 8000, was completed in Zhengzhou
  • Huawei's rotating chairman Wang Tao's assessment: super-node clusters of more than 100,000 cards will become basic configuration by 2027

Going from 2,185 to 9,800 means that over the next four years the country must build 3.5 times the current installed base of intelligent computing facilities. The plan also calls for "orderly deployment of 10,000-card, 100,000-card and larger intelligent computing clusters" and "stepping up efforts to adapt domestic compute chips" — this is the most certain demand base for domestic AI chips over the next five years.

Target 2: Cumulative Information Infrastructure Investment of RMB 3.8 Trillion​

Some have compared this with the "14th Five-Year" figure of RMB 3.7 trillion and concluded that "investment has peaked." Official interpretations explicitly reject this claim: incremental capital is shifting from "scale expansion" to "quality-and-efficiency gains," precisely targeted at intelligent computing infrastructure, 10G optical networks, 6G, and other new tracks. The industry chain pull effect is changing accordingly — AI servers, high-speed optical interconnects, domestic compute chips, and new smart terminals were named as beneficiaries across the whole chain.

Target 3: PUE of New Large Compute Facilities Reduced Below 1.2​

At the end of the "14th Five-Year" period this figure was 1.25, and it must fall another 0.05 within five years — it doesn't sound like much, but with AI servers running at high load year-round and per-rack power density generally exceeding 20kW, every point of PUE reduction is a hard fight. The technical path laid out in the plan is very clear: liquid cooling.

  • The "15th Five-Year" Plan for ICT Industry Development: guide compute facilities to adopt high-efficiency energy-saving equipment and advanced technologies such as liquid cooling
  • The "15th Five-Year" Plan for Electronic Information Manufacturing Development (jointly issued by the MIIT and the NDRC on September 15): lists liquid cooling alongside high-bandwidth memory pooling and all-optical switching as key technologies to be broken through

Industry-side data confirms the trend: according to Omdia, liquid cooling's share of the global data center cooling market climbed from about 20% in 2024 to about 37% in 2025, with penetration expected to exceed 45% by 2029. The ceiling on compute expansion is shifting from chip supply to power supply, and the energy-efficiency constraint of "compute up, energy down" will be the norm for the next five years.

Target 4: Advanced Storage Capacity of 1,700 EB, More Than 2x Growth​

While intelligent compute grows 5x, advanced storage grows more than 2x — storage-compute coordination is given equal weight. The logic: training checkpoints for large models, and the intermediate states and contextual memory of agent inference, are all stored in layers within high-performance storage. As of the end of June, national storage capacity totaled about 2,021 EB, of which advanced storage accounted for about 32%; the 1,700 EB advanced storage target for 2030 means structural upgrading matters more than total-volume growth.

A direct implication for chip selection: KV cache and long context are eating the memory budget — when evaluating AI servers, memory capacity and memory bandwidth should carry more weight than peak compute.

Target 5: The Agent Interconnection Network Enters the Plan for the First Time​

This is the most forward-looking part of the plan: the "agent interconnection network" gets its own dedicated column, deploying four areas of work — building an agent network identifier system, accelerating the construction and application of agent network infrastructure, promoting global interconnection, and establishing a space governance system. It also explicitly states "launching 6G commercial use in a timely manner."

As AI shifts from "applications for people" to "agent infrastructure," the role of the communications network upgrades from connecting people to connecting agents — providing a national-level narrative for the distributed deployment of inference compute (edge inference, compute scheduling).

Three Judgments for Compute Practitioners​

JudgmentBasis
Domestic compute demand is highly certainThe plan explicitly states "stepping up efforts to adapt domestic compute chips" + the 100,000-card cluster build cycle has begun
Energy-efficiency targets become hard siting constraintsPUE below 1.2 + liquid cooling named as a key technology; high-density liquid cooling solutions take priority
Inference compute sinks to the edge on demand"Deploy inference compute facilities as needed for scenarios"; edge and regional compute centers enjoy a policy window

Under a 5x compute expansion target, what has always been scarce is not planning but chips, power, and delivery capability. For buyers, the supply window remains tight; for solution selection, we recommend using the TCO Calculator to convert PUE differences into electricity costs — the gap between 1.25 and 1.2 amounts to tens of millions on the electricity bill of a 10,000-card cluster.

Summary​

The "15th Five-Year" Plan writes intelligent compute into a national-level project: 5x compute, 2x storage, RMB 3.8 trillion in investment, PUE 1.2, and the agent internet. Looking back five years from now, the wave of domestic 100,000-card clusters in the second half of 2026 may well prove to be the starting point of this curve.

(Plan data is cited from MIIT documents and official media interpretations; market data is cited from statistics by CAICT, Omdia, and other institutions.)

Vera Rubin 全面量产:100% 全液冷 + 800V 直流供电,AI 数据中心基础设施范式重构

· 6 min read
Industry Research Team

2026 年 9 月初,供应链信息确认:英伟达 Vera Rubin 平台已于 8 月正式量产、9 月启动批量出货,无延期、无卡顿。与 Blackwell 迭代初期的产能波折不同,这次量产节奏异常平稳——谷歌云、微软 Azure、CoreWeave、甲骨文云等头部云厂商已启动机架部署。但真正值得产业记住的,不是"又一代 GPU 量产了",而是 Rubin 把液冷从"可选配置"变成了"硬性前置条件"。


1. 量产节奏:史上最平稳的一次平台切换​

根据产业链调研与券商跟踪信息:

  • 2026 年 8 月:Vera Rubin 正式量产;
  • 2026 年 9 月:批量出货启动;
  • 2026 下半年:CoreWeave、谷歌云、微软 Azure、甲骨文云机架部署落地;
  • 2026 年:上代 GB 架构机柜出货量有望达 6 万台(同比翻倍);
  • 2027 年:GB 与 Rubin 两代平台合计出货体量有望接近 10 万台,Rubin 新机柜远期产能目标为每天 1000 个 NVL72 机柜。

需求侧同样在加码:华尔街报告披露,英伟达管理层表示 FY28 同比增长 70% 的目标并非需求上限——若供应不受限,增速可能超过 100%。当前主要约束已从需求端转向先进晶圆与 HBM 供应。

2. 单卡 2300W:风冷时代的终结​

Rubin 平台与前代最根本的差异不在算力,而在功耗密度:

指标H100GB300Rubin
单 GPU TDP700W~1400W2300W
机柜功耗~40kW~140kW190–230kW
散热方案风冷为主风液混合100% 全液冷
供电架构48V48V800V 高压直流

单芯片 TDP 从 700W 升至 2300W、单机柜功率密度突破风冷物理极限——这意味着 液冷不再是高端算力的选配升级,而是运行 Rubin 服务器的先决条件。英伟达官方将 Rubin 全液冷架构定义为"数据中心历史上最重要的能效突破之一",并已写入 DSX AI 工厂参考设计:所有跟随英伟达技术路线的云厂商和数据中心运营商,都必须采用全面液冷方案。

三个关键架构变化:

  1. 无风扇整机:GPU、CPU、交换机、DPU 全部器件强制采用直接冷板式液冷,45℃ 温水冷板成为出厂标配;
  2. 液冷边界延伸:散热覆盖范围从 GPU 冷板延伸至 CPU、DPU、交换机乃至光模块(液冷 Cage/鼠笼开始从"可选"变"刚需"),整套液冷系统价值量较 GB300 提升约 40%;
  3. 800VDC 供电:替代传统 48V 机架配电,整机电源 BOM 价值增长 30% 以上,PSU 电源模块从 5.5kW 向 18.3kW 迭代,固态变压器、高压直流 CDU 成为数据中心新增核心设备。

3. 对产业链的三重传导​

第一重:液冷从"配套"变"主角"。 2026 下半年以小规模部署验证为主,真正的放量窗口在 2027 年——Rubin 机架大规模铺货后,冷板、快速接头、CDU、液冷泵进入业绩兑现期。台系供应链 7 月数据已率先验证:AVC 奇鋐 7 月营收 185.9 亿新台币创历史新高(同比 +57.4%),双鸿、健策 7 月同比分别 +116.7%、+91.0%。

第二重:国产液冷供应链进入核心 BOM。 国内厂商由外围冷源和代工环节逐步进入芯片平台、服务器 ODM 和海外云厂商供应体系,替代路径从 Manifold、管路推进至高可靠快接头和冷板。英维克 26H1 海外收入占比 71.4%,飞龙股份液冷泵小功率平台订单超 5 万台——液冷全核心零部件自主可控正在成为现实。

第三重:供电与散热边界融合。 800VDC 架构下,电源模块、PDB 配电单元、高速交换芯片自身发热也达到很高水平,部分电源组件同样需要液冷辅助散热——电源与温控两条产业链正在合并成一条。

4. 需求矩阵扩容:云厂商之外,太空算力入场​

Rubin 的客户矩阵已从传统云厂商扩展至三个层次:

  • 全球云厂商:谷歌云、Azure、甲骨文云、CoreWeave;
  • AI 科技巨头:马斯克公开披露 2027 年 8GW 超大规模 IDC 建设规划;SpaceX 将 Vera Rubin 架构定义为"最优 AI 计算架构",计划地面与太空双向部署,支撑 "Starmind" 卫星算力项目;
  • 主权与边缘:远期 Rubin Ultra 及 2027 年后更高功耗机型单机柜有望冲击 600kW+。

普华永道预计全球数据中心累计投资到 2035 年将达 31.6 万亿美元。AI 基础设施建设的确定性,已经从"是否建设"变成"多快建设"。

5. 对采购方的启示​

  • 机房规划前置:2027 年起采购 Rubin 级算力,液冷改造(单千瓦改造成本较高)或按全液冷标准新建,必须在预算周期一开始就纳入;
  • 看 PUE 也看水温:45℃ 温水直冷允许更高进水温度,可利用自然冷源压低 PUE——选址时人工冷源依赖度成为新的评估维度;
  • 供应商组合即风险对冲:HBM 与先进封装供应是当前核心瓶颈(详见本站 HBM4 竞速分析),供应链多元化比单点性能更重要。

相关链接​

参考资料​


本文基于 2026 年 9 月初供应链调研、券商研报与英伟达官方披露整理。出货量与功耗数据为产业链预估口径,实际以英伟达及客户正式披露为准。

AI Hardware Enters the "Era of Deployment": Five Major Shifts of 2026 and the Rules for Survival

· 9 min read
Industry Research Team

In 2026, the AI hardware market is undergoing a fundamental shift from the "training race" to "deployment as king." As large models move from technology demos to large-scale commercial deployment, hardware form factors, technology roadmaps, and the competitive landscape are undergoing systematic change.

Publisher: CSHIA Research (中智盟咨询) Author: Zhou Jun

Trend 1: Shift in compute demand structure — inference becomes the main engine of growth​

The biggest change in the 2026 AI hardware market is the shift in the center of gravity of compute demand from training to inference.

According to market data:

  • In 2026, global AI inference compute demand is expected to grow over 60% year-over-year
  • Inference compute will exceed training compute for the first time, becoming the dominant workload of AI infrastructure

This shift stems from AI applications moving from "model development" into the "large-scale deployment" stage — enterprises no longer train large models frequently, but instead transform AI capability into real business value through high-frequency inference calls.

Key manifestations​

  1. Inference chip market explosion: Shipments of dedicated inference chips (ASICs) are expected to grow 129%, with their share of AI servers rising from under 20% in 2025 to 27.8%.

  2. Cost structure optimization: NVIDIA's Rubin platform reduces inference token cost to 1/10 of the previous generation, pushing inference applications from "luxury" to "commodity."

  3. Workload characteristics change: Inference tasks show "high-frequency, long-pipeline, low-latency" characteristics, demanding higher real-time responsiveness from hardware.

Latest GTC 2026 developments (June 1, Taipei)​

NVIDIA CEO Jensen Huang announced several major inference compute advances at GTC 2026 Taipei:

  • Vera Rubin platform enters full production: The NVL72 rack system delivers agentic throughput 10× that of the previous-generation Grace Blackwell, designed for Agentic AI
  • Vera CPU officially launched: 88-core Armv9.2 custom Olympus architecture, highest single-thread IPC in the world, 1.5TB LPDDR5X memory, 1.2 TB/s bandwidth, native FP8 support
  • RTX Spark AI PC chip: Co-developed with MediaTek (codename N1X), Blackwell-architecture GPU with 1 PFLOP AI compute, 128GB unified memory, TSMC 3nm, reshaping the Windows PC ecosystem
  • AI Factory platform DSX: Four components — DSX Sim (digital-twin simulation), DSX OS (resource orchestration), DSX MaxLPS (power optimization), DSX Flex (grid coordination)

This trend means the competitive focus for hardware vendors is no longer "peak single-card compute" but "inference energy efficiency" and "system-level optimization capability."


Trend 2: Edge and on-device AI — the scaled deployment of compute moving downstream​

2026 is the pivotal year for edge AI hardware moving from proof-of-concept to scaled deployment.

As cloud inference cost pressure rises and privacy compliance requirements tighten, compute is accelerating its migration toward data sources, spawning explosive growth in hardware form factors such as edge servers, AI terminals, and smart devices.

Three deployment scenarios​

ScenarioHardware formCore characteristics2026 market size forecast
Edge serversCompact cabinets, edge compute nodesPower density 40-80kW/cabinet, liquid cooling supportedGlobal shipments grow 28%
AI terminalsAI phones, AI PCs, smart glassesOn-device NPU compute 60+ TOPS, offline inference1.5 billion units shipped
IoT devicesSmart cameras, sensors, robotsLow-power chips, real-time responseMarket size exceeds $1.5 trillion

Technology breakthroughs​

  1. On-device model compression: Through quantization, distillation and other techniques, models with tens of billions of parameters are compressed to run on-device.

  2. Heterogeneous compute architecture: CPU+NPU+GPU coordination maximizes performance under power constraints.

  3. Memory bandwidth optimization: Application of HBM technology in edge chips alleviates the "memory wall" problem.

The edge AI explosion means hardware design must balance "performance density" with "power efficiency," and traditional general-purpose chips face specialization challenges.


Trend 3: Dedicated chips and heterogeneous computing — breaking the monopoly of a single architecture​

In 2026 the AI chip market will show a "one superpower, many strong players, a hundred flowers blooming" competitive landscape.

Although NVIDIA maintains its advantage in training, in segmented markets such as inference, edge, and specific scenarios, dedicated chips (ASICs) and heterogeneous computing solutions are rising rapidly.

Major technology roadmap comparison​

Chip typeRepresentative vendorsCore advantageApplicable scenarios
General-purpose GPUNVIDIA, AMDMature ecosystem, flexible programmingCloud training, complex inference
Dedicated ASICGoogle TPU, CambriconHigh energy efficiency, cost advantageLarge-scale inference, specific algorithms
Compute-in-memoryMultiple startupsBreaks the "memory wall," low latencyEdge inference, real-time processing
FPGA/DPUXilinx, HuaweiReconfigurable, high flexibilityNetwork acceleration, data preprocessing

Market landscape changes​

  1. Domestic substitution accelerates: China's AI chip vendors raise their share in inference, edge and other scenarios to over 30%.

  2. Open-source ecosystem rises: Open-source frameworks such as ROCm and OpenML lower the barrier to dedicated-chip development.

  3. Chiplet technology popularizes: Integrating chips of different process nodes through advanced packaging achieves a balance of performance and cost.

  4. GTC 2026 new products accelerate deployment (June 1, Taipei):

    • Vera Rubin platform: NVL72 rack system, agentic throughput 10× Grace Blackwell
    • Vera CPU: 88-core Olympus custom architecture, designed for Agentic AI low latency
    • RTX Spark: In partnership with MediaTek and Microsoft, reshaping the Windows PC ecosystem, 1 PFLOP AI compute
    • Nemotron 3 Ultra: SSM+MoE hybrid architecture, 5× faster inference, 30% lower cost

The core logic of this trend is: no single chip can dominate all AI scenarios; scenario fragmentation spawns technology-roadmap diversification.


Trend 4: Energy efficiency and thermal management — from technical challenge to business bottleneck​

As AI chip power consumption breaks the kilowatt level (NVIDIA Rubin GPU reaches 2300W), energy efficiency and thermal management have been upgraded from "supporting technology" to "core bottleneck."

In 2026, single-cabinet power density will exceed 240kW, traditional air cooling completely fails, and liquid cooling changes from "optional" to "mandatory."

Key data​

  • Power cost share: The share of power cost in AI data center operating cost rises from 15% to 35%
  • Thermal value increases: A single GB300 server's liquid-cooling components are worth about $50,000, 15-20% of hardware cost
  • PUE optimization: Liquid-cooled data centers can bring PUE down to under 1.1, but upfront investment rises 30%

Technology evolution directions​

  1. Tiered liquid cooling: Cold-plate (mainstream), immersion (high density), two-phase cooling (frontier)

  2. Power architecture upgrade: From 12V to 48V/800V high-voltage DC, reducing conversion losses

  3. Intelligent thermal management: AI predictive cooling, dynamically adjusting cooling strategy based on load

This trend means a hardware vendor's competitiveness depends not only on chip performance but more on "system-level energy efficiency optimization capability"; the importance of supporting technologies such as thermal management, power delivery, and cabinet design rises substantially.


Trend 5: AI-native hardware ecosystem — from "compatibility" to "reconstruction"​

In 2026, AI hardware is undergoing a paradigm shift from "adapting to AI" to "built for AI."

Traditional general-purpose hardware architectures struggle to meet the unique demands of AI workloads, spurring the rise of AI-native hardware design philosophy.

Three reconstruction directions​

1. Compute architecture reconstruction​
  • Memory hierarchy optimization: HBM4 memory bandwidth breaks 3TB/s, compute-in-memory architecture reduces data movement
  • Interconnect upgrade: NVLink 6.0 reaches 1.8TB/s bandwidth, supporting direct GPU-to-GPU communication
  • Heterogeneous integration: Through advanced packaging, CPU, GPU and memory are stacked to boost bandwidth and reduce latency
2. Software-defined hardware​
  • Reconfigurable logic: FPGA and DPU support dynamic algorithm loading, adapting to different AI models
  • Compiler optimization: AI compilers (e.g., MLIR) automatically optimize hardware resource allocation
  • Hardware abstraction layer: Unified programming interfaces shield underlying hardware differences
3. Ecosystem co-evolution​
  • Model-hardware co-design: Large-model architectures account for hardware constraints (e.g., sparsification, quantization)
  • Open-source hardware design: Application of RISC-V in AI chips lowers the development barrier
  • Vertical integration: Cloud vendors' self-developed chips (e.g., AWS Graviton, Google TPU), software-hardware co-optimization

The essence of this trend is: the characteristics of AI workloads (matrix operations, high parallelism, memory sensitivity) are redefining hardware design principles, and the universality advantage of traditional x86 architecture is weakened in AI scenarios.


Key Conclusions and Outlook​

The inference demand explosion drives edge deployment, edge scenarios spawn dedicated chips, high power consumption forces an energy-efficiency revolution, and all changes ultimately point to the reconstruction of the AI-native hardware ecosystem.

The core driver of this round of change is AI moving from "technology demo" to "commercial deployment"; hardware must satisfy the industry requirements of "scale, low cost, high reliability."

2. Opportunity windows for industry participants​

For industry participants, the opportunities in 2026 lie in:

  • ✅ Capture the inference dividend: Deploy inference-specific chips and system optimization
  • ✅ Deepen vertical scenarios: Customize hardware solutions for specific industries/applications
  • ✅ Break the energy-efficiency bottleneck: Liquid cooling, high-voltage DC, AI thermal management and other technologies
  • ✅ Build an open ecosystem: Open-source frameworks, open standards, cross-industry collaboration

Vendors that can provide "end-to-end solutions" rather than "single-point chips" will gain an advantageous position in this reshuffle.

3. Dynamic adjustment and continuous evolution​

The above analysis is based on early-2026 market data and industry forecasts; actual development may adjust dynamically due to factors such as technology breakthroughs, policy adjustments, and market demand changes.


Industry Implications​

2026 is a watershed year for the AI hardware industry:

  • From "compute race" to "deployment as king"
  • From "single-point breakthroughs" to "system optimization"
  • From "general-purpose architecture" to "dedicated customization"
  • From "performance first" to "energy efficiency balance"

Vendors that can keenly capture trends, rapidly adjust strategy, and sustain technological innovation will seize the initiative in the AI hardware "era of deployment."


References:

  • CSHIA Research, "2026 AI Hardware: Five Transformations and the Rules for Survival"
  • "AI Hardware Enters the 'Era of Deployment'," Sohu Tech, February 10, 2026