Skip to main content

Huawei Ascend 960

Product Overview​

Huawei Ascend 960 is the fifth-generation Ascend AI chip, officially announced on September 17, 2026 by Rotating Chairman Wang Tao at HUAWEI CONNECT 2026 (Shanghai). The biggest difference from its predecessor: the 960 is no longer a roadmap teaser, but a complete commercial combination of dual versions (960DT training / 960PR inference) + dual SuperPods (Atlas 860 air-cooled / Atlas 960 liquid-cooled), and both chips are ready ahead of schedule — the 960DT three quarters early, the 960PR one quarter early.

Also announced: the world's first SuperPod using NPO (near-package optics) technology — the Ascend 960 SuperPod: 4096 interconnected cards per node, 8 EFLOPS FP8, 1PB HBM, built on the "UnifiedBus + Hi-ONE optical engine". Huawei also confirmed the one-chip-per-year cadence: Ascend 970 in 2028, Ascend 980 in 2029.

⚠️ Information note: single-chip specs on this card follow the official HC2026 disclosures. In August 2026, the Digital China Summit reported that "the Atlas 960 SuperPoD supports 15,488 Ascend cards"; in September the official figure was 4096 cards per SuperPod (a revision of the earlier 15,488 claim; this card follows the official press release).

Core Specifications​

The following are the Ascend 960DT (training version, officially disclosed values):

ParameterValue
ArchitectureDa Vinci v6 (Ascend 6th generation)
FP8 Compute2 PFLOPS
FP4 Compute4 PFLOPS
Memory288 GB (in-house HBM, max configuration)
Memory Bandwidth9.6 TB/s
TDP700 W (presumed, not officially disclosed)
Debut2027 Q2

📌 Compared with the 950 series (FP8 1 PFLOPS / FP4 2 PFLOPS), every key metric of the 960 doubles: FP8 2→x2, FP4 4→x2, memory 288GB vs 144GB (950DT), bandwidth 9.6TB/s vs 4TB/s. Huawei calls this continuous doubling path "Tao's Law".

⚠️ TDP is a third-party estimate; the official figure has not been published. The existence of an air-cooled version (Atlas 860) implies per-card power must stay within what air cooling can handle.

960DT vs 960PR Dual Versions​

Dimension960DT (training)960PR (inference)
FP8 Compute2 PFLOPSNot separately disclosed
Ready2027 Q1 (3 quarters early)2027 Q3 (1 quarter early)
Launch2027 Q22027 Q3
Paired SuperPodAtlas 860 (air-cooled)Atlas 960 (liquid-cooled)
PositioningTen-trillion-parameter model training + high-concurrency inferenceLarge-scale inference deployment

Ascend 960 SuperPod (world's first NPO SuperPod)​

MetricAscend 950 SuperPodAscend 960 SuperPod
Cards per node1024 cards (shown at WAIC 2026)4096 cards
FP8 Compute8 EFLOPS (8192-card full config)8 EFLOPS (single node)
FP4 Compute16 EFLOPS (8192-card full config)16 EFLOPS (single node)
Total HBM1152 TB (8192-card full config)1 PB (4096 cards)
Interconnect RTT3 μs2 μs (domain-wide D2D)
Optical interconnectUnifiedBus 2.0NPO optical engine Hi-ONE (industry's first mass production)

NPO (Near-Package Optics) Key Figures​

  • Hi-ONE optical engine: 7.2T transmission capacity per engine, the industry's first mass-produced NPO product, the highest current transmission capability, and the only one with a built-in light source
  • 5500 Hi-ONE units replace 48000 800G optical modules, cutting power by more than 550 kW
  • System mean time between failures doubled, overall availability at 99.8%
  • Why NPO over CPO: NPO keeps optical engines independent while enabling short-distance opto-electric handoff, avoiding CPO's reliability, manufacturability, and serviceability problems, and preserving the existing optical module industry ecosystem
  • MFU gain (Huawei Markov Lab simulation): in a 100k-card cluster of 4K SuperPods, MFU improves 2.75x versus 8-card server networking — in traditional architectures, intra-cluster communication takes over 40% of training time

Cluster Scaling Path​

  • Multiple 960 SuperPods interconnected via UnifiedBus or RoCE: up to 512,000 cards (two-tier CLOS four-plane fabric)
  • With multi-rail topology: up to 1 million cards in an Ascend SuperPod cluster
  • Kunpeng SuperPod upgraded in step: all-optical networking up to 4096 nodes, 256TB unified memory pool

Annual Cadence Roadmap (Tao's Law)​

YearChipStatus
2026Ascend 950DT / 950PRIn volume commercial use
2027Ascend 960DT / 960PRReady ahead of schedule, Q2 / Q3 launches
2028Ascend 970Announced, compute specs continue to double
2029Ascend 980Announced, memory bandwidth / capacity / interconnect bandwidth all up sharply

Ecosystem and Deployment Status​

  • Deployment base: over 1000 Ascend 910C SuperPods deployed and the Ascend 950 SuperPod in volume commercial use — the 960 climbs from this two-step foundation
  • CANN ecosystem crosses the tipping point: external developers now account for 61% (surpassing internal teams for the first time), monthly active developers past 5200, and CANN is fully open-sourced with regular community operations
  • Kunpeng + Ascend: over 7.8 million developers gathered, with 20,000+ industry solutions incubated
  • Storage companion: the OceanStor M900 cluster (PB-scale KV cache one hop away over the UnifiedBus) released in step

Competitor Comparison​

MetricAscend 960DTNVIDIA B200NVIDIA Rubin R200AMD MI455X
FP8 compute2 PFLOPS4.5 PFLOPS12.5 PFLOPS2.3 PFLOPS
FP4 compute4 PFLOPS9 PFLOPS50 PFLOPS4.6 PFLOPS
HBM capacity288 GB192 GB288 GB288 GB
HBM bandwidth9.6 TB/s8 TB/s16 TB/s12 TB/s
ProcessSMIC domestic (unconfirmed)TSMC 4NPTSMC 4NPTSMC 3NM
Launch2027 Q2AvailableAvailableAvailable

The 960's single-card absolute performance still lags NVIDIA's flagship by a generation (Rubin R200 FP8 is about 6x higher); Huawei's strategy is to close the gap with SuperPod system-level capabilities (unified memory addressing, 2.75x MFU, NPO optical interconnect) — evaluating domestic compute should not come down to a single-chip PFLOPS comparison.

Use Cases​

  • ✅ Ten-trillion-parameter LLM training (4096-card SuperPod + unified memory addressing)
  • ✅ High-concurrency inference (inference latency down ~70% and training throughput up ~2.3x vs the predecessor, official figures)
  • ✅ Standard server room deployment (Atlas 860 air-cooled SuperPod, no liquid cooling retrofit needed)
  • ❌ Single-card purchase (delivered as SuperPod / cluster)
  • ❌ Deployment within 2026 (chips launch in 2027)

Vendor Information​

ParameterValue
ManufacturerHuawei Technologies Co., Ltd. (HiSilicon)
Websitehttps://www.hiascend.com
Announced2026-09-17 (HUAWEI CONNECT 2026, Shanghai)
960DT launch2027 Q2
960PR launch2027 Q3

References​