Skip to main content

2 posts tagged with "MI400"

AMD Instinct MI400 series

View all tags

HBM4 Mass-Production Year One: Samsung Yield Breaks 80%, Three Giants Pass NVIDIA Certification, the Last Bottleneck of AI Compute Supply

· 6 min read
Industry Research Team

If 2025 was the year of HBM3E capacity ramp-up, then 2026 is year one of HBM4 mass production. With NVIDIA Vera Rubin and AMD MI400 — two generations of flagship — both betting on HBM4, this "memory on the AI chip" has for the first time become a strategic commodity that dictates the delivery pace of entire racks. The yield and certification data disclosed densely in August is rewriting the global HBM supply map.


1. Golden Yield Breakthrough: Samsung Jumps from Under 60% to 80% in Six Months​

Per South Korea's Seoul Economic Daily on August 9, Samsung Electronics' HBM4 yield officially crossed the 80% "golden yield" threshold in early August — more than four months ahead of its original year-end target.

TimelineSamsung HBM4 YieldNotes
Feb 2026 (mass production start)Under 60%Line ramp-up period
Early Aug 2026~80%Crosses the mass-production / stable-profit watershed

The semiconductor industry has long held that "80% yield is the golden yield" — it is both a yardstick of foundry competitiveness and the financial break-even point for large-scale commercial supply. The key to this leap was Samsung's breakthrough in Thermal Compression Non-Conductive Film (TC-NCF) bonding, plus the stable base of its underlying 1c DRAM yield, already above 80%. In the same period, Samsung's HBM4E reliability test yield also broke 70%.

Industry assessments suggest SK Hynix's HBM4 yield has likewise entered the 80% range. The gap between the two giants in production quality is being rapidly erased.


2. Supply Map: SK Hynix Holds 60–70% of Rubin Allocation​

At a Seoul event on June 5, Jensen Huang publicly confirmed: Samsung, SK Hynix, and Micron have all passed HBM4 certification for Vera Rubin — the first time three memory makers have simultaneously received public certification for the same platform.

But certification is just the "entry ticket" — allocation share is where the real voice lies:

Vendor2026 Rubin HBM4 Allocation (est.)Notes
SK Hynix60%–70%Based on HBM3/3E-era customer relationships and MR-MUF packaging
Samsung25%–30%Rapid share gains after yield leap
MicronRemainderLimited HBM4 exposure, relatively stable share

Counterpoint Research forecasts the 2026 HBM4 market as SK Hynix 54% / Samsung 28% / Micron 18%. Samsung has set staged catch-up targets: Q3 HBM4 revenue up 3× QoQ, HBM4 exceeding 60% of total HBM revenue in H2, and year-end overall HBM market share approaching 38%.


3. The Real Bottleneck: From Wafers to "Back-End Stacking"​

As front-end yield stabilizes, the rhythm of the AI accelerator supply chain no longer depends on "how many wafers can be made," but on the speed of back-end stacking, bonding, testing, and shipment.

  • Industry analysts rank HBM stacking as the second-most severe bottleneck in the AI chip supply chain, second only to TSMC's CoWoS advanced packaging capacity.
  • HBM accounts for roughly 25% of 2026 DRAM wafer output; each HBM wafer consumes about 3–4× the resources of a standard DRAM wafer (extra TSV and stacking steps), so every wafer redirected pulls 3–4 units of commodity memory off the spot market.
  • Samsung is considering relocating part of its legacy memory back-end lines (Cheonan, Onyang) to Vietnam to free up HBM back-end capacity — a side confirmation that back-end throughput is now the tightest link in the chain.

4. HBM4 Spec Snapshot: Generational Leap in Bandwidth and Efficiency​

SpecHBM4 (12-Hi / 16-Hi)HBM4E
Per-stack capacity36 GB / 48 GB—
Pin rate11.7–13.0 Gbps16 Gbps
Per-stack bandwidthup to 3.3 TB/sup to 3.6 TB/s
Bus width2048-bit—
Energy efficiency+40% vs HBM3E—
Thermal resistance / cooling+10% improvement / +30%—

Samsung HBM4 entered mass production in Feb 2026; its 11.7 Gbps pin rate already exceeds the 8 Gbps industry baseline required for Vera Rubin compatibility; HBM4E samples were first shipped to major customers on May 29.


5. Pricing Power Extends Into 2027: Supply Remains Tight Balance​

TrendForce judges that HBM suppliers' pricing power will run through 2027, because supply remains constrained:

  • 2027 HBM bit shipments are expected to grow 50%–60% YoY, but will still lag demand growth, keeping the market tight;
  • The industry already anticipates significant price increases;
  • For NVIDIA and AMD, a stronger Samsung means more supply options and more comfortable lead times — in a market where memory is the tightest link in AI servers, the mere existence of second and third suppliers is itself a buffer.

For entire racks, HBM cost is already the biggest driver: the Rubin Ultra rack carries an estimated price tag as high as $21 million, with HBM making up a substantial portion.


6. Lessons for China: HBM Export Controls Accelerate Domestic Iteration​

HBM is one of the core fronts of current AI chip controls. As the overseas HBM4 arms race intensifies, domestic HBM technology iteration is being pushed forward in sync — Huawei's Ascend roadmap has explicitly written "drive domestic HBM technology iteration" into its product cadence (the 950 series advances domestic HBM pairing, with the 960/970 series planned for gradual rollout in 2027–2028).

In the short term, HBM4 scarcity will directly transmit to the delivery cadence of Rubin / MI400; in the long term, whoever can lock in stable HBM4 supply holds the valve on 2027 AI compute expansion.

References​


This article is compiled from August 2026 public reports by TrendForce, Seoul Economic Daily, TechTimes, etc. HBM allocation shares and market shares are third-party estimates, not official vendor-confirmed data.

AMD MI455X Stuns at CES 2026: AI Chip Performance Up 1000x in 4 Years

· 6 min read
Industry Research Team

On January 5, 2026, on the opening day of CES 2026 (Consumer Electronics Show), AMD Chair and CEO Dr. Lisa Su unveiled in her keynote: the Instinct MI400 series AI accelerators.

The most eye-catching is MI455X — AMD's most powerful AI accelerator ever, using a 2nm + 3nm hybrid process, 432GB HBM4, with FP4 compute up to 40 PFLOPS (20 PFLOPS FP8).

Key highlights​

  • MI455X: FP4 40 PFLOPS, FP8 20 PFLOPS, 10× over MI355X
  • MI450: cost-performance version, FP4 28 PFLOPS, 288GB HBM4
  • Process upgrade: world's first AI chip with 2nm + 3nm hybrid process (GCD on 2nm, MCD on 3nm)
  • Memory upgrade: from MI350X's 288GB HBM3e to 432GB HBM4 (MI455X)
  • Bandwidth upgrade: from MI350X's 8 TB/s to 19.6 TB/s (2.45×)
  • Architecture upgrade: from CDNA 4 to CDNA 5
  • Mass production: MI455X Q4 2026, MI450 Q3 2026

Full MI400 series specs​

📌 Important correction (2026-06-16): After official spec verification, MI455X memory is 432GB HBM4 (not the earlier reported 288GB), and FP4 compute is 40 PFLOPS. Corrected herein.

ModelPositioningMemoryFP4 computeFP8 computeTDP (est.)
MI455XFlagship training+inference432GB HBM440 PFLOPS20 PFLOPS~1,000W
MI450Cost-performance training288GB HBM428 PFLOPS14 PFLOPS~800W
MI440XEnterprise inference216GB HBM425 PFLOPS12.5 PFLOPS~600W
MI430XHPC / scientific computing192GB HBM420 PFLOPS10 PFLOPS~500W
MI400XGeneral / edge inference128GB HBM412 PFLOPS6 PFLOPS~400W

Key upgrades (vs MI350 series):

  • Memory: HBM3e → HBM4, capacity +50% (432GB vs 288GB)
  • Bandwidth: 19.6 TB/s (vs MI350's 8 TB/s, +2.45×)
  • Compute: FP4 40 PFLOPS (vs MI355X's 20 PFLOPS, +2×)
  • Process: 2nm + 3nm hybrid (GCD on 2nm, MCD on 3nm)
  • Architecture: CDNA 5 (vs MI350's CDNA 4)

Performance vs. MI355X​

MetricMI355X (2025)MI455X (2026)Improvement
FP4 compute20 PFLOPS40 PFLOPS2×
FP8 compute10 PFLOPS20 PFLOPS2×
Memory capacity288GB HBM3e432GB HBM41.5×
Memory bandwidth8 TB/s19.6 TB/s2.45×
ProcessTSMC 3nm2nm + 3nm hybridNew gen
ArchitectureCDNA 4CDNA 5New gen
TDP800-1000W~1,000WFlat

Lisa Su at CES 2026:

"Four years ago, MI250's AI performance was X. Today, MI455X's performance is 1000× that. That's the pace of AI chip progress."

CDNA 5 architecture in detail​

The MI400 series adopts the CDNA 5 architecture (MI355X uses CDNA 4):

Key upgrades​

  1. Matrix Core upgrade: FP8/INT8/FP16 support, sparsity acceleration
  2. HBM4 controller: 12-layer HBM4 (vs HBM3e's 8 layers)
  3. Infinity Fabric 4.0: 50% higher die-to-die / die-to-GPU bandwidth
  4. Native sparsity support: MoE Expert-Parallel optimization
  5. Long-context optimization: 1M+ token KV Cache acceleration

vs. NVIDIA Blackwell / Rubin​

MetricAMD MI455XNVIDIA B200NVIDIA Rubin R200 (2026 Q4)
FP4 compute40 PFLOPS20 PFLOPS (45 sparse)~40 PFLOPS (est.)
FP8 compute20 PFLOPS10 PFLOPS (22.5 sparse)~20 PFLOPS (est.)
Memory432GB HBM4192GB HBM3e288GB HBM4
Memory bandwidth19.6 TB/s8 TB/s13 TB/s
TDP~1,000W700-1000W~1,000W
Process2nm + 3nm hybridTSMC 4npTSMC 3nm
Mass production2026 Q42024 Q42026 Q4
Software ecosystemROCmCUDACUDA
StrengthMemory capacity, open ecosystemMost mature ecosystemNext-gen architecture
WeaknessSoftware ecosystem gapSmaller memoryNot yet launched

Conclusion: MI455X leads B200 in FP4/FP8 compute and memory capacity/bandwidth, but software ecosystem remains a weak point. Versus Rubin R200, paper specs are close, but Rubin has the CUDA ecosystem moat.

Production timeline​

TimeEvent
June 12, 2025MI400 series specs first announced at Advancing AI
January 5, 2026MI455X/MI450/MI440X formally launched at CES 2026
2026 Q3MI450 sampling begins
2026 Q4MI455X mass production
2026 Q4MI440X (enterprise inference) launched
2027 Q1MI430X/MI400X (HPC/edge inference) launched
2027MI500 series (next gen)

AMD AI chip roadmap (2025-2027)​

TimeProductProcessNotes
Q4 2024MI325XTSMC 5nmHBM3e upgraded
Q3 2025MI355X (MI350 series)TSMC 3nmCDNA 4, 288GB HBM3e
Q4 2026MI455X (MI400 series)2nm + 3nm hybridCDNA 5, 432GB HBM4
Q1 2027MI500 seriesTSMC 2nm (est.)Next gen, further gains

Software ecosystem: ROCm's progress and challenges​

✅ Progress​

  • PyTorch 2.5+: native MI300X/MI455X support
  • Hugging Face Transformers: official AMD GPU support
  • vLLM 0.8+: MI300X inference support (experimental)
  • JAX: AMD adapting (vs Google TPU)

⚠️ Challenges​

  • Framework optimization: PyTorch on AMD GPUs still below NVIDIA
  • Operator coverage: some niche operators need hand-written HIP
  • Multi-card communication: RCCL (vs NCCL) still lags
  • Developer ecosystem: tutorials, cases, community activity far below NVIDIA

Competitive comparison​

VendorProductFP4 computeMemoryMass productionStrengthWeakness
AMDMI455X40 PFLOPS432GB HBM42026 Q4Largest memory, open ecosystemSoftware gap
NVIDIAB20020 PFLOPS192GB HBM3e2024 Q4Most mature ecosystemSmaller memory
NVIDIARubin R200~40 PFLOPS288GB HBM42026 Q4Next-gen architecture, CUDAExpensive
HuaweiAscend 910C~1.6 PFLOPS64GB HBM2026 Q2China-localizedExport-controlled
GoogleTPU 8t~9.2 PFLOPS~256GB HBM3eLate 2027Gemini-integratedGoogle Cloud only

Industry impact​

1. Impact on NVIDIA​

On paper, AMD MI455X has already caught up to B200 (FP4 40 PFLOPS vs 20 PFLOPS), even leading substantially in memory capacity (432GB vs 192GB).

But:

  • NVIDIA has the CUDA ecosystem moat
  • NVIDIA has the Vera Rubin platform (full solution, 2026 Q4)
  • AMD only sells cards/nodes, NVIDIA sells AI factories
  • MI455X mass production (2026 Q4) coincides with Rubin R200 — head-on competition

2. Pressure on domestic chips​

MI455X's launch means: mainstream international AI chips enter the 2nm + HBM4 era in 2026.

Domestic chips (Huawei Ascend, Cambricon, MetaX, etc.) need to:

  • Catch up to 5nm + HBM3e by 2026-2027
  • Otherwise the gap widens from "1 generation" to "2 generations"

3. Significance for cloud providers​

MI455X gives cloud providers a second option beyond NVIDIA:

  • Microsoft Azure: already deployed MI355X, may follow with MI455X
  • Google Cloud: in-house TPU, won't use AMD
  • Amazon AWS: in-house Trainium/Inferentia, won't use AMD
  • Alibaba Cloud, Tencent Cloud: may procure MI455X as NVIDIA alternative

References​


This article is compiled from AMD CES 2026 official announcements, Baidu Baike, and Zhihu on-site reports; specs verified against official sources. Updated 2026-06-16: corrected MI455X memory (288GB → 432GB) and compute (FP8 6 PFLOPS → FP4 40 PFLOPS).