Skip to main content

Meta MTIA 400 (Custom AI Accelerator · Recommendation + GenAI Dual Mandate)

Product Overview​

MTIA 400 is the fourth generation of Meta's custom MTIA AI accelerator family, systematically disclosed for the first time at Hot Chips 2026 (2026-08-23~25).

The biggest change in this generation is the mandate expansion: from its inception, MTIA served only the single workload of Recommendation & Ranking; the MTIA 400 also takes on generative AI (GenAI), becoming a "recommendation + GenAI dual-mandate" chip. Meta's production figure given in the talk: hundreds of thousands of MTIA chips already in production.

Architecturally, the MTIA 400 is fully chipletized — 2 compute dies + 1 SoC die + 2 network dies, plus HBM. This is the first time Meta has split its custom accelerator into a multi-die combination, a generational watershed compared with the earlier single-die MTIA 200.

Core Specifications​

ParameterValue
ArchitectureMTIA 400 (2 compute dies + 1 SoC die + 2 network dies + HBM, fully chipletized)
PE Array8 × 6 processing element array (with redundant rows)
ProcessNot disclosed
FP4 Compute12 PFLOPS
FP16/BF16 ComputeNot disclosed (15x the MTIA 200)
Memory CapacityNot disclosed
Memory BandwidthNot disclosed (DRAM bandwidth 46x the MTIA 200)
On-chip SRAM BandwidthNot disclosed (5x the MTIA 200)
TDPNot disclosed
Scale-up Domain72 MTIA 400s per domain
Design PartnerBroadcom (multi-generation MTIA collaboration)
Announced2026 (disclosed at Hot Chips 2026)

⚠️ Caveat: the "15x / 46x / 5x" figures in the table are multiples relative to the MTIA 200, not absolute specifications. Meta consistently does not publish full MTIA parameters; do not mix these relative multiples with absolute compute figures.

Generational Comparison with the MTIA 200​

MetricMTIA 200MTIA 400Change
MandateRecommendation / rankingRecommendation + GenAIExpanded
PackageSingle die2 compute + 1 SoC + 2 network diesFully chipletized
FP4 ComputeNot disclosed12 PFLOPSNewly disclosed
FP16 ComputeBaseline—15x
DRAM BandwidthBaseline—46x
SRAM BandwidthBaseline—5x
Scale-up Domain—72 chips—

Key insight: DRAM bandwidth up 46x and FP16 compute up 15x — the bandwidth multiple far exceeds the compute multiple. This ties directly to the MTIA 400's new GenAI mandate: the decode stage of generative inference is a memory bandwidth bottleneck, and Meta has clearly weighted its resources toward data movement.

MTIA Roadmap (300 / 400 / 450 / 500)​

Meta's published roadmap shows four generations — MTIA 300 / 400 / 450 / 500 — iterating over the next two years:

GenerationPositioningKey Changes
MTIA 300Recommendation / GenAI / inferenceRoadmap starting point
MTIA 400Recommendation + GenAI dual mandate12 PFLOPS FP4, fully chipletized
MTIA 450GenAI / inferenceHBM bandwidth doubled vs the MTIA 400
MTIA 500GenAI / inferenceHBM bandwidth up another 50%; per-chip power up to 1700 W

Accompanying cooling roadmap: Meta targets rack power density of 80 kW or even above 120 kW, using an air-assisted liquid cooling (AALC) + Sidecar CDU architecture. Meta's core aim is not maximum cooling efficiency, but bringing liquid cooling capability to existing air-cooled data centers as quickly as possible — Sidecars can be deployed per rack, scale quickly, and isolate faults, making them better suited to rapidly launching inference services.

Vendor Information​

ItemDetails
CompanyMeta Platforms, Inc.
HeadquartersMenlo Park, California, USA
Custom Silicon RoadmapMTIA (recommendation → GenAI → inference)
Design PartnerBroadcom
AvailabilityInternal use only (not sold externally)
FoundryTSMC (specific node not disclosed)

Use Cases​

  • ✅ Recommendation systems / ranking models (MTIA's traditional home turf)
  • ✅ Generative AI inference (new MTIA 400 mandate)
  • ✅ Mixed deployment with NVIDIA / AMD GPUs (Meta pursues a multi-chip strategy)
  • ❌ External sales / third-party procurement
  • ❌ Frontier model pretraining (Meta still relies on NVIDIA GPUs)

References​