Skip to main content

Moore Threads MTT S4000 (2023)

Product Overview​

MTT S4000 is Moore Threads' large model AI computing accelerator released in December 2023, based on self-developed Qiyuan GPU architecture (third-generation MUSA core architecture), equipped with 48GB GDDR6 memory (bandwidth 768 GB/s), FP32 compute 25 TFLOPS, TF32 compute 50 TFLOPS, INT8 compute 200 TOPS, customized optimization for training, fine-tuning and inference of 100-billion-parameter large language models, combined with advanced graphics rendering capabilities, video encoding/decoding capabilities and ultra-high-definition 8K HDR display output.

Positioning: Full-function meta-computing card (training+inference integration + graphics rendering), core component of KUAE AI computing center solution.

Core Specifications​

ItemParameter
ArchitectureSelf-developed Qiyuan GPU (third-generation MUSA core architecture)
ProcessNot disclosed (estimated 7nm/6nm)
FP3225 TFLOPS
TF3250 TFLOPS
INT8200 TOPS
FP16/BF16100 TFLOPS (official documentation figures)
Memory Capacity48 GB GDDR6
Memory Bandwidth768 GB/s
TDP450 W
InterconnectMTLink (x8 Serdes, up to 56Gbps PAM4)
InterfacePCIe 5.0 x16, 4× DisplayPort
PowerCPU 8-pin × 1
ReleaseDecember 2023
Mass ProductionSince 2024
Software StackMUSA software stack (CUDA compatible)

MUSA Architecture Evolution​

ArchitectureCoreRepresentative ProductRelease
First-generation MUSAChunxiaoMTT S80/S70 (consumer)2022
Second-generation MUSAQuyuan (improved)MTT S30002023
Third-generation MUSAQiyuan GPUMTT S40002023.12

Comparison with MTT S3000​

MetricMTT S3000MTT S4000Improvement
ArchitectureSecond-generation MUSAThird-generation MUSA (Qiyuan GPU)New generation
MemoryNot disclosed48GB GDDR6Larger
BandwidthNot disclosed768 GB/sHigher
FP32Not disclosed25 TFLOPSValue disclosed
TDPNot disclosed450WData center grade
Release20232023.12Same period improvement

KUAE AI Computing Center Solution​

MTT S4000 is the core component of Moore Threads' KUAE AI computing center solution:

  • 100-billion-parameter large model training, fine-tuning, inference full-stack support
  • MTLink multi-card high-speed interconnect (x8 Serdes, 56Gbps PAM4)
  • MUSA software stack fully supports PyTorch/DeepSpeed and other mainstream frameworks
  • CUDA compatibility layer, reducing model migration costs

Application Scenarios​

  • ✅ 100-billion-parameter large model training (customized optimization)
  • ✅ Large model inference as a service (INT8 200 TOPS)
  • ✅ Graphics rendering + AI hybrid workloads (full-function GPU)
  • ✅ Video encoding/decoding (8K HDR display output)
  • ✅ Domestic AI computing center (KUAE solution)
  • ❌ Ultra-high FP16 training compute (25 TFLOPS FP32 lower than H100)
  • ❌ Ultra-large-scale clusters (MTLink TBD vs NVLink)

Product Matrix​

SeriesPositioningRepresentative Product
MTT S SeriesServer GPU (data center)S3000, S4000, S5000
MTT S Series (Consumer)Desktop GPUS80, S70
KUAEAI computing center solutionS4000 + MTLink + MUSA software stack

References​