Skip to main content

Enflame CloudBlaze S60

Enflame's third-generation AI inference accelerator, based on the self-developed GCU320 (Suiyuan 320) chip, released in March 2024, with a maximum power of 300W, targeting large-scale data center deployment and large-model inference (full-spec DeepSeek all-in-one appliance solutions).


Core Specifications​

SpecificationValue
ArchitectureGCU320 (Suiyuan 320)
ProcessNot disclosed
TDP300 W (official product manual, maximum power)
Memory48 GB (official product manual)
Memory Bandwidth672 GB/s (official product manual)
Precision SupportFP32 / FP16 / BF16 / INT8 (peak per-precision compute not published by the vendor)
InterfacePCIe Gen5 x16, full-height full-length dual-slot
Video DecodeUp to 256 channels
ECC / SecurityECC, Secure Boot, SR-IOV (4 VF)
Release2024-03
PriceNot disclosed

📌 Data correction (2026-09 cross-validation): This page previously recorded "1.6 TB/s bandwidth / FP16 100 TFLOPS (estimated)", which does not match the official manual. Enflame's official S60 product manual confirms: 48GB memory, 672 GB/s, PCIe Gen5 x16, maximum power 300W, support for FP32/FP16/BF16/INT8, but does not publish peak per-precision compute — corrected per the "no estimates when undisclosed" principle.


Technical Highlights​

  • Third-generation inference card: Based on the Suiyuan 320 (GCU320) chip, Enflame's third-generation AI inference product
  • Large-model inference optimization: Optimized for large language model inference scenarios such as LLaMA and DeepSeek; the vendor offers an all-in-one appliance supporting the full 671B model
  • Search/ads/recommendation support: Sustains tens of billions of daily calls for Tencent's search, ads, and recommendation workloads (per the IPO prospectus)
  • Easy migration: Broad model coverage and strong usability, supporting smooth migration from NVIDIA GPUs
  • High-density deployment: Maximum power 300W with passive air cooling, suitable for large-scale data center deployment

Product Positioning​

The CloudBlaze S60 is Enflame's new-generation AI inference accelerator for large-scale data center deployment, targeting the NVIDIA L4/L40. As a third-generation product, the S60 significantly improves on memory capacity (48GB) and video decoding (256 channels) over the previous-generation CloudBlaze i20 (16GB), and is a representative product of Enflame's "card–model–system" closed loop (full DeepSeek model adaptation).


Application Scenarios​

  • Large language model inference (LLaMA, DeepSeek, ChatGLM, etc.)
  • Search, advertising, and recommendation system inference
  • Computer vision inference (CV, 256-channel video decode)
  • Natural language processing inference (NLP)
  • Large-scale data center inference deployment

Reference Price​

  • Official pricing: Not disclosed


References​