Skip to main content

Sophgo BM1684X

Product Overview​

The BM1684X is the fourth-generation Tensor Processor (TPU) released by SOPHGO in 2022, the iterative flagship of the BM1684. Built on a 12nm process, it integrates an 8-core ARM Cortex-A53 and a proprietary NPU (Bernoulli architecture), delivering peak compute of 32 TOPS INT8 / 64 TOPS INT4 / 16 TFLOPS FP16/BF16 / 2 TFLOPS FP32, with strengthened post-processing engines such as NMS/SORT.

The BM1684X targets edge large-model inference and high-density video analytics: it can deploy large models such as Llama3, ChatGLM, and Qwen locally, while supporting 32-channel 1080P video decoding and 16-channel full-pipeline AI analysis. It is often paired with main controllers such as the RK3588 in heterogeneous edge boxes, benchmarked against the NVIDIA Jetson Orin NX.

Specification Correction Note: Some early or unofficial sources list the BM1684X's process as 16nm and its memory as 32GB LPDDR. After checking specification sheets from Sophgo and its partners (Firefly, Tianqi, IOTDT), the actual figures are a 12nm process and on-board memory of up to 16GB LPDDR4X (6/12/16GB options); 32GB appears in multi-node server systems rather than single-chip configurations.

Core Specifications​

ParameterValue
ArchitectureSophgo fourth-generation proprietary TPU (Bernoulli architecture), integrating 8-core ARM Cortex-A53 @ 2.3GHz
Process Node12nm (TSMC; some sources incorrectly list 16nm, corrected)
FP16 / BF16 Compute16 TFLOPS
INT8 Compute32 TOPS (up to 64 TOPS at INT4)
FP32 Compute2 TFLOPS
Memory Capacity6 / 12 / 16GB (max 16GB LPDDR4X)
Memory TypeLPDDR4 / LPDDR4X
Memory BandwidthAbout 68.3 GB/s (16GB configuration, 128-bit @ 4266 Mbps)
TDP≤ 18–20 W (full load, passive fanless cooling)
InterconnectDual Gigabit Ethernet, PCIe 3.0 (16 lanes), multi-chip cascading
InterfaceSoC on-board / PCIe 3.0 (module, computing box, micro-server)
Launch2022
Mass Production/AvailabilityMass production in 2022

Key Features​

  • Edge large-model deployment: supports local private inference of Llama3-8B, ChatGLM2/3-6B, Qwen-7B, Qwen2.5-VL-7B, and more.
  • Mixed precision: full-stack INT8 / FP16-BF16 / INT4 / FP32 precision, with automatic quantization by the TPU-MLIR compiler.
  • High-density video: 32-channel 1080P@25fps decoding, 12-channel encoding, with independent physically isolated VPU and TPU.
  • Full-stack frameworks: PyTorch, TensorFlow, PaddlePaddle, ONNX, Caffe, Darknet, MXNet.
  • Complete toolchain: SophonSDK one-stop compilation/quantization/inference, Docker containerized management.
  • Industrial wide temperature: stable operation from -20℃ ~ +60℃, metal fanless enclosure.

Vendor Information​

ParameterValue
CompanySOPHGO
HeadquartersBeijing, China
Founded2019

Use Cases​

  • ✅ Edge large-model inference, intelligent security, smart cities, intelligent transportation, industrial quality inspection, multi-stream video structuring
  • ❌ Standalone display main controller (no GPU; requires pairing with a main-control SoC)

References​