Skip to main content

7 posts tagged with "Cambricon"

Cambricon AI chips and accelerators

View all tags

国产 AI 芯片半年报大检阅:寒武纪净赚 23 亿、壁仞营收暴涨 20 倍、燧原登板,"抢芯大战"白热化

· 6 min read
Industry Research Team

8 月底至 9 月初,国产 AI 芯片厂商半年报密集披露,加上燧原科技 9 月 2 日启动 IPO 申购,被称为"国产四小龙"的摩尔线程、沐曦、壁仞、燧原全部完成上市,加上早已在科创板的寒武纪——国产 AI 芯片的资本市场拼图就此补齐。更重要的是,财报数字第一次集体印证了一件事:国产算力正从"能用"跨向"好用",商业化拐点已现。


1. 半年报成绩单:增长是主旋律,盈利是分水岭​

厂商2026H1 营收同比盈利状态技术路线
寒武纪59.96 亿元+108.1%归母净利 23.11 亿自研 MLU 指令集(DSA)
摩尔线程17.36 亿元+147.4%净亏 1156 万(收窄)全功能 GPU(兼容 CUDA 路线)
壁仞科技12.36 亿元+1997.6%未盈利通用 GPU
沐曦13.24 亿元+44.7%净利 6.12 亿(首次扭亏)通用 GPU
燧原11.20 亿元+279.1%未盈利DSA 专用架构(TopsRider)

几个值得注意的细节:

  • 沐曦率先跨过盈利线:8 月 31 日披露的半年报显示净利润 6.12 亿元、同比扭亏(上年同期亏损 1.86 亿);不过扣非净利润仍为 -4900 万(亏损收窄 75.8%)——含金量仍在爬坡。
  • 壁仞低基数暴增:近 20 倍的同比增速来自上年同期极低的收入基数,但 12.36 亿的绝对体量已与燧原、沐曦同量级,第二梯队座次重新洗牌。
  • 研发强度惊人:沐曦研发费用占营收 39.7%、摩尔线程 44.3%、壁仞高达 65%——高研发投入是全员未完全盈利的根本原因,也是未来竞争力的来源。
  • 摩尔线程毛利率承压:从 78% 降至 57%,成本增速(+245%)远超营收增速(+147%)——大规模量产期的品控与爬坡成本开始显现。

2. 燧原登板:四小龙资本拼图补齐​

燧原科技 9 月 2 日启动公开申购,发行 4303 万股新股(约占上市后总股本 10%),募资目标 60 亿元,采用 DSA 专用架构 + 自研 TopsRider 软件平台(不兼容 CUDA)。至此:

  • 科创板:寒武纪(2020 年)、摩尔线程、沐曦、壁仞(2026 年)
  • 燧原:2026 年 9 月完成上市

国产 AI 芯片第一梯队全部进入公开市场,融资通道打开后,研发投入的"军备竞赛"将进一步升级。

3. 需求端:百万卡缺口,订单排到三年后​

财报爆发的另一面,是需求端的极度饥渴:

  • 产能缺口:行业调研显示,2026 年国产 AI 芯片需求规模约 400 万颗,实际交付约 300 万颗,存在百万级缺口;
  • 订单周期:内蒙古乌兰察布远景星河基地(规划支撑百万卡级并行算力)表示"在手订单已排到三年后";
  • 供给创新:算力紧缺催生"集装箱式算力中心"——20/40 英尺标准集装箱为载体,插电接水即可在 24 小时内完成部署,把传统"一年工期"压缩到"一天上线";
  • 市场空间:IDC 数据显示 2025 年中国 AI 加速卡出货约 400 万片,国产占约 41%(165 万片);CIC 预测中国 AI 加速器市场 2028 年将超万亿元,其中国产方案份额约 90%。

4. 两种路线的两种赌注​

市场研究机构预测 2026 年中国高端 AI 芯片市场中,国产方案份额将接近 90%,其中华为约 62%、寒武纪约 14%,剩余由摩尔线程、沐曦、壁仞等共同占据。而寒武纪与摩尔线程这对"双龙头",正在押注两种截然不同的未来:

寒武纪:用现金买确定性。 半年报资产负债表上,存货 82.48 亿 + 预付款 29.14 亿,两项合计 111.6 亿、占总资产 61%——在实体清单限制下"产能即订单",提前锁定未来 12-18 个月的晶圆产能,代价是经营现金流净额同比下滑 66%,且上半年已计提存货跌价损失 3.97 亿。底气来自订单确定性:57 亿元股权激励计划的考核目标是 2026 年营收不低于 135 亿、2026-2028 年累计不低于 1000 亿。

摩尔线程:用亏损买时间。 全功能 GPU 的"大而全"路线(AI 计算 + 图形渲染 + 物理仿真 + 视频编解码)前期投入巨大,需要在现金流耗尽之前跨过规模化的门槛。上半年亏损已收窄至千万级,IPO 募资到位后,时间窗口正在变宽。

5. 软件生态:被忽视的胜负手​

需求外溢(英伟达 H20 系列在中国市场遇冷)推动 DeepSeek、智谱 GLM、月之暗面 Kimi 等头部模型纷纷适配国产加速器(昇腾、沐曦、摩尔线程等),这给了国产芯片难得的"真实负载打磨机会"。但正如行业共识:芯片可以三年一代,生态需要十年之功。寒武纪自研指令集的封闭高效、摩尔线程 MUSA 对 CUDA 生态的兼容路线、燧原 DSA 的专用极致——哪条路能沉淀出真正的开发者生态,才是决定 2030 年格局的关键变量。


相关链接​

参考资料​


本文基于上市公司半年报、Pandaily 与央视财经等公开报道整理。财务数据为公司披露口径,市场份额为研究机构预测,不构成任何投资建议。

Domestic Big Three 2026 H2: Localization Rate Crosses 40% Toward 60%, Ascend 960 Roadmap, MLU690 and S5000 Ecosystems Ramp Up

· 6 min read
Industry Research Team

In 2026, China's AI chip market landscape has shifted from "NVIDIA unipolar dominance" to "overseas vendors leading, domestic multi-route catch-up." According to industry research, China's overall AI accelerator market was ~4M units in 2025, of which 1.65M were domestic, with share first breaking 40%; as products iterate and fabs follow up, the localization rate is expected to rise to 60%-70% by 2027. This article focuses on the latest H2 2026 progress of Huawei Ascend, Cambricon, and Moore Threads — the domestic "Big Three."


1. Huawei Ascend: 950 Capacity Fully Booked, 960 Roadmap Unveiled​

Ascend's core advantage is "architecture + full-stack ecosystem synergy," with ~800K units shipped in 2025, capturing 50% of the total domestic vendor share. The product iteration cadence is clear:

TimeProductNote
2025 Q1Ascend 910CMain transitional model
2026 Q1Ascend 950PRInference flagship
2026 Q4 (planned)Ascend 950DTTraining flagship, drives domestic HBM iteration
2027-2028Ascend 960 / 970Roadmap products

950 series capacity has entered a "fully booked" state: 950PR entered mass production in April 2026; June monthly capacity jumped to 500K-600K units (nearly 10x MoM), with a full-year target of 1.2M units at 100% certainty; ByteDance locked in 350K units for $5.6B, while Tencent / Alibaba / Baidu combined locked in 400K units.

Ascend 960 roadmap specs (per roadmap disclosure):

MetricAscend 960
ArchitectureAscend 6th gen (Da Vinci v6)
FP8 compute~4 PFLOPS
Memory288GB
Memory bandwidth9.6 TB/s
Super-nodeAtlas 960 SuperPoD, 15,488 cards, Lingqu optical-electrical converged bus
Debut2027 Q4 (roadmap)

The previous-gen Ascend 384 super-node has cumulatively shipped over 750 sets, deployed across 20+ industries including internet, operators, finance, education, and healthcare — Huawei calls it "the only domestic super-node that has trained a SOTA model."


2. Cambricon MLU690: H2 Mass Production, Entering ByteDance Bidding Window​

Cambricon is the core domestic compute leader in the absence of an Ascend IPO, with the technology gap continuously narrowing:

  • Siyuan 590 (7nm): Performance equivalent to 80% of A100, already supports DeepSeek, continuously adapting to mainstream large models like Qwen 3 and GLM
  • Siyuan 690 series: Will enter mass production in H2 2026, expected to achieve order scale-up during ByteDance's H2 bidding window
  • Revenue certainty: Equity incentive targets show >100% revenue growth for the next 3 years: 2026 revenue target 13.5B RMB, 2027 27B RMB, 2028 60B RMB

Cambricon fully benefits from the industry dividend of "domestic CSP capex + full adaptation of domestic large models and domestic chips," making it the most direct elasticity play on rising localization rate.


3. Moore Threads MTT S5000: Full-Function GPU + Ecosystem Breakthrough​

Moore Threads takes a differentiated "full-function GPU" route, with the flagship MTT S5000 based on the 4th-gen "Pinghu" MUSA architecture:

MetricMTT S5000
Dense AI compute1000 TFLOPS
Memory80GB
Memory bandwidth1.6 TB/s
Inter-card interconnect784 GB/s
PrecisionFP8 to FP64 full precision (training + inference)
SecurityFirst batch to pass national "Safe and Reliable Evaluation" (Level I)

Its engineering capability is verified: the Kuae (KUAE) intelligent computing cluster based on S5000 achieves 95% training linear scaling efficiency, with compute efficiency loss within 5% at ten-thousand-card scale; supports checkpoint-resume training with effective training time ratio >90%; and has trained a MoE-236B base model with >25 trillion tokens of corpus from scratch.

The ecosystem is Moore Threads' deepest moat: MUSA has achieved 100% core math library compatibility, 3000+ PyTorch operator compatibility, covers 55 categories of core AI operators, has official vLLM and SGLang support, Day-0 adaptation of mainstream models, and 800K+ developers. Its PD heterogeneous-disaggregation solution achieves equivalent replacement of international high-end GPUs at a 2:1 ratio with S5000, significantly reducing inference cost.

The 5th-gen "Huagang" architecture (released 2025-12) supports FP4 to FP64 full precision, with 50% higher compute density and 10x better energy efficiency than the previous gen, supporting 100K+ card clusters; cumulative R&D investment in the "Huashan" (train-infer integrated) and "Lushan" (graphics rendering) new chips based on this architecture exceeds 900M RMB.


4. Software Ecosystem Decides: Day-0 Adaptation Becomes Routine​

Beyond hardware, software ecosystem realization is the watershed for domestic compute in 2026:

  • Huawei's CANN heterogeneous computing architecture and MindSeries suite are fully open-sourced, with the community incubating 67 projects, 12.44M+ lines of code, and 3,500+ monthly active developers
  • The "release-and-adapt" closed loop between domestic large models and domestic chips has basically formed: Tencent Hunyuan T3 (295B), DeepSeek-V4, and GLM-5.2 all completed Day-0 adaptation
  • 2026 is regarded as the "first year of domestic super-nodes"; Huatai Securities estimates China's super-node architecture market will reach 341.4B RMB by 2028, with a 2026-2028 CAGR of 194%

5. Industry Judgment: From "Can It Be Built" to "Can It Be Used Well"​

The domestic Big Three are converging along three paths:

  1. Huawei: Locks government/enterprise and internet big customers with super-node system-level capability + full-stack software
  2. Cambricon: Impacts the revenue inflection point by narrowing the training-side gap + scaling up via big-customer bidding
  3. Moore Threads: Covers cloud-edge-end full scenarios with full-function GPU generality + mature CUDA-compatible ecosystem

The common shortcoming of all three remains advanced process and HBM supply — precisely the core link of overseas controls. But as domestic HBM iterates and fabs follow up, a realistic path to 60%-70% localization by 2027 exists.

References​


This article is compiled from public industry research, broker views, and corporate announcements as of August 2026. Some shipment and market-share figures are third-party estimates, not officially confirmed data.

Domestic GPU IPO Wave: The "Four Little Dragons" Assemble on Capital Markets, Moore Threads MTT S5000 Benchmarks Against H100

· 5 min read
Industry Research Team

From December 2025 to July 2026 — just half a year — at least 6 AI chip companies have listed or are about to list on capital markets. Together with already-listed Cambricon, Hygon, and Iluvatar, the domestic GPU corps' total market cap is approaching ¥2 trillion. This marks the critical climb from domestic GPUs being "usable" to "useful."

1. The "Four Little Dragons" assemble on capital markets​

CompanyListing statusRaise / issue priceSponsor
Moore ThreadsListed (STAR Market sh688795, 2025-12-05)Issue price ¥114.28, raised ¥8BCITIC Securities
MetaXIPO accepted (2026-06-30)¥3.904B (total investment ¥5B)Huatai United
EnflamePassed review (2026-06-15)¥6B—
BirenHKEX / sprinting——

Already-listed camp: Cambricon (sh688256, STAR Market 2020-07-20), Hygon, Iluvatar (HKEX). Moore Threads turned a book profit of ¥29.35M in Q1; MetaX narrowed losses 57.7% and gave a 2026 breakeven timeline.

2. Moore Threads MTT S5000: benchmarking against H100​

Moore Threads announced its flagship AI train+inference GPU MTT S5000 successfully completed full-pipeline adaptation validation of Zhipu's new-generation large model GLM-5 — measured performance "breaks the domestic compute ceiling":

MetricMTT S5000
Architecture4th-gen "Pinghu" architecture
FP8 compute1 PFLOPS (1,000 TFLOPS)
Memory bandwidth1.6 TB/s
PositioningFull-function train+inference GPU, benchmarks against NVIDIA H100
ProductionMass-produced; clusters online supporting trillion-parameter training

Deployment validation: jointly completed full-pipeline training of embodied-brain model RoboBrain 2.5 with BAAI; partnered with SiliconFlow for high-performance DeepSeek-V3 inference, single-card speed near international top products. IPO funds go to three directions: next-gen AI train+inference chip, next-gen graphics chip, next-gen AI SoC chip.

WAIC 2026 new progress: Moore Threads showcased the MTT C256 SuperNode (first-of-its-kind single-layer Scale-up 256-card full interconnect, sub-microsecond latency) and three AI-factory solutions — "model training factory / token production factory / agent production factory"; the company pre-announced H1 2026 revenue of ¥1.65B-1.75B, up 135%-149% YoY.

3. Cambricon: dual flagships MLU590/690​

ChipProcessComputeMemoryCustomer / status
MLU590 (思元590)7nm ChipletINT8 512 TOPS / FP16 345 TFLOPS96 GB HBM2eByteDance inference mainstay, ~80% of A100 overall, mass shipments early 2026
MLU690 (思元690)5nm-class (SMIC N+2)FP16 700+ TFLOPS / INT8 2800+ TOPS196 GB HBM3 (3.35 TB/s)Dual-die packaging, MLU-Link 890 Gbps; ~70% of H100 (80-90% pure inference); ByteDance largest customer, mass production early 2026

Cambricon is the only domestic AI chip vendor with a "unified edge-cloud architecture" — one MLU instruction set spans 思元 220 (edge) → 370 (border) → 590/690 (cloud), with one NeuWare toolchain across compute tiers.

Capital and performance double explosion: Cambricon's total market cap exceeded ¥1 trillion on June 30, 2026, becoming the STAR Market's first "trillion-yuan stock," up 75%+ YTD. On performance, Q1 2026 revenue ¥2.885B (+160% YoY), deducted net profit ¥934M; full-year 2025 revenue ¥6.497B (+453% YoY), net profit attributable to parent ¥2.059B, ending long-term losses. ByteDance has cumulatively deployed over 100k 思元 590/690, its largest customer.

4. DeepSeek-V4 effect: changing the expectation coordinate system​

On April 24, 2026, DeepSeek released the trillion-parameter flagship DeepSeek-V4. Unlike a year earlier when V3's launch sparked debate over "can domestic chips even run large models," this time multiple domestic chips — Huawei Ascend, Cambricon, Hygon, MetaX, Moore Threads, Kunlun, T-Head, Iluvatar — completed adaptation on launch day.

The evaluation coordinate system is shifting: from "what percentage of NVIDIA's same-generation product performance" to "can it carry the real workloads of top-tier large models."

Industry interpretation​

  1. Capital ammunition in place: dense IPOs provide ample funding for domestic GPU R&D iteration and capacity expansion, moving from "technology breakthrough" to "commercial virtuous cycle."
  2. Train+inference becomes the mainstream route: Moore Threads takes the full-function GPU route (graphics+AI+general compute), differentiating from Huawei Ascend's "AI-focused."
  3. Software ecosystem is the decider: Day-0 adaptation and the maturity of unified software stacks (MUSA / NeuWare / MXMACA) are replacing raw peak compute as the core yardstick of domestic GPU "usability."

References​


This article continuously tracks the domestic GPU listing process and product iteration.

2026 H1 AI Chip Industry Review: Blackwell Ultra, the Domestic Big Three, and the Inference Era

· 11 min read
Industry Research Team

In the first half of 2026, the AI chip industry underwent a historic turning point — the center of gravity shifted from the "training race" to "inference efficiency," domestic chip market share broke 40% for the first time, NVIDIA built higher barriers with Blackwell Ultra, and the inference-specific chip track bloomed in diversity.


I. Compute Doubles Again: NVIDIA Blackwell Ultra Launch (June 1)​

On June 1, 2026, NVIDIA CEO Jensen Huang unveiled the new-generation AI chip Blackwell Ultra at Computex 2026 (Taipei), setting a new starting line for the AI infrastructure race over the next two years.

Key Specs​

MetricBlackwell UltraB200Improvement
FP8 compute20 petaFLOPS~10 petaFLOPS100%
ArchitectureBlackwell UltraBlackwellUpgrade
Expected delivery2027 Q12026 Q1—
PositioningHyperscale training + inferenceTraining + inferenceFlagship

Industry Significance​

  1. Direct impact of doubled compute: 20 petaFLOPS FP8 means training time for hundred-billion-parameter models drops sharply; trillion-parameter model training moves from "scientific experiment" to "engineering routine"
  2. System-level balance: Blackwell Ultra is not just a chip but a system-level engineering breakthrough across NVLink, HBM, cooling, and power delivery
  3. Roadmap certainty: The Q1 2027 delivery timeline lets cloud vendors and AI labs plan infrastructure budgets 18 months ahead

Challenges​

  • Energy crisis: Doubled performance comes with sharply higher power; datacenter power and cooling design face extreme challenges
  • Accessibility: Top-tier compute goes first to top cloud vendors; how smaller developers and research institutes reach compute at reasonable cost via cloud services
  • Software stack adaptation: New hardware needs matching CUDA versions and framework support; software ecosystem maturity becomes the key bottleneck for compute conversion

II. Domestic AI Chips: The Tipping Point from "Usable" to "Good"​

On June 16, 2026, Xinchuang World published "2026 China Domestic AI Chip Vendor Capability Quadrant", clearly outlining the overall domestic landscape.

2.1 Capability Quadrant Ranking​

QuadrantRepresentative Vendors
Leader quadrantHuawei Ascend, Hygon, Cambricon, Alibaba T-Head, Moore Threads
Visionary quadrantBaidu Kunlunxin, Biren, Enflame, Iluvatar, HardyVision
Contender quadrantTSINGMICRO, Black Sesame, SemiDrive, Lisuan, Houmo
Challenger quadrantDenglin, Zhicun, VeriSilicon, Rockchip, Intellifusion

2.2 Huawei Ascend: The Anchor of Domestic Compute​

Market Position​

  • In 2025, Ascend series shipped 812,000 units, capturing 49% of the domestic AI accelerator card share, firmly No.1 domestically
  • Ascend 950PR single-card FP8 compute reaches 1P (PetaFLOPS), FP4 compute reaches 2P
  • Inference performance is about 2.87x that of NVIDIA H20, priced at only 72,000-75,000 RMB, a significant price/performance advantage

Full-Stack Advantage​

Huawei's "device-network-cloud-chip" integrated strategy is Ascend's core moat:

  • Chip design: Da Vinci 3.0 architecture iterating continuously
  • OS: HarmonyOS/Euler OS deeply optimized
  • Networking: Euler network protocol stack
  • Cloud: Huawei Cloud ModelArts platform seamlessly integrated

Latest Progress​

  • On June 5, 2026, Shenzhen Hetao College, together with HIT (Shenzhen) and Huawei, completed full-parameter post-training of a 1.6-trillion-parameter DeepSeek V4 Pro model on an Ascend 910C cluster
  • This is the first time domestic AI chips completed trillion-parameter-level model training, marking "domestic substitution" moving from inference to training

2.3 Cambricon: The First Profitable Domestic AI Chip Benchmark​

Performance Explosion​

MetricFull-year 20252026 Q1YoY Growth
Revenue6.497B RMB2.885B RMB+453% / +160%
Net profit2.059B RMB (first annual profit)1.013B RMB— / +185%

Core Product: Siyuan 590​

  • In DeepSeek R1 inference scenarios, TPS reaches 942, about 50% higher than H20
  • Years of joint optimization with ByteDance; strongest short-term cloud inference deployment capability
  • Of 2.885B RMB Q1 2026 revenue, Siyuan 590 contributed over 70%

Potential Risks​

Absent from the 2nd 2026 "Safe and Reliable Evaluation Results Announcement"; the reason is unclear and will affect its domestic government/enterprise market performance.

2.4 TSINGMICRO: The "Third Route" of Reconfigurable Chips​

Technical Route​

TSINGMICRO adopts a reconfigurable dataflow architecture同源 with Groq LPU, finding a balance between GPU generality and ASIC extreme efficiency.

MetricTSINGMICRO TX81Traditional GPUAdvantage
Inference costBaseline+100%Reduced 50%
Energy efficiencyBaselineBaseline3x improvement
ArchitectureReconfigurable dataflowSIMT/SIMDBetter for inference

Deployment Progress​

  • Cumulative shipments of reconfigurable chips exceed 30 million units
  • Scaled deployment in a dozen-plus thousand-card-scale intelligent computing centers nationwide
  • Has begun A-share IPO tutoring; likely to become the "first reconfigurable chip stock"

III. The Inference Chip Track: Core Signal of the Industry Shift​

On June 4, 2026, TrendForce published a deep report "The Era of Inference Economy: The Rules of AI Chips Are Being Rewritten," pointing out that the compute competition center of gravity is shifting from training to inference.

3.1 Why Now?​

Cost Structure Changed​

  • Training is a one-time cost: Once a model is trained, marginal cost approaches zero
  • Inference is a recurring cost: Every API call, every generated token represents compute consumption and gross-margin pressure
  • Per-unit inference cost and energy efficiency directly affect gross margin and scale-expansion capability

Model Compression Tech Matured​

  • 1.58-bit quantization and weight pruning let models maintain inference accuracy at extremely low memory footprint
  • MoE (Mixture of Experts) architecture activates only a few expert sub-networks per inference via "partial wake-up," greatly reducing actual computation
  • The rise of slimmed models provides commercial viability for hard-wired inference chips

3.2 NVIDIA's $20B Bet: Acquiring Groq (December 2025)​

On December 24, 2025, NVIDIA acquired Groq's inference technology license and core team for $20 billion, one of NVIDIA's largest M&A/tech acquisitions ever.

Strategic intent:

  1. Fill the inference gap: NVIDIA GPU is unshakable in training, but inference efficiency was never its strongest suit
  2. Counter specialized inference chips: Cerebras, Taalas, SambaNova and other startups are eroding the inference market
  3. Position for Agentic AI: Agentic AI needs extremely low-latency, high-throughput inference

3.3 Taalas HC1: Proof of Concept for Hard-Wired Inference​

On February 20, 2026, Canadian AI chip startup Taalas launched inference chip Taalas HC1, directly etching Meta's open-source AI model Llama 3.1 8B into the chip.

Key Metrics​

MetricTaalas HC1NVIDIA B200 (throughput optimized)Advantage
Inference rate16,960 tokens/s/userBaseline~4-5x
Cost per million tokens0.75 cents3.79 centsReduced 80%
Power~250W~700WReduced 64%
ProcessTSMC N6TSMC 4nmMore mature
HBM❌ Not used✅ HBM3eLower cost

Technical Principle​

Taalas HC1 uses an aggressive Computing-in-Memory (CIM) implementation:

  • Model weights directly固化 in Mask ROM (fully hardware-defined)
  • On-chip SRAM handles dynamic data (KV cache and LoRA fine-tuning weights)
  • Only 2 mask layers need modification to produce a dedicated chip for another AI model; turning an AI model into a physical chip takes only 2 months

Limitations​

  • Lack of flexibility: Hard-wiring cannot cope with rapidly iterating model updates
  • Ecosystem barrier: The current cloud market still relies on general-purpose platforms; customers may prefer flexible solutions that upgrade with models
  • NRE cost: High one-time engineering cost, requiring sufficient deployment scale to amortize

3.4 Cerebras: The IPO Path of Wafer-Scale Integration​

On May 14, 2026, Cerebras Systems officially listed on NASDAQ, becoming the first wafer-scale AI chip company to go public.

Core Technology: Wafer-Scale Integration (WSI)​

  • WSE-3 (third-gen wafer-scale engine): An entire 12-inch wafer as a single chip
  • 44GB on-chip SRAM: No external HBM, eliminating the memory bandwidth bottleneck
  • 21 PB/s bandwidth: On-chip communication bandwidth, thousands of times that of GPUs
  • Partnership with OpenAI: Signed a 3-year, 750MW, $20B+ compute cooperation agreement

IPO Significance​

Cerebras's listing marks the maturation of the inference-specific chip track:

  1. Capital markets begin pricing such companies
  2. Proves "non-GPU" technical routes have commercial viability
  3. Provides valuation references for other inference chip startups (Groq, SambaNova, Taalas, etc.)

3.5 Inference Chip Landscape: Multiple Technical Routes Coexist​

CompanyTechnical RouteCore AdvantageRepresentative Product
TaalasHard-wired (Mask ROM)Extreme inference efficiency, low costHC1
CerebrasWafer-scale integration (WSI)Ultra-high bandwidth, large-model inferenceWSE-3
GroqSRAM-first architectureDeterministic latency, high throughputLPU (acquired by NVIDIA)
d-MatrixDigital in-memory compute (DIMC)More flexible than hard-wiringCorsair
EtchedHard-wired TransformerWorks for all Transformer modelsSohu
Axelera AIDigital in-memory compute (D-IMC) + RISC-VHigh energy efficiencyMetis AIPU

TrendForce predicts:

  • General-purpose GPUs still dominate training and multi-model environments
  • But in mature, predictable scenarios, general-purpose GPU profit margins will be compressed
  • The industry shifts from general compute monopoly to a dual-track structure of general + specialized coexistence

IV. Overall Domestic AI Chip Landscape in H1 2026​

4.1 Industry Enters Scale-Up Phase​

Metric20252026 Q1Trend
Domestic AI accelerator shipments1.65M units (41% share)—Rising
Total China AI accelerator shipments~4M units——
Hygon revenue growth—Doubled↑
Cambricon revenue growth—+160%↑
Moore Threads revenue growth—Doubled↑

Leading vendors collectively entered the revenue realization channel, moving from "technical validation" to "scale commercialization."

Trend 1: Capitalization Wave Reshapes the Landscape​

  • Late 2025 to early 2026: Moore Threads, Iluvatar listed on the STAR Market
  • Biren listed on the Hong Kong stock exchange
  • Enflame STAR Market IPO accepted
  • Kunlunxin, T-Head initiated listing processes
  • TSINGMICRO, HardyVision and others advancing IPOs

Capitalization brings dual effects:

  • ✅ Positive: Supports R&D and ecosystem building
  • ⚠️ Negative: Valuation bubbles and revenue realization pressure

Trend 2: Capacity Becomes the Biggest Constraint Variable​

The contradiction between explosive domestic AI chip demand and limited advanced-process capacity is sharpening:

VendorAdvanced-process capacity needActually obtained
Huawei Ascend15K wafers/month (7nm-class)Priority guaranteed
SMIC total capacity~20K wafers/month (7nm-class)—
Other vendors~5K wafers/month combinedExtremely tight

Whether stable wafer capacity can be secured directly determines vendor survival. Cambricon's 75.4% inventory-to-revenue ratio is essentially a lock on capacity.

Trend 3: Competition Shifts from "Usable" to "Good"​

Early competition focused on "can it run the model"; now it's about "runtime efficiency, deployment cost":

Dimension"Usable" era"Good" era
Hardware performanceCan it run the modelRuntime efficiency, energy efficiency
Software stackBasic adaptationMaturity, framework breadth
EcosystemExistenceDeveloper community activity
Deployment costInsensitiveCore competitive factor

V. H2 2026 Outlook​

5.1 Upcoming Key Events​

TimeEventImpact
2026 Q3NVIDIA Rubin architecture details revealedNext-gen flagship specs unveiled
2026 Q3Huawei Ascend 950PR/950DT formally launchedNew benchmark for domestic inference chips
2026 Q4AMD MI350X scaled deliveryNVIDIA Blackwell competitor
2026 Q4Cambricon Siyuan 690 launch (est.)New-gen training chip
2027 Q1NVIDIA Blackwell Ultra deliveryNew compute benchmark lands

5.2 Key Competitive Factors Over the Next Three Years​

  1. Wafer capacity access: Advanced-process capacity is a scarce resource; vendors tied to SMIC and TSMC have inherent advantages
  2. Capital operation efficiency: The IPO window is limited; raising enough capital on the market determines R&D sustainability
  3. Software ecosystem depth: Hardware performance is only the entry ticket; software stack maturity, framework adaptation breadth, and developer community activity are the core moat

VI. Conclusion: A Diverse Ecosystem Will Eventually Form​

In H1 2026, the AI chip industry is undergoing a historic transition from "one dominant player" to "pluralistic coexistence."

  • NVIDIA builds higher training barriers with Blackwell Ultra while laying out inference efficiency via the Groq acquisition
  • Huawei Ascend holds the domestic compute baseline with full-stack capability; 950PR begins to surpass H20 in inference
  • Cambricon proves the commercial viability of domestic AI chips by turning profitable first; Siyuan 590 surpasses international rivals in specific scenarios
  • Cerebras, Taalas and other inference-specific chip companies opened a "non-GPU" third route
  • TSINGMICRO's reconfigurable architecture provides a diversified technical route choice for China's AI chips

Over the next three years, the domestic AI chip endgame will form a pluralistic ecosystem where GPU, ASIC, and reconfigurable computing three technical routes coexist, with cloud and edge developing in coordination. "Domestic substitution" is no longer a slogan, but an industrial reality happening now.


Data sources:

  • Xinchuang World "2026 China Domestic AI Chip Vendor Capability Quadrant" (2026-06-16)
  • TrendForce "The Era of Inference Economy: The Rules of AI Chips Are Being Rewritten" (2026-06-04)
  • RayByte "Compute Doubles! NVIDIA Blackwell Ultra Chip Launched" (2026-06-02)
  • Official financial reports and announcements of each company

Related reading:


China's Domestic AI Chip Triopoly (2026): Ascend, Cambricon, Moore Threads — Who Is the "China H100"?

· 7 min read
AI Hardware Analyst

Against the backdrop of U.S. export controls, China's AI chip market is forming a "three-way standoff." This article compares the technical routes, product specs, software ecosystems, and commercial progress of the three major domestic AI chip vendors: Huawei Ascend, Cambricon MLU, and Moore Threads MTT.


Key Points​

  • Huawei Ascend: leader in domestic AI training chips; Ascend 950 in mass production; most mature software ecosystem
  • Cambricon MLU690: the "China H100," compute close to H200, clear efficiency advantage
  • Moore Threads MTT S5000: full-function GPU route; achieved Day-0 support for Qwen3.5 and GLM-5.2 in June 2026
  • Shared challenge: affected by U.S. export controls, primarily aimed at the Chinese market, limited internationally

I. Vendor Overview​

VendorFoundedFounderListed2025 RevenueMain Customers
Huawei Ascend2018 (division)Ren Zhengfeiprivate (wholly owned by Huawei)~¥20B (est.)Chinese gov, SOEs, military
Cambricon2016Chen Tianshi (CAS)2020-07 (STAR Market 688256)~¥5.2BByteDance, Alibaba, Baidu
Moore Threads2020Zhang Jianzhong (ex-NVIDIA China)2023-12 (STAR Market 688495)~¥1.5B (est.)gov, SOEs, gaming cos.

Strategic Positioning​

VendorTech routeCore strengthMain challenge
Huawei AscendAI-training-specific (Da Vinci)co-optimized HW/SW, carrier channelssanctions, process limits
CambriconAI-training-specific (MLUarch)high efficiency, competitive priceimmature ecosystem
Moore ThreadsFull-function GPU (MUSA)graphics + AI + general compute, Day-0 supportcompute below dedicated AI chips

II. Flagship Product Comparison​

1. Huawei Ascend 950DT (2026 flagship)​

ItemSpec
BF16 compute1,000 TFLOPS
Memory144GB HiZQ 2.0 (in-house HBM)
Memory bandwidth4 TB/s
TDP400W
ProcessN+2 (improved 7nm)
Released2026-04
Mass production2026-Q2
Unit price~¥80,000 (est.)

Strengths:

  • ✅ High large-model inference throughput: 144GB memory friendly to DeepSeek R1 (671B MoE)
  • ✅ Most mature ecosystem: CANN ~85% operator coverage, supports PyTorch, TensorFlow
  • ✅ Strong carrier channel: China Mobile, China Telecom large purchases

Weaknesses:

  • ❌ Process limited: N+2 below TSMC 4nm
  • ❌ Mediocre efficiency: 400W TDP, 2.5 TFLOPS/W

2. Cambricon MLU690 (2026 flagship)​

ItemSpec
BF16 compute600 TFLOPS
Memory64GB HBM3
Memory bandwidth2 TB/s
TDP280W
ProcessTSMC 7nm
Released2025-Q4
Mass production2026-Q1
Unit price~¥140,000 (est.)

Strengths:

  • ✅ Best efficiency: 280W TDP, 2.14 TFLOPS/W (1.5x H100)
  • ✅ Competitive price: ~$20,000, 33% cheaper than H100
  • ✅ Top-tier customer orders: ByteDance, Alibaba, Baidu

Weaknesses:

  • ❌ Small memory: 64GB limits large-model training scale
  • ❌ Immature ecosystem: NeuWare ~75–85% coverage; complex LLMs need manual tuning

3. Moore Threads MTT S5000 (2025 flagship)​

ItemSpec
FP16 compute~1,000 TFLOPS (est.)
Memory80GB GDDR6X
Memory bandwidth1.6 TB/s
TDP~350W
ProcessTSMC 4nm (est.)
Released2025-02
Mass production2025-Q2
Unit price~¥50,000 (est.)

Strengths:

  • ✅ Full-function GPU: graphics + AI + general compute, broader scenarios
  • ✅ Strong Day-0 support: June 2026 Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • ✅ Lowest price: ~¥50,000, high cost-performance

Weaknesses:

  • ❌ Compute below dedicated AI chips: FP16 ~50% of H100
  • ❌ Low memory bandwidth: 1.6 TB/s (48% of H100), limits large-model training

III. Compute Comparison (BF16/FP16)​

ChipBF16 computeMemoryBandwidthTDPEfficiency
Huawei Ascend 950DT1,000 TFLOPS144GB4 TB/s400W2.5 TFLOPS/W
Cambricon MLU690600 TFLOPS64GB2 TB/s280W2.14 TFLOPS/W
Moore Threads MTT S5000~1,000 TFLOPS80GB1.6 TB/s~350W~2.86 TFLOPS/W
NVIDIA H100989 TFLOPS80GB3.35 TB/s700W1.41 TFLOPS/W
NVIDIA H200989 TFLOPS141GB4.8 TB/s700W1.41 TFLOPS/W

Key insights:

  1. Ascend 950DT has the highest compute (1,000 TFLOPS) but mediocre efficiency
  2. Cambricon MLU690 has the best efficiency (2.14 TFLOPS/W), TDP only 280W
  3. Moore Threads MTT S5000 wins on full-function versatility but low bandwidth

IV. Software Ecosystem​

VendorStackFramework supportCoverageMaturity
Huawei AscendCANNPyTorch, TensorFlow, MindSpore~85%⭐⭐⭐⭐ (4/5)
CambriconNeuWarePyTorch-Cambricon, TensorFlow-Cambricon~75–85%⭐⭐⭐ (3/5)
Moore ThreadsMUSIFYPyTorch, TensorFlow, ONNX~70%⭐⭐⭐ (3/5)
NVIDIACUDAall~99%⭐⭐⭐⭐⭐ (5/5)

Ecosystem Maturity Assessment​

Huawei Ascend CANN:

  • ✅ Strength: highest operator coverage, supports MindSpore (in-house framework)
  • ❌ Weakness: steep learning curve, incomplete docs

Cambricon NeuWare:

  • ✅ Strength: PyTorch/TensorFlow compatible, low migration cost
  • ❌ Weakness: complex LLMs need manual tuning

Moore Threads MUSIFY:

  • ✅ Strength: strong Day-0 support, ONNX support
  • ❌ Weakness: lowest operator coverage, dual graphics+AI engine complexity

V. Commercial Progress​

Vendor2026 commercial progressMain customersShipments
Huawei AscendAscend 950 mass production; China Mobile large purchaseChina Mobile, China Telecom, gov~100K/yr (est.)
CambriconMLU690 mass production; ByteDance, Alibaba ordersByteDance, Alibaba, Baidu~50K/yr (est.)
Moore ThreadsMTT S5000 mass production; Day-0 Qwen3.5gov, SOEs, gaming cos.~30K/yr (est.)

Latest as of June 2026​

Huawei Ascend:

  • ✅ Ascend 950DT fully ramping
  • ✅ ¥1B procurement agreement with China Mobile

Cambricon:

  • ✅ MLU690 in volume shipment
  • ✅ ByteDance order ~20K units

Moore Threads:

  • ✅ Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • ✅ MTT S5000 2nd-gen released

VI. Selection Advice​

Scenario 1: Trillion-parameter training (GPT-4 class)​

Recommended: Huawei Ascend 950DT

  • ✅ 144GB large memory supports super-large models
  • ✅ Most mature ecosystem (~85% coverage)
  • ✅ Strong carrier channel, Chinese government backing

Alternative: Cambricon MLU690 (high efficiency, but small memory)

Scenario 2: Tens-to-hundreds-of-billions parameter training​

Recommended: Cambricon MLU690

  • ✅ Best efficiency (2.14 TFLOPS/W), low TCO
  • ✅ Competitive price (~$20,000)
  • ✅ Validated by top customers (ByteDance, Alibaba)

Alternative: Huawei Ascend 920 (more compute, mediocre efficiency)

Scenario 3: Cloud AI inference​

Recommended: Huawei Ascend 950PR (inference-specific)

  • ✅ Well-optimized inference throughput
  • ✅ 128GB memory friendly to MoE models
  • ✅ Mature stack, low deployment cost

Alternative: Moore Threads MTT S5000 (full-function GPU, inference + graphics)

Scenario 4: Edge AI / on-device inference​

Recommended: Moore Threads MTT S5000

  • ✅ Full-function GPU, graphics + AI
  • ✅ Lowest price (~¥50,000)
  • ✅ Strong Day-0 support

Alternative: Huawei Ascend 310 (low power, 8W TDP)

Scenario 5: Domestic substitution (gov, SOEs)​

Recommended: Huawei Ascend 950DT

  • ✅ Chinese government first choice, carrier bulk buys
  • ✅ Co-optimized HW/SW, stable performance
  • ✅ Supported by national semiconductor fund

Alternative: Cambricon MLU690 (high efficiency, competitive price)


VII. Future Roadmap​

Vendor2026 H220272028
Huawei Ascend950DT ramp960 (FP8 ~2 PFLOPS)970 (N+3 process)
CambriconMLU690 rampMLU790 (5nm, BF16 ~1,000 TFLOPS)MLU890 (3nm)
Moore ThreadsMTT S5000 2nd-genMTT S6000 (HBM3, FP16 ~1,500 TFLOPS)MTT S7000

VIII. Summary: Who Is the "China H100"?​

DimensionAscend 950DTMLU690MTT S5000
Compute⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Memory⭐⭐⭐⭐⭐ (5/5)⭐⭐ (2/5)⭐⭐⭐ (3/5)
Efficiency⭐⭐⭐ (3/5)⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐⭐ (4/5)
Ecosystem⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Price⭐⭐⭐ (3/5)⭐⭐⭐⭐ (4/5)⭐⭐⭐⭐⭐ (5/5)
Overall⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)

Final conclusion:

  • Huawei Ascend 950DT is the domestic AI training chip closest to H100, strongest overall
  • Cambricon MLU690 is the most efficient domestic AI chip, lowest TCO
  • Moore Threads MTT S5000 is the cheapest full-function GPU, suited to edge AI and graphics+AI

References​


Disclaimer: Data based on public sources; actual specs per vendor official. MirrorFrog continuously updates domestic AI chip data — corrections welcome.

Changelog: 2026-06-23 initial release

Cambricon MLU690 vs NVIDIA H100: In-Depth Comparison — Can a Domestic AI Chip Replace the H100?

· 6 min read
AI Hardware Analyst

In 2026, against the backdrop of U.S. export controls on AI chips to China, Cambricon's MLU690 has drawn intense attention as a "China-made H100." This article compares the two in depth across compute, memory, power, software ecosystem, measured performance, and price to help you make a selection decision.

Core Verdict (Read This First)​

DimensionMLU690H100WinnerGap
BF16 compute600 TFLOPS989 TFLOPSH100+65%
Memory capacity64GB HBM380GB HBM3H100+25%
Memory bandwidth2 TB/s3.35 TB/sH100+68%
TDP280W700WMLU690-60%
Energy efficiency2.14 TFLOPS/W1.41 TFLOPS/WMLU690+52%
Software ecosystemNeuWare (~75% coverage)CUDA (100% coverage)H100large gap
Price~¥140,000~¥200,000MLU690-30%
Availabilitydomestic spot stockexport-controlledMLU690✅

One-line summary: MLU690 delivers roughly 60% of H100's compute, but at only 40% of the power and 70% of the price — a strong fit for AI training and inference in the Chinese market.


1. Detailed Spec Comparison​

1.1 Compute​

PrecisionMLU690H100 SXM5H200 SXM5Note
FP8~300 TFLOPS (est.)3,958 TFLOPS3,958 TFLOPSH100 supports FP8; MLU690 likely does not
BF16/FP16600 TFLOPS989 TFLOPS989 TFLOPSH100 leads by 65%
FP32~150 TFLOPS (est.)60 TFLOPS60 TFLOPSMLU690 estimate; H100 actually higher
INT81,200 TOPS1,979 TOPS1,979 TOPSH100 leads by 65%

Key findings:

  • ✅ MLU690 reaches 60% of H100's BF16 compute
  • ⚠️ H100 supports FP8 (4-bit); MLU690 likely does not (needs confirmation)
  • ⚠️ H100's higher INT8 compute favors inference scenarios

1.2 Memory​

ItemMLU690H100H200Note
Capacity64GB HBM380GB HBM3141GB HBM3eH200 largest
Bandwidth2 TB/s3.35 TB/s4.8 TB/sH200 highest
TypeHBM3HBM3HBM3eH200 uses latest HBM3e

Key findings:

  • ⚠️ MLU690 has 20% less memory than H100 (64GB vs 80GB)
  • ⚠️ MLU690 bandwidth is 40% lower than H100 (2 TB/s vs 3.35 TB/s)
  • ❌ When running 70B+ parameter models, MLU690 may run out of memory (model parallelism required)

1.3 Power​

ItemMLU690H100H200
TDP280W700W700W
Efficiency (FP16/W)2.14 TFLOPS/W1.41 TFLOPS/W1.41 TFLOPS/W
8-card server power~3.5kW~6kW~6kW
Annual electricity (¥0.6/kWh)~¥18,400~¥36,800~¥36,800

Key findings:

  • ✅ MLU690 draws only 40% of H100's power, sharply cutting data-center electricity cost
  • ✅ MLU690 leads efficiency by 52%, better suited to large-scale deployment
  • ✅ For power-sensitive inference, MLU690 has a clear edge

2. Software Ecosystem​

2.1 Framework Support​

FrameworkMLU690 (NeuWare)H100 (CUDA)Note
PyTorch✅ (PyTorch-Cambricon)✅ nativeMLU690 needs an extra plugin
TensorFlow✅ (TensorFlow-Cambricon)✅ nativesame
JAX⚠️ partial✅ nativeMLU690 limited
ONNX⚠️ partial✅ nativesame
vLLM⚠️ in progress✅ nativeMLU690 awaits community port

2.2 Operator Coverage​

CategoryMLU690H100Note
Basic operators✅ 95%✅ 100%conv, matmul, etc.
Transformer operators✅ 85%✅ 100%Attention, LayerNorm, etc.
Custom operators⚠️ hand-written✅ CUDA C++MLU690 harder to develop
LLM inference opt.⚠️ basic✅ mature (FlashAttention, PagedAttention)H100 leads

Key findings:

  • ⚠️ NeuWare is only 5–6 years old, with ~75–85% operator coverage
  • ❌ Complex LLMs (e.g., GPT-4, Claude) may need manual optimization
  • ✅ Common models (Llama, Qwen, GLM) are essentially already supported

3. Measured Performance​

3.1 Training​

ModelMLU690 (time)H100 (time)Speedup
Llama 7B~48 h (est.)~30 h1.6x
Llama 70B~7 days (est.)~4.5 days1.6x
Qwen 72B~8 days (est.)~5 days1.6x

Note: above figures are estimates; real performance depends on software optimization.

3.2 Inference​

ModelMLU690 (tok/s)H100 (tok/s)Note
Llama 7B~80 tok/s (est.)~120 tok/sH100 +50%
Llama 70B~20 tok/s (est.)~35 tok/sH100 +75%
Qwen 72B~18 tok/s (est.)~30 tok/sH100 +67%

Key findings:

  • ⚠️ H100 leads inference by 50–75%
  • ✅ But MLU690 draws only 40% the power, with better efficiency
  • ✅ For cost-sensitive inference, MLU690 is more economical

4. Price​

4.1 Hardware Procurement​

ItemMLU690H100H200
Per-card (domestic)~¥140,000~¥200,000~¥300,000
8-card server (turnkey)~¥1,200,000~¥1,800,000~¥2,600,000
Cost gap-+50%+117%

4.2 TCO (3 years)​

ItemMLU690H100Note
Hardware¥1,200,000¥1,800,000MLU690 33% cheaper
Electricity (3y)¥55,200¥110,400MLU690 50% cheaper
Facility¥150,000¥250,000MLU690 40% cheaper
TCO (3y)¥1,405,200¥2,160,400MLU690 35% cheaper

Key findings:

  • ✅ MLU690's TCO is 35% lower than H100's
  • ✅ For large-scale deployment (100+ cards), the cost advantage is pronounced

5. Selection Advice​

5.1 Choose MLU690 if...​

  • ✅ Your business is primarily in the Chinese market
  • ✅ You are affected by U.S. export controls and cannot buy H100/H200
  • ✅ You are power-sensitive (edge data centers, high electricity-cost regions)
  • ✅ Your models use common architectures (Llama, Qwen, GLM)
  • ✅ You have domestic-substitution requirements (government, SOEs, military)

5.2 Choose H100/H200 if...​

  • ✅ Your business is global
  • ✅ You need to train frontier models (GPT-4 class)
  • ✅ Your models use complex operators (need the CUDA ecosystem)
  • ✅ You demand extreme performance (low-latency inference)
  • ✅ You can legally procure H100/H200
ScenarioRecommended
TrainingH100 (high perf) + MLU690 (low-cost scale-out)
InferenceMLU690 (cost-sensitive) + H100 (low-latency)
Domestic projectall MLU690
International marketall H100/H200

6. Outlook​

6.1 MLU690's weaknesses​

  • ⚠️ Immature software ecosystem: 75–85% operator coverage; complex models need manual tuning
  • ⚠️ Small memory: 64GB limits support for 70B+ parameter models
  • ⚠️ Weak interconnect: Cambricon Link bandwidth below NVLink
  • ⚠️ Limited international market: affected by U.S. export controls

6.2 MLU690's improvement path​

  • 📅 MLU790 (2027): expected 5nm process, ~2x compute
  • 📅 Memory upgrade: next gen may adopt HBM3e, capacity up to 128GB
  • 📅 Software: NeuWare ecosystem improving, operator coverage target 95%

7. Summary​

DimensionMLU690H100Recommended scenario
Compute⭐⭐⭐⭐⭐⭐⭐⭐⭐H100 for top-tier training
Memory⭐⭐⭐⭐⭐⭐⭐H100 for large models
Power⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for inference
Ecosystem⭐⭐⭐⭐⭐⭐⭐⭐H100 for complex models
Price⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for large-scale deployment
Domestic⭐⭐⭐⭐⭐❌MLU690 for Chinese market

Final recommendation:

  • 🇨🇳 Chinese market: prefer MLU690 (domestic + low cost)
  • 🌍 International market: prefer H100/H200 (performance + ecosystem)
  • 💡 Hybrid: train on H100, infer on MLU690

References​


Disclaimer: Data in this article is based on public sources and reasonable estimates; actual performance is subject to vendor official testing. MLU690's software ecosystem is evolving rapidly — watch NeuWare updates.

Last updated: 2026-06-23

China AI Chip Landscape 2025: Ascend, Cambricon, Hygon — Who Will Dominate?

· 5 min read
Industry Research Team

Escalating U.S. export controls are forcing China's AI chip industry to accelerate self-reliance. By 2025, the discussion around domestic Chinese AI chips has shifted from "are they usable?" to "which one should I choose?"

This article systematically reviews the major players, core products, and actual deployment status of domestic AI chips, helping developers and procurement decision-makers understand the competitive landscape.