HomeAI HardwareHuawei Ascend Claim Raises a Narrow Question for AI Chip Buyers

Huawei Ascend Claim Raises a Narrow Question for AI Chip Buyers

A Huawei-linked research team says it completed full-parameter post-training on DeepSeek’s V4-Pro, a 1.6-trillion-parameter model, using a cluster said to include at least 1,000 Ascend 910C chips. That account, attributed to Shenzhen government materials, has not been independently verified.

For AI infrastructure buyers, the important point is not simply the chip count. The practical question is whether domestic Chinese accelerators can handle more of the training pipeline without relying on Nvidia hardware. On that point, the reported result is interesting, but still incomplete.

What the claim appears to show

The team describes the work as full-parameter post-training. In plain terms, that means the reported run updated the model’s weights directly rather than only attaching a smaller adapter layer. Post-training is the stage after pre-training that shapes model behavior through instruction following, alignment work, and task-specific data.

That distinction matters because inference and training stress hardware in different ways. Inference is the serving stage, where a finished model answers prompts. Training and post-training place heavier demands on interconnects, memory movement, software stability, and cluster scheduling.

The reported group included Huawei, the Shenzhen Loop Area Institute, the Shenzhen campus of Harbin Institute of Technology, and the Shenzhen Research Institute of Big Data. That collaboration has also not been independently verified from the source material provided.

What it does not prove

Even if the post-training account is accurate, it does not prove that Ascend clusters can pre-train a frontier model from scratch. Pre-training is the larger and more expensive stage, and it is where hardware, networking, and software bottlenecks become much harder to hide.

The source also says DeepSeek documentation puts the V4-Pro pre-training corpus above 32 trillion tokens, but that figure should be treated as a reported documentation claim rather than an independently verified benchmark.

Question for buyers What the report says What remains unclear
Workload type Full-parameter post-training is claimed No proof of frontier-scale pre-training
Cluster size At least 1,000 Ascend 910C chips are reported Utilization, stability, and scaling efficiency are not disclosed
Performance Earlier DeepSeek testing reportedly put Ascend 910C inference near 60% of Nvidia H100 performance No comparable post-training benchmark is supplied
Software stack Huawei’s CANN stack is positioned as the CUDA alternative Tooling maturity and failure rates are not quantified

The buyer-relevant gap is evidence

The most useful missing details are the ones procurement teams would need before treating Ascend as a training-class alternative: run time, cost, cluster utilization, failure rate, interconnect behavior, and comparison against an equivalent Nvidia setup.

The source says earlier reports described DeepSeek struggling with Ascend-based training for an R2 model, citing unstable performance, slow chip-to-chip interconnects, and gaps in Huawei’s CANN software stack. Those claims are not independently verified here, so they should be read as reported context rather than settled fact.

The same caution applies to the claim that DeepSeek later relied on Nvidia GPUs for training while using Ascend for inference, and to the statement that DeepSeek-V4-Pro was the first DeepSeek model built around Ascend from the outset.

Verdict for infrastructure decisions

This is a notable claim for Huawei’s AI hardware ambitions, but it is not yet enough to support a confident platform decision. A successful post-training run would be meaningful, especially at 1,000-chip scale, but buyers should separate that from proof of reliable, efficient, frontier-scale training.

Until independent benchmarks, run details, and efficiency numbers are available, the safest reading is narrow: the reported Ascend cluster may have handled an important post-training workload, but the evidence does not yet show whether it can replace Nvidia hardware for the hardest parts of large-model training.

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

- Advertisment -

Most Popular

POPULAR TAGS

- Advertisment -