AI News

DeepSeek plans at least 160,000 Huawei Ascend chips in Inner Mongolia

Bloomberg says DeepSeek wants at least 160,000 of Huawei's Ascend 950DT accelerators for a data centre at Ulanqab in Inner Mongolia. Huawei sells that chip for decode and training. DeepSeek's plan uses it for decode, and keeps Nvidia for the training runs.

Listen to this articleListen

Editorial collage of a Huawei Ascend accelerator card on newsprint, with the DeepSeek whale mark and a torn map of Inner Mongolia, under the headline DEEPSEEK and the line 160,000 ASCEND 950DT · ULANQAB.

Bloomberg reported on 4 September that DeepSeek intends to install at least 160,000 of Huawei’s Ascend 950DT accelerators at a data centre it is building at Ulanqab, in Inner Mongolia. The report attributes the plan to unnamed people, puts delivery at more than a year, says shortages of high-bandwidth memory will hold 950DT output to the low hundreds of thousands across this year, and says DeepSeek has asked Beijing to help persuade Huawei to allocate more chips to it, and sooner. DeepSeek intends to run its models on those chips, and to keep training them on Nvidia.

As of 5 September that report is the only account of the order, and both companies have left it to stand on its own. The chip, the site and the rule that governs the alternative are all documented, so those are the parts a reader can check.

What Huawei sells the 950DT for

Huawei’s rotating chairman Eric Xu set out the Ascend roadmap in his keynote at Huawei Connect 2025 in Shanghai on 18 September 2025. He split the 950 generation across two parts on a shared die: “So we’ll have the Ascend 950PR chip for prefill and recommendation, and the Ascend 950DT chip for decode and training.”

The DT part is the one DeepSeek is reported to want. In Xu’s words it “is optimized for both the decode stage of inference and for model training”, and its specification is built around the bandwidth those two jobs need: Huawei’s own HiZQ 2.0 memory, 144GB of it, 4TB/s of memory access bandwidth and 2TB/s of total interconnect. He dated it to the fourth quarter of 2026.

Ascend part Job Huawei names for it Memory Availability
Ascend 950PR Prefill and recommendation HiBL 1.0 Q1 2026
Ascend 950DT Decode and model training HiZQ 2.0, 144GB at 4TB/s Q4 2026
Ascend 960 Training and inference Twice the 950’s capacity and bandwidth Q4 2027
Ascend 970 Training and inference Specs still being set, bandwidth at least 1.5x the 960 Q4 2028

Both 950 parts carry 2TB/s of interconnect, which Xu put at 2.5 times the Ascend 910C, and both deliver 1 PFLOPS in FP8 and 2 PFLOPS in MXFP4.

Two passages from Huawei's published keynote by Eric Xu at Huawei Connect 2025. The first reads: Fourth, different stages of inference have disparate needs for computing power, memory capacity, and memory access bandwidth. The needs of recommendation systems and model training vary too. To address these diverse needs, we will offer two proprietary HBMs for the Ascend 950 chips: HiBL 1.0 and HiZQ 2.0. These HBMs will be separately packaged with the Ascend 950 Die. So we'll have the Ascend 950PR chip for prefill and recommendation, and the Ascend 950DT chip for decode and training. Let me show you the details. The second reads: The next chip is the Ascend 950DT, which is optimized for both the decode stage of inference and for model training. These two scenarios have high requirements for interconnect bandwidth and memory access bandwidth. That's where our HiZQ 2.0 HBM comes in, delivering 144 GB of memory and 4 TB/s memory access bandwidth. The chip's total interconnect bandwidth will reach 2 TB/s. It will also provide additional support for FP8, MXFP8, XMFP4, and HiF8. The Ascend 950DT chip will be available in the fourth quarter of 2026.
Huawei's own description of the 950DT, from the published text of Eric Xu's Huawei Connect 2025 keynote, 18 September 2025. Captured 5 September 2026.

Why does the training stay on Nvidia?

Because the two workloads break differently when the hardware is unfamiliar. A training run is one job held together across every chip for weeks: a fault, a numerical difference or a slow collective anywhere in the cluster propagates into the run, and a restart costs whatever was between checkpoints. Serving a model is thousands of short independent requests, so a chip that is a little slower or a little quirkier costs latency on some fraction of them and nothing else. The cheapest place to prove new silicon is therefore the serving fleet.

DeepSeek’s published numbers show what the training side has cost it so far. Its V3 technical report records a cluster of 2,048 Nvidia H800 GPUs, 2.788 million GPU hours in total, and a figure of $5.576m at an assumed $2 per GPU hour. Its later work, including DeepSeek V4, has stayed on that side of the house.

The serving side, by contrast, already has Ascend history, and Huawei is the one that said so. Xu opened the same 2025 keynote by crediting DeepSeek for the year Huawei had just had: “Huawei Cloud has worked around the clock to support DeepSeek’s fast-growing user base and traffic. Between January and April 30, our AI R&D teams worked closely to make sure that the inference capabilities of our Ascend 910B and 910C chips can keep up with customer needs.”

Three jobs, the parts Huawei names for each, and where the reported plan puts them. Looping.

The site sits inside a hub Beijing designated in 2021

Ulanqab is not an improvised location. On 29 December 2021 the National Development and Reform Commission published reply 发改高技〔2021〕1843号, agreeing to start construction of a national hub node of the integrated national computing network in the Inner Mongolia Autonomous Region. The reply plans a Horinger data centre cluster whose start-up zone is bounded by the Horinger New Area and the Jining Big Data Industrial Park, Jining being Ulanqab’s central urban district. It sets the cluster two hard numbers: average rack occupancy of at least 65 per cent, and power usage effectiveness held below 1.2.

The document also states the reasoning, which is climate and energy rather than proximity: the hub is to use the region’s advantages in climate, energy and environment to build low-carbon data centre clusters, serving Beijing-Tianjin-Hebei for latency-sensitive work and regions such as the Yangtze River Delta for everything else.

The National Development and Reform Commission's published reply 发改高技〔2021〕1843号, dated 29 December 2021, agreeing to start construction of a national hub node of the integrated national computing power network in the Inner Mongolia Autonomous Region. Clause three plans the Horinger data centre cluster with a start-up zone bounded by the Horinger New Area and the Jining Big Data Industrial Park. Clause five sets an average rack occupancy of no less than 65 per cent and power usage effectiveness below 1.2.
The NDRC reply that made Inner Mongolia one of China's national computing hub nodes, on the commission's own site. Captured 5 September 2026.

Other companies are building at gigawatt scale in the same city, on their own account. Envision announced on 6 August 2026 that it had commissioned its own Galaxy Campus at Ulanqab, a 120,000 square metre site designed to scale beyond 2GW and to support up to one million AI accelerators, which its own release puts at one million PFLOPS of AI compute at full build-out. “The next frontier of AI is infrastructure,” Ricky Zheng, general manager of Envision’s AIDC, said in that announcement. “As AI models become larger and more compute-intensive, the limiting factors are increasingly power availability, network performance and energy efficiency.” That release is Envision’s own project and names no customer and no chip supplier, so it stands beside DeepSeek’s reported plan rather than behind it. What the two share is the city, and the reason to be in it.

The scale, measured in Huawei’s own units

Huawei sells the 950DT in pods. Its Atlas 950 SuperPoD, dated to the same fourth quarter of 2026, holds up to 8,192 Ascend 950DT chips across 160 cabinets and about 1,000 square metres, and 64 of those pods make an Atlas 950 SuperCluster of more than 520,000 chips rated at 524 EFLOPS in FP8.

Put the reported order into those units and it is a little under twenty Atlas 950 SuperPoDs, or roughly 30 per cent of one SuperCluster. On Huawei’s per-chip figure of 1 PFLOPS in FP8, 160,000 chips come to about 160 EFLOPS. Those three are our arithmetic on Huawei’s published numbers rather than figures Huawei or Bloomberg has printed.

The constraint Bloomberg names is memory. High-bandwidth memory is the part of an accelerator China has had the hardest time sourcing, which is why Huawei packaged its own HiZQ 2.0 with the 950 die in the first place, and an order of this size sits against a year’s output rather than a quarter’s.

What the US rule currently allows

Nvidia is available to Chinese buyers on narrow terms. A BIS final rule effective 15 January 2026, published at 91 FR 1684, moved exports of certain accelerators to China and Macau from a presumption of denial to case-by-case review. The rule draws the line by performance rather than by product name: it covers commodities with a total processing performance below 21,000 and total DRAM bandwidth below 6,500 GB/s, “such as the NVIDIA H200 or AMD MI325X”.

Four conditions come with it. The rule requires that the chip is commercially available in the United States when the rule publishes, and that the exporter certifies sufficient US supply, that production for export to China will not divert foundry capacity from US customers, that the recipient has demonstrated sufficient security procedures, and that the item passes independent third-party testing in the United States to verify its performance. Anything above those two thresholds stays under a presumption of denial.

That is the route by which any new Nvidia silicon reaches a Chinese buyer now, and Washington can narrow it whenever it chooses. DeepSeek’s own published training hardware, the 2,048 H800s of the V3 report, was bought long before the rule existed. A serving fleet on domestic chips holds its value on those terms: it is the half of the workload that keeps running when a licence is refused.

Infographic titled DeepSeek and the Ascend 950DT, dated 5 September 2026. Reported by Bloomberg on 4 September 2026, from unnamed people: at least 160,000 Ascend 950DT accelerators; a data centre at Ulanqab, Inner Mongolia; the chips are for running models, with training staying on Nvidia; delivery could take more than a year; 950DT output is held to the low hundreds of thousands this year by high-bandwidth memory shortages; neither DeepSeek nor Huawei has commented. Confirmed by Huawei's own keynote of 18 September 2025: the Ascend 950DT is optimised for the decode stage of inference and for model training; HiZQ 2.0 memory of 144GB; 4TB per second of memory access bandwidth; 2TB per second of total interconnect; 1 PFLOPS in FP8 and 2 PFLOPS in MXFP4; available in the fourth quarter of 2026; an Atlas 950 SuperPoD holds up to 8,192 of the chips; an Atlas 950 SuperCluster holds more than 520,000 and is rated 524 EFLOPS in FP8. Confirmed by the National Development and Reform Commission reply of 29 December 2021: Inner Mongolia approved as a national computing hub node; the Horinger cluster start-up zone includes the Jining Big Data Industrial Park; average rack occupancy at least 65 per cent; power usage effectiveness below 1.2. Confirmed by DeepSeek's own V3 technical report: 2,048 Nvidia H800 GPUs; 2.788 million GPU hours; 5.576 million dollars at an assumed 2 dollars per GPU hour. Confirmed by the BIS final rule at 91 FR 1684, effective 15 January 2026: case-by-case review for exports to China and Macau; the threshold is total processing performance below 21,000 and total DRAM bandwidth below 6,500 GB per second, such as the NVIDIA H200 or AMD MI325X; four exporter certifications required.

How far the evidence goes

The chip count, the location, the split between serving and training, the delivery timetable, the memory constraint and the approach to Beijing all rest on Bloomberg’s 4 September report and its unnamed sources. Everything under them is documented: the 950DT’s design, specification and ship quarter come from Huawei’s own published keynote; Ulanqab’s status as part of a national computing hub comes from the NDRC’s own reply; the training figures come from DeepSeek’s own technical report; and the licensing position comes from the Federal Register. Bloomberg.com answers a datacentre request with a 403, so this desk read the report through pickups carrying its wording rather than from the page itself.

Memory supply decides whether this becomes a cluster

Huawei has a chip designed for exactly the job DeepSeek wants done, dated to a quarter that starts next month, sold in pods that would take twenty of themselves to fill the order. Memory sets the pace from here: HiZQ 2.0 output governs how many 950DTs exist at all, which is why Bloomberg’s delivery window runs past a year, and why, on the same report, DeepSeek has asked Beijing to help persuade Huawei to allocate it more chips and sooner.

The test from here is industrial. Huawei has already shown it can carry DeepSeek’s traffic on Ascend, and said so in its own keynote. Whether it can hand one customer the better part of a year’s production of its newest part, while every other Chinese buyer wants the same silicon, is the question the fourth quarter answers.

Sources

  1. Bloomberg: DeepSeek Plans Big Huawei AI Chip Order to Power New Data Center (4 September 2026)bloomberg.com
  2. Huawei: Groundbreaking SuperPoD Interconnect, keynote by Eric Xu, Huawei Connect 2025huawei.com
  3. NDRC reply 发改高技〔2021〕1843号 approving the Inner Mongolia national computing hub nodendrc.gov.cn
  4. BIS final rule, Revision to License Review Policy for Advanced Computing Commodities, 91 FR 1684federalregister.gov
  5. Department of Commerce revises license review policy for semiconductors exported to Chinabis.gov
  6. DeepSeek-V3 Technical Report (arXiv:2412.19437)arxiv.org
  7. Envision commissions Galaxy Campus in Ulanqab (6 August 2026)prnewswire.com