Field brief · hardware · September 2026

DeepSeek did not beat America with $6 million

They beat a kneecapped interconnect. V4.1 Flash is the latest exhibit — and the offshore-GPU rumor is the weakest part of the story.

DeepSeek H800 V4.1 Flash export controls 10 Sep 2026
Named V3 training job
2,048 H800
Final run, not the company. Paper: 2.788M GPU-hours.
V4.1 Flash KV
890 B/tok
8B active in / 16B out · 45T tokens · no SKU named
Huawei 950DT plan
160k
Inference, Ulanqab · Bloomberg · late 2027 / 2028

1. The story that ate January 2025

DeepSeek-V3's paper said the final training run cost 2.788 million H800 GPU-hours. At a $2 rental they printed $5.576 million. The internet heard "$6 million to match GPT-4o." NVIDIA lost $600 billion of market cap in a day. The narrative wrote itself: a hedge-fund side project in Hangzhou had broken scaling laws with a rounding error of compute.

That number was never the company. The paper says so in the same table. It is the official pretrain + context extension + a tiny SFT/RL tag. It excludes ablations, failed runs, data, salaries, and the cluster they already owned. SemiAnalysis called the hardware pile ~$1.6 billion of server CapEx. Both can be true. One is a run. One is a lab.

V4.1 Flash shipped 10 September 2026. Same lab, same habit: a model card that is extremely specific about architecture and silent about the rack. 45 trillion tokens. No GPU named. That silence is the tell. When they were proud of 2,048 H800s, they printed 2,048 H800s.

2. The constraint that is actually real

The H800 is an H100 with the interconnect kneecapped for export control. Same tensor cores. Same FP8. NVLink inside the node drops from 900 GB/s to 400 GB/s. DeepSeek put eight 400 Gb/s InfiniBand NICs on each node to compensate. Cross-node is still slower than the NVLink they were not allowed to have.

Mixture-of-experts training is an all-to-all problem. Every token has to find its experts on other GPUs. On a healthy H100/H200/B200 node, a lot of that stay on NVLink. On H800, a lot of it spills onto InfiniBand. The V3 paper's own ratio is ugly: compute : communication ≈ 1:1 if you do nothing clever.

So they did something clever.

DualPipe overlaps a forward chunk of one micro-batch with a backward chunk of another, and hides the all-to-all inside the math. Node-limited routing forces each token onto at most four nodes, so the IB trip is taken once and NVLink fans the token out inside the box. They reserved 20 of 132 SMs on each H800 as a software NIC. They wrote PTX by hand so the communication kernels would not thrash L2.

That is not "we are more talented than OpenAI." That is "our NVLink is 44% of yours, so we will not wait for it."

The second constraint is simpler: they cannot buy B200 or GB300 in volume. Blackwell is the US serving story. China-legal Hopper plus Huawei is theirs. H20 — the memory-fat, compute-thin China SKU — is a decode chip. 96–141 GB, 4 TB/s, anemic FLOPs. If your model is memory-bound at decode, H20 is not a consolation prize. If your model is a dense 400B that wants tensor cores, it is.

So they stopped writing dense 400Bs.

3. What the constraint bought

Every DeepSeek generation is a smaller KV cache and a sparser forward.

GenerationThe trickWhy it existsWho else had a version
V2 (2024)MLA — compress KV into a latentDecode is an HBM tax. Smaller cache = more users per GPUGQA / MQA at Google and Meta. MLA is the more aggressive compression
V2–V3Fine-grained MoE + auxiliary-loss-free routingActivate 37B of 671B. Training FLOPs collapseGShard, Switch, Mixtral. DeepSeek made the experts smaller and the load-balance less destructive
V3FP8 mixed-precision training at 671BHopper tensor cores are wasted in BF16. They published the recipeUS labs were already doing FP8. DeepSeek was first to open it at this scale
V3DualPipe, node-limited routing, custom IB+NVLink kernelsH800 NVLink is cutNobody else had to. That is the point
V3 servingPrefill / decode split. 32 GPUs for prefill, 144–320 for decodePrefill is compute-bound. Decode is HBM-bound. Mixing them wrecks bothPart 2 of the serving series. DeepSeek ran it in production in 2025
V3.2DSA — sparse attention, O(Lk) not O(L²)Long context was going to bankrupt them on H800Sparse attention papers exist. They continued-trained it onto a production model
V4CSA + HCA, 1M context, FP4 expertsSame tax, one more zero on the windowHybrid attention is now a fashion. They shipped it
V4.1 FlashCED — 8B active on input, 16B on output. FP4 KV. 890 B/tokenAgent loops are input-heavy. Prefill was the billThis one is new. Asymmetric encoder/decoder MoE is not GPT-5.6's public story

US labs had MoE. They had GQA. They had speculative decoding. They had prefill/decode disaggregation. What they did not have is a regulator deleting 55% of their scale-up bandwidth and then asking them to match Gemini anyway.

Give Anthropic the same H800 island and they would have invented a DualPipe. Give DeepSeek a GB300 NVL72 and they would have trained a denser model and cared less about 890-byte KV. The architecture is a fossil of the export control.

That is also why the efficiency is ahead of the absolute quality.

MLA and CED make serving cheap. They do not automatically make the model smarter than Opus 5 on a messy 12-hour debug. V4.1 Flash's own card puts it ahead of V4-Pro on coding/agent benches and in the mix with GPT-5.6 Sol / Opus 5.0 on some of those, behind on others. Day-one vendor benches. Treat them as directional. The cost claim is the one with receipts: they keep shrinking the bytes you have to keep hot.

Related density brief: Intelligence Density. Laguna S 2.1 is the US-open exhibit of the same idea. DeepSeek has been running that play for two years.

4. Where the GPUs actually are

Five hypotheses. Ranked.

H1 — Legal NVIDIA they already owned (high)

High-Flyer, the quant fund, bought ~10,000 PCIe A100s in 2021, before the A100 ban. That cluster is in the Fire-Flyer paper. DeepSeek spun out of High-Flyer in 2023 and still shares the iron.

V3's named job: 2,048 H800 SXM, 8 per node, NVSwitch inside, InfiniBand across. Confirmed. V2 was also H800; they never printed the count, only 172.8K GPU-hours per trillion tokens. R1-Zero / R1: 512 H800s (64 nodes), not the 2,048. A100s were used for 30B prep, not the 660B runs.

On 28 February 2025 they published a production snapshot: every V3/R1 web, app, and API request on H800. Peak 278 nodes = 2,224 GPUs. That is a serving fleet the size of the training island. Of course it is. Chat is not a side effect.

H2 — H20 for decode, H800 for prefill and train (high)

H20 is the China-legal Hopper part with the memory. Ant Group's public H20 serving numbers, and DeepSeek's own decode-bound math, both point the same way. The Financial Times (August 2025) said that after R2's Huawei training attempt failed, they went back to NVIDIA for training and left Huawei on inference. The Register said the NVIDIA side of that fallback was H20.

This is the boring, correct picture of 2026 serving: H800 (and whatever H100s they have) for the FLOPs, H20 for the tokens.

H3 — Some H100s, "non-compliant cards" (medium)

SemiAnalysis, January 2025: ~10,000 H800, ~10,000 H100, orders for many more H20, plus the A100s. CSIS restated the H20 pile as ~30,000. H100s after October 2022 are not supposed to be in China.

Liang Wenfeng, closed-door investor call, leaked July 2026: "we can buy some non-compliant cards." He also said they currently had about 20,000 H-equivalent, most of it arrived in the last month or two, and the next buys would be "basically all NVIDIA."

I would not build a national-security brief on a speech-to-text transcript. I also would not pretend a founder saying "non-compliant" on a recorded call is nothing. Gray-market Hopper is the honest middle: some H100s, not a secret B200 supercomputer.

20,000 H-equivalent (Liang, mid-2026) and 50,000 Hopper cards (SemiAnalysis, early 2025) are not a contradiction until you convert units. H20 is ~0.15× H100 on FLOPs and better than H800 on memory. A pile of H20s inflates card count and barely moves training FLOPs.

H4 — Huawei for inference now, training later (high for the plan, medium for today's mix)

August 2025: they could not finish an R2 pretrain on Ascend, even with Huawei engineers in the building. Unstable runs, slow chip-to-chip, CANN vs CUDA. They put Ascend on inference.

June 2026: a Huawei-led group (not DeepSeek's paper) showed full-parameter post-training of V4-Pro on ≥1,000 Ascend 910C at 34.22% MFU. That is "domestic silicon can fine-tune a 1.6T MoE." It is not "V4 was pretrained on 910C." Tsinghua's Liu Zhiyuan told MIT Tech Review the pretrain was still probably NVIDIA.

Liang: Huawei had allocated ~16,000 950-class cards, which he called equal to ~4,000 NVIDIA B-series. "Not a big quantity." Four Huawei ≈ one NVIDIA, plus two years. He was buying them to help Huawei's ecosystem, not because they were the training plan.

4 September 2026, Bloomberg: DeepSeek plans at least 160,000 Ascend 950DT at a ~1 GW site in Ulanqab, Inner Mongolia, for operating models, not training. Partial live late 2027 / early 2028. Huawei production is the gate. Street price chatter puts the order around $2.5B. DeepSeek declined to comment.

So: Huawei is the serving sovereignty play. Training stays NVIDIA as long as they can still get NVIDIA.

H5 — Offshore racks in Southeast Asia (low)

This is the one people want. A ghost GB200 cluster in Malaysia. Weights trained in Singapore, served from Hangzhou.

What is actually in the record:

Could a few thousand banned GPUs have moved through a friendly cloud in SEA? Sure. Is that how you get 45 trillion tokens of V4.1 Flash? You would still need the interconnect, the checkpoint fabric, the 3FS cluster, and a year of not leaking. The simpler story fits the papers: Hangzhou / High-Flyer iron, H800/H20, some gray Hopper, Huawei on the way in for decode.

ByteDance, Tencent, and China Mobile show up as investors or as hosts of DeepSeek APIs for their own customers. They are not shown as the landlord of the 2,048-GPU job.

5. Efficiency ahead, ceiling not

Two different leaderboards got glued together in January 2025.

Efficiency. True, and compounding.

That is why the API is cheap, why the free app did not melt, and why OpenRouter's DeepSeek provider can sit at $0.15 / $0.60 with cache at $0.003. Agent loops re-read the prompt. The cache is the product.

Absolute intelligence. Mixed, and always one generation behind the US closed frontier until it isn't, for a week.

R1 matching o1 was real and fast — because reasoning was a new, cheap paradigm, not because 2,048 H800s are a supercomputer. o3 / GPT-5.x / Opus 5 / Gemini 3.x kept moving. DeepSeek's own V3.2 paper says the open-closed gap widened in 2025 on some axes, then they spent RL until V3.2 was "GPT-5-high class" on the benches they chose. V4.1 Flash's card says it beats their own V4-Pro on coding/agents at a lower price. Independent day-one numbers are still vendor-shaped.

The honest sentence: they turned a handicapped cluster into a serving machine that is structurally cheaper than a dense US model of similar vibes. They did not prove that 2,048 crippled GPUs beat a million H100-equivalents at the frontier. They proved you do not need a million if you refuse to pay for KV.

US labs could have done MLA in 2024. They had less reason to. H100 NVLink was fine. H200 memory was fine. GB200 NVL72 makes expert-parallel almost a rack-local problem. When the wire is fat, you scale the dense model and hire more cluster engineers. When the wire is thin, you invent DualPipe.

Give the constraint five years and you get a different field. That is the part of the DeepSeek story that survives contact with the footnotes.

6. V4.1 Flash, specifically

Shipped 10 September 2026. Causal Encoder–Decoder. 20-layer encoder, 20-layer decoder. Prefill only lights 8B. Decode lights 16B. Native vision. 196B Engram parameters sitting in a lookup table, not in the forward FLOPs. DSpark speculative decoding. Reasoning effort 1–100.

They invite anyone with 2,000 GPUs and a storage cluster to talk about deployment. That is a serving partnership, not a confession about training.

OpenRouter, provider = DeepSeek: first-party. The adapter points at api.deepseek.com. Quantization listed unknown. Implicit cache on. Peak/off-peak pricing matches the official API. Third-party rows on the same slug (Novita, DeepInfra, Fireworks, Morph, GMICloud, io.net) are other people's H200/B200/MI350X. Pin the DeepSeek provider if you want Hangzhou's iron. You still will not get a SKU. The last time they told you the serving SKU was February 2025, and it was H800.

FlashMLA's open V4.1 sparse-decode kernel is SM100-only (Blackwell). That is for you, on a B200, with the open weights. It is not proof DeepSeek's API is on Blackwell. They served V3 on SM90 for a year. They can keep doing that with kernels they do not publish.

7. True, false, and the maybes

ClaimVerdictWhy
"V3 cost $6 million."False as a company cost. True as a final-run GPU-hour sticker at $2/hrPaper's own table. Excludes R&D, ablations, the cluster they already had
"They only have 2,048 GPUs."FalseThat is one job. Serving alone peaked at 2,224 H800s in Feb 2025. High-Flyer had 10k A100s in 2021
"Export controls did nothing."FalseH800 NVLink is the whole DualPipe paper. They still cannot buy B200
"Export controls stopped them."FalseThey shipped V3, R1, V3.2, V4, V4.1 Flash anyway
"Constraints made them more efficient than US labs."True for serving and for training FLOPs/tokenMLA, MoE, FP8, DSA, CED, 890 B/token KV
"Constraints made them smarter than US labs."Mostly falseThey catch a generation, then the closed labs spend 10× more RL. Efficiency ≠ ceiling
"They train on Huawei."False for frontier pretrain, so farR2 Ascend pretrain failed. V4 post-train on 910C is a different verb
"They serve on Huawei."Partly true, growingFT split in 2025. 160k 950DT is the 2027–28 plan, not today's API
"They train on AMD."FalseNo DeepSeek paper, leak, or fleet note. Azure/SGLang/InferenceX are third parties
"They have a secret GB200 cluster offshore."UnsupportedRumors exist. NVIDIA said far-fetched. Papers and the leak point at Hopper in China plus gray-market cards
"Liang admitted smuggling."Maybe"Non-compliant cards" on a leaked call. Not a shipping manifest
"Open-source DeepSeek means you know the rack."FalseWeights are MIT. The rack went dark after V3

8. What I would actually believe

If I had to bet a rack:

Training V4.1 Flash happened on NVIDIA Hopper in China. A few thousand to low tens of thousands of H800/H20, maybe some H100s Liang will not put in a paper. Not 2,048. Not 200,000 Huawei. Not a Malaysian GB300. The 45T token run is expensive even with MoE; it is not "we rented 8× B200 for a weekend."

Serving V4.1 Flash on the official API (and therefore OpenRouter's DeepSeek provider) is still mostly Hopper — H800 and H20 — with a rising Ascend 910C slice. The 160k 950DT order is how they get off NVIDIA for decode in 2028, because that is the workload China can actually supply.

The efficiency lead is the real story. DualPipe, MLA, DSA, CED, 890-byte KV. US labs will copy all of it. They already copied MoE and GQA. The copy time is how you measure whether the constraint was a gift. If Gemini 4 ships a CED-like split in six months, DeepSeek donated a paper. If serving costs in the US drop 4× because everyone stole CSA2, export controls accidentally funded the global inference stack.

The lie was never "DeepSeek is good." The lie was "DeepSeek is cheap because they are small." They are good because the wire was bad, and they had already spent a billion dollars of fund money on GPUs before Twitter learned their name.

9. Caveats

10. Sources