# Four to One: Auditing the Leaked DeepSeek Notes

Ryan Cunningham · August 6, 2026

Source: https://x.com/rydcunningham/status/2085434224178303376

> E² edition, September 25, 2026. Original wording through Appendix B; images and embedded posts omitted, captions retained. Appendix C is referenced but absent from the retrieved source. Original citation numbering is retained. Annotations and the calibration audit are separate, at /resources/references/four-to-one/claims.

**My friend **[**@ruima**](https://x.com/ruima)** asked me to weigh in with her on the leaked DeepSeek notes. We published the original piece on the **[Tech Buzz China Substack](https://techbuzzchina.substack.com/p/the-deepseek-leak-and-chinas-ai-hardware)** - this is the more technical deep-dive companion piece, published here on X and Substack.**

**Also, launching Machine Yearning's research tools "Cortex" and "Overclock" to a private beta. If interested, drop a comment or send me a DM.**

---

## TL;DR

**Liang's 4:1 (GB300 : Ascend 950) is a reasonable planning ratio, especially for inference, but it is not a general equivalence. Gaps certainly exist on training throughput due to MFU and goodput deltas - but these are improving.**

> [Figure omitted]

1. **The 16,000-chip allocation would be a very capable inference machine, but suitable only for training tests.** In an Atlas 950 SuperPoD, Overclock projects ~1.9 Ttok/day of generated DeepSeek V4 Flash tokens (~0.5 for V4 Pro), ~33 Flash (~8.6 Pro) Ttok/day of generated tokens at measured 16:1 input mix. This derates to ~15–23 Ttok/day against ~1.1–10 Ttok/day of demand-side estimates. Plenty for serving and post-training, but not a pre-training allocation.
2. **LongCat 2.0's "50,000 domestic accelerators" claim coheres** on 910B silicon: ~1.0×10²⁵ FLOPs, a base case of ~1.6M accelerator-days, 22–98 days of wall-clock on that fleet.
3. **Kimi K3 (~2.2×10²⁵ FLOPs) fits comfortably on a reported 20,000-Hopper cluster (~47 days),** and most likely was not trained on domestic silicon.
4. **A Mythos-scale model for DeepSeek** would take ~99 days on a 200,000-card 950 figure at today's goodput. 50,000 GB300s do it in about a quarter of the time.

The Overclock model reproduces Tensor Economics methods to 2–5%, runs 1.4–2.2x above measured deployments, and independently lands a median 1.81x against SemiAnalysis InferenceX measurements. Full assumptions available in the Methodology section.

---

## Introduction

The leaked notes from the DeepSeek investor meeting provided a lot of interesting fodder for discussion. Rui and I published our notes on Tech Buzz China last week and attempted to address the claims one-by-one, focusing mainly on the "4:1" Ascend-to-Blackwell equivalency ratio.

Leaks have a way of getting repeated until they harden into common knowledge, so it was important to me that I build a replicable and auditable tool for assessments. I also figured this would be relevant for many other projects in my backlog.

Every quantitative claim in this piece is either computed from first principles in Overclock - my model × accelerator × energy simulator, introduced below - or pinned to a first-party source. Where a claim is unconfirmed or rumored, I try to label it.

Fair warning, this is probably better served as a white paper rather than an article. Maybe I'll make one later.

Enjoy.

### Context

First, lets contextualize the claims we're evaluating. Numbers about chip counts and cluster sizes get thrown around a lot, but I imagine for most people (certainly myself), it's not always clear what one can actually **do** with a certain amount.

What constitutes a good training cluster, or inference cluster, or hybrid? Is "best" always better or is "good enough" sufficient?

In Liang's own words, there can be no doubt that there is a measurable difference between "best" and "good" - and that size does matter.[1]

> 如果我要训练跟 AI 同样大的模型，应该需要五万张 GB300，或者华为 950，二十万卡。这只是训练，还没有考虑做研究。所以我们跟美国之间最大的差距是在资源上面。 我们现在的资源，以及我们今年之内、接下来几个月的资源，包括马上大量资源，也只足够我们在几十 B 激活的这个规模里面去做更多的实验。因为我在几十 B 激活的这个规模里面，还有很多实验要做，还有很多事情需要搞清楚的。我们应该离能够训 800 B 的模型还比较远，时间还非常多，没有那么多卡。 【。。。】所以我们现在根本就不会考虑，到这么大的一个规模上去跟美国竞争。我们现在首要，还是在我们能够训练得起、能够用得起的规模上，就几十 B 激活的规模上要先把它做好。再等下一步有更多的资源的时候，再把它做到 150 B、156 B，或者 250 B 激活的这个规模。

> To train a model on the same scale as the leading AI models, we would need something like 50,000 GB300 or Huawei 950 chips - essentially a cluster of 200,000 GPUs. That is just for training, not even accounting for the research component. Consequently, the biggest gap between us and the U.S. lies in resources. Our current resources - and what we anticipate having in the coming months, even with a significant influx - are only sufficient for conducting further experiments at the scale of "tens of billions" of active parameters. There is still a great deal of experimentation and understanding to be gained at this level. We are likely a long way off from being able to train an 800B parameter\* model; we simply lack the necessary hardware capacity. [...] Therefore, we are not even considering competing with the U.S. at that massive scale right now. Our immediate priority is to excel at the scale we can afford to train and utilize - specifically, the range of tens of billions of active parameters. Only when we secure more resources down the line will we look to scale up to 150B, 156B, or 250B active parameters.

\*\* Translator's note: Liang most likely means "800 billion active parameters here. DeepSeek V4 Pro is already a 1.6T total parameter model with ~49B active parameters.\*\*

**Why 800B? While undisclosed, the rumor mill suggests that Anthropic's latest models - Opus 5 and Mythos, of which Fable is derived - are estimated to be 5T and 10T parameter models with 500B – 1.2T active parameters, respectively.[3] Liang's 800B assertion would put the active parameter count of such a model squarely within that band.**

There are two claims to dissect from this.

- First, that Liang's roadmap from 50B active params to "frontier" scale (800B+) requires lots and lots of chips that he does not currently have. This is the heart of the "see? export controls are biting - tighten them!" cudgel from security hawks.[2]
- Second, Liang seems to be emphasizing a roughly 4:1 equivalency ratio between GB300s and Ascend 950s (while he doesn't specify which 950, we can assume the 950DT training variant). Similarly, this seems to affirm the "lack of productive capacity" component of the same cudgel.

In a separate article stating their current chip count, Liang again asserts the 4:1 claim more explicitly:[4]

> DeepSeek已經在跟華為合作，拿到了1.6萬張950卡，雖然只相當於4,000張B系列，做不了下一代模型，但足夠幫華為把生態跑通。

> DeepSeek is already collaborating with Huawei and has secured 16,000 950 chips;\* while this capacity is equivalent to only 4,000 B-series chips - insufficient for training next-generation models - it is enough to help Huawei validate and operationalize its ecosystem.

\*\* Translator's note: This is just a guess, but given the closeness between DeepSeek and Huawei I would not be surprised if these were delivered in their new Atlas 950 SuperPoD form factor. If so, the total chip count would actually be 16,384, since each SuperPoD contains up to 8,192 chips.[5]

~16,000 new accelerators sounds very small against American clusters twenty or thirty times that size. And yet, the team consistently seems to deliver "DeepSeek Moments" - so much so they have an eponymous meme for stock market disruption - the latest arriving when V4 Flash 0731 beat the pareto frontier one day after OpenAI cut its small-model prices 80%.[6]

So what is going on? Is this distillation? Government subsidies? "Access to a secretive cluster of 50,000 chips they shouldn't have?" Weirdly racially charged claims of Chinese engineer black magic?

We'll assess the veracity of his claims by working through, on a bottoms-up basis, what kind of workloads a cluster that size could actually support.

- First, we'll narrow in on his "4:1" assertion and demonstrate how this might be quantified (or objected).
- Second, we'll work through the different training and inference serving use cases for a ~16,000-card cluster.
- Third, we'll triangulate the accuracy of these estimates with parallel cases from Meituan LongCat-2.0 and Kimi K3, as well as grounding with InferenceX by SemiAnalysis.

### On Systems-Thinking

Why write a ~10,000 word white paper to prosecute these claims? Because we live in a world where in order to talk about this stuff, I need to prepare a mini Wikipedia.

> [Embedded post omitted.]

Naïve comparisons between chips are unhelpful without context on how they are used. GroqChips are beasts at inference but complete dogwater for training, due to deliberate memory procurement decisions in their construction. So saying "Company A's chips are 10x better than Company B's" without clarification risks confusion. Further, chip-to-chip comparisons obscure critically important advantages that emerge in systems-level deployments... advantages I would very much like to see American ecosystems adopt.

There's a school of analysis, which I call "nodemaxxing"[7], that treats the process node as the fulcrum upon which all leverage exists. The presumption is that whoever fabs the smallest transistors indisputably secures energy-compute supremacy, that competitively advanced silicon is stalled without EUV ("complete collapse")[8], and that everything else is just commentary.

No doubt advanced process nodes are fundamental to building advanced chips. This is not in dispute. The question is whether or not they are singular to it.

> [Embedded post omitted.]

As Oriental Computing CEO Wei Shaojun has acknowledged, the process node is certainly a factor. But it is not the **only** factor.[9]

An investment thesis or ideology which myopically focuses on one sub-component of the chain will not only keep being surprised when temporary deficits are compensated for, but create the very failure conditions they sought to avoid.

Such ideologues ignore the empirical reality that nothing in the laws of physics prevents engineering efforts from eventually replicating that sub-component, on a tech tree they have absolutely zero influence over.

If a regime is captured by and takes action on that ideology, those negative externalities - like our present memory crunch - create objectively worse economic conditions at home and abroad.

> [Embedded post omitted.]

The idea that whoever builds the fastest railcar - and not the most railroads - would attain pole position in a nü-industrial revolution, has always been absurd to me.[10] But I understand this can be difficult to grok if one reasons from theory-first rather than first-principles. Especially when the local echo chamber affirms said theories must be true by some unmeasurable intrinsic quality, or by recursively referring to other, more "authoritative" theories, for support.

> [Embedded post omitted.]

That very pattern - recursive deference - is downstream of Dan Wang's "lawyerly society vs. engineering state" meme (itself a simplification, but a useful one): some societies build rules. Some build systems.[12]

Ideology, like faith, is highly resistant to adjusting priors. In the face of mounting evidence which challenges belief, adherents double-down, declare heresy, censure, sanction, and risk further damage to themselves and others, rather than ask the question: "what if I'm wrong?" Without empiricals, the only way to honestly measure a strategy's effectiveness.

I try (not always successfully) to hold myself, and any reader, to a higher standard of intellectual honesty.

## Enter Cortex + Overclock

In that spirit, I made a mini Wikipedia. I'm opening up my research tools **Cortex** and **Overclock** to a private beta. I built Overclock precisely to run bottoms-up analyses like this article.

### Cortex

I built **Cortex** to more dynamically track entities that come up in my research rather than dumping them all in a Google Drive. It is a knowledge base evolved from the Silicon Vanguard project I began last year[13], grown to cover every accelerator, interconnect, and first-party scale-up system I can document, with provenance attached to every figure. Data is vendor-published, derived, assumed, or conflicted, each labeled as such.

> [Figure omitted: Sample set from an earlier AI Accelerator Atlas build.]

It also contains rich entity-level detail on organizations, companies, facilities, persons, research institutes, etc.. central to these intersectional ecosystems.

> [Figure omitted: Registry: Wiki-like semi-static information on thousands of known entities in the energy-compute ecosystem.]

> [Figure omitted: Connectome: Visualized edges in the network graph between entities. Shown: related organizations, persons, infrastructure, etc. to MetaX 沐曦.]

> [Figure omitted: Reports: Relationships between entities can be rendered into supply chain maps, geographic dependencies, and other reporting.]

### Overclock

**Overclock** is the simulation engine running on the Cortex data plane. It builds energy-compute digital twins (token factories) from first-principles models of the full chain: training FLOPs, decode rooflines, batching behavior, interconnect collectives, fleet economics, and energy generation/storage/transmission modeling.

> [Figure omitted: Cluster-level analysis projects aggregate token factory throughputs on given energy footprints.]

> [Figure omitted: Auditable provenance data on key model and accelerator inputs - every data point has a trace to first-party or verifiable third-party data. Rumors and projections are flagged.]

> [Figure omitted: Report-ready training and inference throughput charting.]

The models are stress-tested against empirical anchors (SemiAnalysis' InferenceX measurements, turbine de-rating curves, published serving deployments, vendors' own reported MFUs). Post-validation on smaller deployments, we can scale these up to supercluster modeling, like China's "National Unified Computing & Power Network," which carries incredibly interesting takeaways for how to build sustainable energy-compute networks. More on this another time.

> [Figure omitted: An excerpt from the China Energy-Compute project, inclusive of 东数西算 AIDC clusters.]

I grew up in the Aaron Swartz era of the Internet, so I'd rather see OSINT raise the floor of discourse than sit behind a paywall.[14] That said, I am still working out how to responsibly publish a few features, like geospatial layers for key infrastructure and interactive digital twins.

If this would be useful in your own work - research, policy, journalism, capital allocation - leave a comment or DM me. The first cohort will be deliberately small.

## Methodology

Every training and inference figure in this piece is computed in Overclock, but primitives for these figures have strict attributions to original sources, derivations, or triangulations. Conventions follow industry references, primarily the Tensor Economics Substack, Baseten's **Inference Engineering**, Google's **How to Scale Your Model**, and HuggingFace's **Smol Training Playbook**. Where the model diverges, I'll say so explicitly.

Appendices A–C carry full derivations, constants, and calibration so anyone can check the work.

### **Training**

**Wall-clock time** for a run is the total FLOPs bill for pretraining, divided by fleet FLOPS capacity:

$$
\text{wall-clock days} = (\text{FPT}(N_{active}) \space × \space T) \space / \space (C \space × \space TF_{dense}(p) \space × \space MFU_p \space × \space g \space × \space 86,400)
$$

Where:

- **FPT**, FLOPs per token, uses the standard transformer approximation of **6 × N\_active** (2 FLOPs / active parameter on forward pass, 4 / backward)
- **N\_active** is activated parameters, not total. For a dense transformer there is no difference.
- **C** is the accelerator count.
- **p** is bit precision the run trains
- **TF\_dense(p)** is the chip's **dense** throughput on that datapath
- **MFU\_p** is "model flops utilization (MFU)", the fraction of TF\_dense(p) sustained in a training run
- **g** is goodput, the fraction of wall-clock spent making forward progress after failures, restarts, and checkpointing

### **Inference**

Overclock solves a decode roofline per model × system pairing. One decode step:

- Reads the resident weights plus every concurrent request's KV cache from memory
- Conducts ~2 FLOPs per active parameter of matrix math
- Exchanges expert-routing traffic over the fabric
- Step time is whichever is slowest of these three, plus an additive per-step floor to account for kernel-launch serialization, collective latency, and attention overhead.

$$
\text{TPOT} = \max(t_{compute}, t_{memory}, t_{interconnect}) + t_{serial}
$$

Where:

- **t\_serial** is an additive per-step floor (8.1 ms node / 1.8 ms rack / 6.8 ms superpod, fitted against solver reference cubes) covering what a pure roofline omits: kernel-launch serialization across n layers, collective latency, attention overhead.
- Per-user tokens/sec is 1/TPOT.
- Prefill is charged explicitly at a 30% prefill MFU
- Time-to-first-token is one request's uncached prompt FLOPs over the whole serving group's compute, and sustained throughput charges the group for prefilling every request it serves.

### **Batching**

Batch amortizes the same weight read across concurrent requests until compute, fabric, or memory capacity binds instead. For each pairing we sweep a batch ladder (1 → 8,192 per accelerator, ~1.5× steps) against a context grid (4K → 1M, clamped to each model's maximum), and select the largest batch that still meets the per-user SLA.

### **Calibration**

- Overclock reproduces Tensor Economics' published Llama serving arithmetic to 2–5%.
- Against measured production deployments (Aleph Alpha's EP72 DeepSeek numbers, Huawei's own CloudMatrix384 serving figures) it runs 1.4 – 2.2× high, which feels acceptable for a first-principles theoretical ceiling
- Delta against SemiAnalysis InferenceX single-node measurements (DeepSeek V4 Pro on B300/B200) is 1.8x p50, 1.4x - 2.2x p10–p90. This is consistent across concurrency figures.
- One residual gap is disclosed in Appendix B: per-user interactivity estimates degrade too slowly with batch (median 2.8× above measured), so per-user speeds quoted at very large batch are optimistic because of the SLA constraint

### Assumptions

These are the primary assumptions underlying the scenarios in Overclock:

> [Figure omitted]

Most-used references:

- Piotr Mazurek, [@tugot17](https://x.com/tugot17) Tensor Economics[14]
- Philip Kiely, [@philipkiely](https://x.com/philipkiely) **Inference Engineering** (Baseten)[15]
- "How to Scale Your Model" (Google/JAX)[16]
- "The Smol Training Playbook" (HuggingFace)[17]

## **"4:1" - Spec Comparisons**

Let's begin by first comparing Ascend and Blackwell at the chip-level.[19][20]

### **1.** **Compute**

On raw TFLOPS, the Ascend lineup is **just** outside that confidence band for the Blackwell series. They are roughly 2:1 with H200s, and strictly better than the aborted H20 chip lineup, but the 950DT is more like 4.5 - 5.0:1 with Blackwells.

> [Figure omitted]

Regarding the 910B, we have at least 3 reported SKUs in Ascend hardware reports which vary by ±20% in compute performance:

> [Figure omitted: Our training math uses 400 TFLOPS (910B1), which is slightly charitable, though we only use 910B for the Meituan LongCat-2.0 training analysis. Note that outside of Huawei, primary estimate of the 910C performance are third-party die-shot analyses.[23] Treat every 910C figure in this piece, ours included, as an estimate.]

It's also true that the 950PR and DT achieve lower dense 16-bit throughput than their predecessor (910C) - a regression critics have seized on.[30] This is, however, a deep mischaracterization of the purpose of these chips. DeepSeek has made FP8 its native training standard, and much of the Chinese lab ecosystem seems to be following[31]. 910Cs have no native FP8 support at all, and require quantizing FP8 to INT8 in order to host DeepSeek models for inference[28] [32].

This is not especially difficult, but it does add friction for a chipmaker seeking lab adoption for training.

### **2. Memory**

On memory bandwidth, the 950DT (with 4 TB/s planned[24]) would be comfortably within Liang's confidence band. A 950DT is roughly 2:1 with the GB300 on a per-GPU basis, but as a point of occasional confusion, the Grace Blackwell GB200/GB300 series are actually dual-GPU "superchip" which share the same CPU. It's not always clear in conversation if, when people mention the GB-series, they're referring to per-GPU or per-superchip specs.

> [Figure omitted: Per-GPU memory capacity and bandwidth.]

If Liang meant the superchip as the atomic unit instead - his claim still holds up. 50,000 GB300 superchips would require roughly 200,000 950DTs to match in memory capacity and bandwidth.

### **3.  Interconnect**

This is when we start to consider system-level attributes in the comparison.

- NVLink 5.0, introduced with the Blackwell generation, hits 900 GB/s of unidirectional scale-up bandwidth into a 72-GPU coherent domain.
- LingQu 2.0 in the Atlas 950 generation quotes 392-840 GB/s per card depending on chassis, and approaches NVLink-class per-card bandwidth when scaling the domain to 8,192 cards.
- While NVIDIA does market a DGX SuperPOD product, this is not a true high-coherence domain like the Atlas 950 or CloudMatrix 384.[24]

> [Figure omitted]

### 4. System vs. System

Real-world throughput depends a lot on the exact nature of the deployment. Are these 16,000 chips split across 2,000 node-scale Atlas units, or 2 Atlas 950 SuperPoDs?

A single instance of a trillion-parameter MoE lives comfortably inside one high-bandwidth fabric, while node-scale deployments have to lace chassis together over weaker interconnect.[34]. This is part of the reason why running "full-blooded DeepSeek" on a single node was a common marketing meme at WAIC 2025 booths.

You can see these differences in the below parity matrices, which shows 1. System-level performance and 2. Normalized system-level performance, respectively. **Notice how on iso-card node setups (8x cards), Huawei nodes are consistently in the ~4:1 band, but full scale-ups meet or exceed parity with a scale-up NVIDIA system.**

> [Figure omitted: Per-system equivalency ratios. Maximum scale-up systems meet or exceed parity with max scale-up NVIDIA systems on inference and training throughput.]

Scenario assumptions:

- 4K context at the measured ~16:1 input:output mix, 50% prefix-cache hit
- 20 tok/s/user SLA, with each system loaded to the largest batch that still meets it
- Serving groups sized at each system's throughput optimum
- **Sustained** throughput

> [Figure omitted: Per-system performance, normalized by card count.]

This is, of course, obviously not the same when normalized on a per-card basis (Atlas 950 SuperPoDs can be deployed in 1,024 to 8,192x card setups), but Liang himself frames substitutability at the **supernode** level, not the chip level (emphasis mine) [1]:

> 我觉得在这一点上，英伟达是在掘自己的坟墓。华为的**超节点**，华为的 950 **超节点**，在性能和价格上可以完全平替英伟达的 GB200、GB300。价格肯定要贵，但贵得有限。价格贵百分之五十、百分之一百，贵百分之一百无所谓，贵百分之两百都无所谓。

> I feel like Nvidia is digging its own grave at this point. Huawei's **super node**, Huawei's 950 **super node**, can completely replace NVIDIA's GB200 and GB300 in terms of performance and price. The price is definitely expensive, but only to a limited extent. It doesn't matter if the price is 50% or 100% more expensive. It doesn't matter if it's 100% more expensive. It doesn't matter if it's 200% more expensive.

### 5. Energy

Finally, there is the energy budget to consider. As SemiAnalysis reported last year[35], the CloudMatrix384 carries a 2.5x performance-per-watt penalty against the GB200 NVL72, which is a brute-force trade in its larger scale-up domain (384 vs. 72 chips).

> [Figure omitted: Per-system equivalency ratios on energy-compute and energy-memory. "Open Image in New Tab" if the ratio is cropped.]

The Atlas 950 SuperPoD reduces that gap by a large jump, achieving near-parity on FLOPS/Watt **and** delivered tokens per joule with rack-scale NVIDIA systems (1.3 - 1.5x), and actually **exceeding** NVIDIA in energy-memory efficiency (0.5 - 0.6x).

These advantages, as well as Liang's comments, speak to the prevalence of maximal scale-up systems at WAIC 2026. The conference was not really a showcase of new chips (with some exceptions[9]), but of many existing players demonstrating capabilities in greater system sizes (Huawei counts 750+ Ascend supernodes already deployed commercially[36])... provided you have the energy budget to support it. At full-kit, the Atlas 950 SuperPoD is a beefy boi drawing ~4.1 MW IT power / 5 MW all-in power.

Having said that, tapping the sign again...

> [Figure omitted]

At any rate, Liang's 4:1 assertion loosely holds across card and system scale-up domains, depending on the measurement. Arguably, it's even conservative in energy-compute terms. But it would be dishonest to state as much on training costs right now.

## Assessments

Thems the specs. Let's work through some bottoms-up assessments.

### Inference: Theoreticals

If 16,000 950DT chips were used as an inference fleet, they would yield up to the following throughput capacity while still meeting 20 tok/s/user SLA:

- V4 Flash: ~32 Ttok total, ~1.9 Ttok output tokens  / day
- V4 Pro: ~14 Ttok total,  ~0.5 Ttok output tokens / day

> [Figure omitted: Aggregate throughput is dominated by input tokens due to a 16:1 IO ratio from OpenRouter. For batching, we load the fleet to the largest batch that fits memory and meets SLA. [38] [39]]

The V4 Pro throughput is remarkably similar to their R1 / V3.2 model series, speaking to the team's emphasis on scalable architectures rather than just making V3 / R1 bigger. V4 Flash yields triple the inference throughput that their previous models could.

> [Figure omitted]

### Inference: Empiricals

Let's check our estimates. Would this be enough to support current demand?

**Top-Down: Tokenomics**

We can use OpenRouter for a top-down triangulation. OpenRouter's public insights put DeepSeek at ~20% of all platform tokens through H1 2026 - its top model since mid-May. Reports indicated DeepSeek's "top three models" comprised ~1.1 Ttok of routed tokens per day, across input, reasoning, and output tokens. However, we don't know what proportion of DeepSeek's total that is. [40] [41]

API prices and revenue estimates are useful. Against an estimated annual run-rate of $400-$500M[42] at published API pricing[43], this implies ~3-10 Ttok per day across API traffic, assuming their 70% Flash / Pro split. This is, of course, only partial as there are tons of free web and mobile chat users that constitute unpaid token traffic (Similarweb alone counts 335M quarterly web visits[44]).

✅ Even discounting our projections by our calibration bands (1.4-2.2x, see Appendix B) that's up to ~23 Ttok of deliverable capacity per day. This is comfortably above the demand range, but not ridiculously so. Seems in the right ballpark.

**Bottoms-Up: SemiAnalysis Actuals**

We also have InferenceX by SemiAnalysis for a more grounded comparison. These are helpful since the tests are auditable, but we have to align our scenarios first.

In the best case scenario, InferenceX shows a highly optimized DeepSeek V4 Pro deployment (FP4, SGLang, wide expert parallelism) running at 12,741.1 tok/s/GPU on a GB300 NVL72 rack-scale system. 60 of the 72 GPUs were allocated for prefill, and 12 for decode-only. However, since this highlighted deployment was "disaggregated," it more effectively balances MFU in a system rather than throw all resources at one problem or the other. [45]

For calibration purposes, we'll compare our single-node theoreticals to stay more apples-to-apples. We use the default InferenceX case of an 8:1 IO ratio (~8K input tokens), single-node setup, and run their selected concurrency values.

> [Figure omitted: Single-node (DGX B300/B200) Overclock projections against InferenceX actuals. Overclock overstates throughput capacity by ~1.8x.]

> [Figure omitted: Single-node throughput, interactivity (tok/s/user) basis. We restrict scenarios to 20 tok/s/user SLAs, hence the divergence in greater concurrencies.]

✅ In the single-node analysis, Overclock estimates ~1.81x above InferenceX empiricals at the median (1.39 - 2.18x, p10-p90). This is consistent between both Blackwell chip variants. There are some residual gaps in per-user interactivity, especially with our strict adherence to a 20 tok/s/user SLA (InferenceX drops below 5 at the largest batches), but generally these seem tight enough for a first attempt.

### Post-Training

What about post-training? Well, most of post-training more strongly resembles inference workloads:

- RL rollouts, reward-model scoring, evaluation harnesses, and synthetic data generation are inference workloads wearing a training hat. They're decode-bound, spread across replicas, and inherit every advantage the fleet already has.
- Modern RL post-training spends most of its wall-clock generating rollouts, not applying gradients.
- SFT and full-parameter continued training require gradients, optimizer state, and an all-reduce across the whole cluster every step. This inherits training risks - the goodput problem and the all-to-all stability question - that we'll get into with LongCat.

Of course, all these qualifications may be moot, since in the time between Rui's original publishing and this piece, DeepSeek V4 Flash 0731 benchmarks became public, alongside which they revealed their full-parameter, end-to-end post-training of DeepSeek V4 Pro on a CloudMatrix 384 (34.22% MFU).[46]

> [Embedded post omitted.]

✅ So yeah, a 16,000-chip Ascend fleet is probably more than sufficient for inference and post-training by this point. DeepSeek has not yet made this claim for pretraining runs though.

### **How to Train Your LongCat 2.0**

Let's go right into our empirical training case with Meituan LongCat-2.0. Here's what we know:

- LongCat 2.0 is a mixture-of-experts (MoE) model with 1.6 trillion total parameters, activating ~48B/token (33B-56B dynamically, depending on token complexity) - a 97% sparsity ratio[48][49].
- Reportedly it was trained on ~50,000 "domestic accelerators"[50], which is a considerable achievement in domestic training capability.
- It's a competent model, though not a frontier one - roughly Gemini 3.5 Flash-Lite or Nemotron 3 Ultra in general intelligence (Artificial Analysis' Intelligence Index puts it at 33, top-third of open-weights peers).[51]
- At 48B active against 35T tokens, Overclock puts the run at ~1.0 x 10²⁵ FLOPs[52]. Large, but not frontier-large. Had LongCat been a dense 1.6T model, the same corpus would have cost ~3.4 x 10²⁶, thirty-three times more[53].

As it happens, a LongCat-2.0 training run on 910Bs lands naturally with their claim of "millions of accelerator-days" of training, as well as their machine count ("up to 48 machines each, with high-bandwidth all-to-all connectivity internally and RoCE fabric between superpods"). We assume 910Bs because the accelerators "have significantly less per-device memory than an H800 (80 GB)", which basically leaves the 910B as the primary option. [48] [54] [55]

> [Figure omitted: Comparative accelerator-days for a LongCat-2.0 training run at 30% MFU and a 40-75% goodput band. On 910B-class chips the run costs 1.1-4.9M accelerator-days (base ~1.6M), or 22-98 wall-clock days across 50,000 chips.]

Interestingly, Meituan also claims "no rollbacks or irrecoverable loss spikes" across the full run[48], backed by described mechanisms like deterministic operators and bit-flip detection. If true, that would indicate genuinely mature training operations on Ascend chips.

- Huawei has separately reported a 70% reduction in daily failure rate and a 1.5x MFU improvement on Ascend training[56], though without a quantified baseline neither figure tells you where they started. We've seen anywhere from 30 to 50% MFU across the DeepSeek and Pangu literature.[57][58][59]

**⚠️ One thing to note for this and all other training run estimates - we are only modeling the final training run here, which is the smallest line item on a Gantt chart. Training labs undergo ablations, false starts, multi-stage training, and post-training stages, which are incremental to whatever the ultimate final training run ended up costing. Lab needs are better measured by fleet-months rather than run-days.[33]**

✅ On balance, these claims are technically feasible and not obviously exaggerated - FLOPs, fleet size, memory constraint and topology all cohere on one hardware generation. Having said that, there's more we could dig for, especially on failure categorization. In my opinion, Meta's root-cause categorization in their 2024 Llama 3 paper is the transparency baseline here.[60]

### Training: "Who the Fuck is Kimmy"

> [Embedded post omitted.]

The second triangulation case is Moonshot's Kimi K3, re-derived here on a full grid of possible model x accelerator combinations. Rumors have popped up surrounding how it was trained (Hopper? Blackwell?) - I will not adjudicate them here, but I will run the numbers.[62][63]

Here are our inputs and underlying assumptions:

- All model parameters derived from the HuggingFace model card[64].
- Training tokens not disclosed; we assume 35T, matching the LongCat 2.0 corpus
- 6 × 104B active params = 6.24×10¹¹ FLOPs per token, ~2.2×10²⁵ for the run. **(Note: Our original piece printed 2.4×10²⁵, which carried a +10% attention uplift the 6N convention retires; K3's KDA layers are linear-attention anyway, making any quadratic uplift second-order at training sequence length. See Appendix A)**

We sweep this over 8K–100K accelerator fleets across several chip species, including Huawei's roadmap projections.

> [Figure omitted]

✅ A 20,000-card H200 cluster certainly fits a reasonable timeline (~47 days) for a final training run, as would any smaller Blackwell cluster. 910Cs are a **possible** basis, but not as likely given the window between Kimi 2.5 and K3 releases.

### Training: A Future "DeepSeek Moment"

For our last assessment, let's run the numbers on a hypothetical DeepSeek mega-model - call it **DeepSeek V5** for now - a hypothetical successor architected like the V4 family but scaled to Mythos-class, at 800B active parameters on 10T total.

At the 6N convention and the V4 Pro token basis (32T), the final pre-train bill is 6 × 800B × 32T ≈ 1.5×10²⁶ FLOPs.

> [Figure omitted]

50,000 GB300s for a final training run of a Mythos-class model (~53D) seems like a perfectly reasonable pill to swallow, especially when adding the other aforementioned training swim lanes. 50,000 950DTs? Not so much. Epoch's data puts the economically rational ceiling on a single run near ~9 months - past that, their estimate is that waiting for better hardware and algorithms beats continuing to train[33].

DeepSeek would require a larger allocation - perhaps 200,000 950DTs, or 100,000 960s - to train a Mythos-class model in an acceptable timeframe.

## Can China Train Frontier Models on Domestic Chips?

The underlying question for this exercise is, "could Chinese labs plausibly train next year's frontier models on purely domestic chips?"

Maybe. First, we have to align on what constitutes "frontier":

- Frontier in the chips used to train it? **All else equal, I do not think this matters to downstream users.**
- Frontier in size? **This matters, but is no longer a guarantee of supremacy.**
- Frontier in performance? **This matters to most observers.**
- Frontier in performance-per-dollar? **This matters the most.**

### Size Matters

The instinct of the size camp is to anchor on a FLOPs threshold. Biden's 2023 Executive Order picked 10²⁶ FLOPs as its reporting trigger[65], and export-control advocates have leaned on it as shorthand for "frontier-class" ever since. If you ask about size, Meituan LongCat-2.0 is our clearest case so far, but still not comparable due to extreme sparsity.

The caveat is that LongCat is frontier-size without being frontier-FLOPs or frontier-performance, and lands closer to Gemini 3.5 Flash-Lite or Nemotron 3. Today's actual front-runners - DeepSeek, Moonshot - still reserve Nvidia for pre-training, which I expect them to continue for a little while longer, as Liang suggests.

### Performance Matters

If you ask about performance, one could argue we're already there given DeepSeek's V4 Flash 0731 technical report. Though V4 wasn't **pre**-trained on Ascend NPUs, it's not like the rules of physics prevent that from eventually happening. It is a question of diagnostics (MFU, goodput), scale, and time.

- In August 2025, DeepSeek reportedly stumbled trying to train their next model on Huawei silicon[67].
- By June 2026, Meituan had put LongCat-2.0 through a full pretraining run on ~50,000 domestic accelerators and reported no rollbacks or irrecoverable loss spikes[50].
- By July 2026, DeepSeek itself was running full-parameter post-training of V4 Pro end-to-end on a CloudMatrix 384 SuperPoD at 34.22% MFU [46].

In less than a year, we've gone from "can't finish a run" to "finished a trillion-parameter run cleanly" to "have successfully post-trained a model on 910C NPUs to the current frontier of performance-per-dollar."

> [Embedded post omitted.]

## **Verdict**

✅ A predominantly domestic frontier-**performance** run in the next twelve months is plausible under the following assumptions:

- 50,000-100,000 950-class chips reserved for most of a year (~3-5 month training runs)
- High-sparsity architectures designed around the hardware's limitations, like LongCat

❌ A frontier-**scale** run in the next twelve months is less plausible. Our DeepMythos wants ~219,000 950DTs for a 90-day run, or ~54,000 for thirteen months, which exceeds the theoretical 9-month Epoch ceiling[33].

- While 950DTs are not 4:1 with Blackwell on training, the vendor-announced 960 specs do narrow the gap to that point. This, along with Huawei chip production schedules, would put late 2027 / early 2028 more likely.

Bottom line, there's no universal minimum chip count for a frontier run. There is a budget equation for it, and how long you're willing to wait.

## Closing Thoughts

Does size matter? I think so. But it is not a guarantee of supremacy. I think this is more closely correlated to a model's performance-per-dollar - especially if they're already at the performance frontier on an absolute basis. We've already seen in other industries (NEVs in particular) how well this objective function performs in 内卷 dynamics.

It seems to me an ecosystem like this creates conditions suitable for energy-compute apex predators: organisms that survive brutal domestic competition on thin margins, but with abundant infrastructure resources.

Our domestic ecosystem, contrastingly, breeds the opposite - large organisms with margin abundance, but energy scarcity that drastically reduces the organism's velocity.

This is not a full-throated endorsement of either path, there are tradeoffs which must be studied and weighed.

But we have to start by asking ourselves: are we investing in ecosystems conducive to the creation of apex predators? Or in the amalgamation of resources and regulatory capture toward monopoly?

If we are tunnel-visioned into "arms race" memes, ideological labels, or myopic nodemaxxing... these questions get lost in the sauce. And with them, the opportunity to improve ourselves.

> There is a way that seems right to a man, but its end is the way of death. - Proverbs 14:12

## Footnotes

[1] Leaked DeepSeek investor-meeting notes, original Chinese transcript: [https://x.com/i/status/2080183070842638813](https://x.com/i/status/2080183070842638813).

[2] Transformer (2026-07-24) — uses Liang's leaked remarks to argue chip export controls are binding and should be tightened further: [https://www.transformernews.ai/p/deepseek-ceo-liang-wenfeng-export-controls-china](https://www.transformernews.ai/p/deepseek-ceo-liang-wenfeng-export-controls-china)

[3] Parameter estimates for Anthropic's frontier models are unconfirmed third-party inference, flagged as such by the source itself: [https://aithinkerlab.com/claude-opus-5-trillion-parameters/](https://aithinkerlab.com/claude-opus-5-trillion-parameters/). Treat as rumor-tier.

[4] 經濟日報 via udn (2026-07-24), quoting Liang: 16,000 Ascend 950s "只相當於4,000張B系列" — equivalent to only 4,000 B-series. [https://udn.com/news/story/6811/9647545](https://udn.com/news/story/6811/9647545)

[5] Huawei MWC 2026 announcement — the Atlas 950 SuperPoD's 8,192-NPU coherent domain (64 NPUs per cabinet). [https://www.huawei.com/en/news/2026/3/mwc-superpod-ai](https://www.huawei.com/en/news/2026/3/mwc-superpod-ai)

[6] Interconnects, "Latest open artifacts #23" — V4 Flash 0731 "beat Luna at the pareto frontier," released a day after OpenAI's 80% small-model price cut: [https://www.interconnects.ai/p/latest-open-artifacts-23-laguna-s21](https://www.interconnects.ai/p/latest-open-artifacts-23-laguna-s21)

[7] [https://x.com/rydcunningham/status/2080446874344148994](https://x.com/rydcunningham/status/2080446874344148994)

[8] [https://x.com/jordanschneider/status/1580889347846713344](https://x.com/jordanschneider/status/1580889347846713344)

[9] WAIC 2026 observations: [https://x.com/rydcunningham/status/2079998255224828013](https://x.com/rydcunningham/status/2079998255224828013)

[10] Jeffrey Ding, **Technology and the Rise of Great Powers** (Princeton University Press) — the diffusion-over-leading-sector argument this section leans on: [https://press.princeton.edu/books/paperback/9780691260341/technology-and-the-rise-of-great-powers](https://press.princeton.edu/books/paperback/9780691260341/technology-and-the-rise-of-great-powers)

[11] "AI is a Five Layer Cake," Jensen Huang, NVIDIA Blog, 2026-03-10: [https://blogs.nvidia.com/blog/ai-5-layer-cake/](https://blogs.nvidia.com/blog/ai-5-layer-cake/)

[12] [https://x.com/a16z/status/1975290317005070737](https://x.com/a16z/status/1975290317005070737)

[13] "China's Silicon Vanguard," Machine Yearning [https://www.machineyearning.io/p/chinas-silicon-vanguard](https://www.machineyearning.io/p/chinas-silicon-vanguard)

[14] [https://en.wikipedia.org/wiki/Aaron\_Swartz](https://en.wikipedia.org/wiki/Aaron_Swartz)

[15] Tensor Economics: [https://www.tensoreconomics.com/](https://www.tensoreconomics.com/)

[16] Philip Kiely, **Inference Engineering** (Baseten Books, 2026): [https://www.baseten.co/inference-engineering/](https://www.baseten.co/inference-engineering/)

[17] The 2-forward / 4-backward convention and its caveats (notably that 6N excludes the quadratic attention term): [https://jax-ml.github.io/scaling-book/transformers/](https://jax-ml.github.io/scaling-book/transformers/)

[18] "The Smol Training Playbook," HuggingFace — source of the 20–30% MoE MFU reference and the GEMM-level FP8 measurements cited in Appendix A: [https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook](https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook)

[19] NVIDIA marketing materials typically show Blackwell-class compute in "sparse" FLOPS, roughly 2× dense; we convert everything to dense for apples-to-apples. The GB-series are two-GPU "superchips"; GB specs are shown per GPU unless noted.

[20] NVIDIA B300/GB300 figures from the DGX B300 and Blackwell Ultra datasheets: [https://resources.nvidia.com/en-us-dgx-systems/dgx-b300-datasheet](https://resources.nvidia.com/en-us-dgx-systems/dgx-b300-datasheet); [https://resources.nvidia.com/en-us-blackwell-architecture/blackwell-ultra-datasheet](https://resources.nvidia.com/en-us-blackwell-architecture/blackwell-ultra-datasheet)

[21] HiSilicon Ascend 910 product page (archived 2022-12-01): [https://web.archive.org/web/20221201202510/https://www.hisilicon.com/en/products/Ascend/Ascend-910](https://web.archive.org/web/20221201202510/https://www.hisilicon.com/en/products/Ascend/Ascend-910). First-generation part: 256 TFLOPS FP16, 32 GB HBM2, 1,228 GB/s, 300 W.

[22] Jacob Feldgoise and Hanna Dohmen, "Pushing the Limits: Huawei's AI Chip Tests U.S. Export Controls," CSET, 17 June 2024: [https://cset.georgetown.edu/publication/pushing-the-limits-huaweis-ai-chip-tests-u-s-export-controls/](https://cset.georgetown.edu/publication/pushing-the-limits-huaweis-ai-chip-tests-u-s-export-controls/). Appendix Tables 1–2 enumerate the first- and second-generation SKUs; the highest first-generation part "appears to be capable of 320 FP16 TFLOPS" at 1,228 GB/s (four HBM stacks at 307 GB/s each), and the highest 910B "appears to be capable of 400 FP16 TFLOPS" with four 16 GB stacks (64 GB) and a maximum 1,600 GB/s. CSET publishes no TDP figures.

[23] Lennart Heim, "Huawei's Ascend 910C," 12 March 2025: [https://blog.heim.xyz/huawei-ascend-910c/](https://blog.heim.xyz/huawei-ascend-910c/). Explicitly an estimate — "I'd expect the 910C to achieve ~800 TFLOP/s at FP16 and ~3.2 TB/s memory bandwidth" — derived from the dual-die package rather than any Huawei disclosure.

[24] Huawei Connect 2025 Ascend roadmap (2025-09-18): 950PR Q1 2026 (prefill/recommendation, first with in-house HBM), 950DT Q4 2026 (decode/training); the 950 series adds FP8/MXFP8/HiF8 at 1 PFLOPS and MXFP4 at 2 PFLOPS per chip — 8,192 × 1 PF is the SuperPoD's 8.2 EF FP8 aggregate. Via TrendForce: [https://www.trendforce.com/news/2025/09/18/news-huawei-unveils-ascend-950-with-in-house-hbm-in-2026-touts-superpod-to-rival-nvidia/](https://www.trendforce.com/news/2025/09/18/news-huawei-unveils-ascend-950-with-in-house-hbm-in-2026-touts-superpod-to-rival-nvidia/); see also Tom's Hardware ([https://www.tomshardware.com/tech-industry/semiconductors/huawei-unveils-ascend-roadmap-backed-by-in-house-hbm](https://www.tomshardware.com/tech-industry/semiconductors/huawei-unveils-ascend-roadmap-backed-by-in-house-hbm)) and Rui's contemporaneous thread ([https://x.com/ruima/status/1968592824380878931](https://x.com/ruima/status/1968592824380878931)).

[25] Huawei Connect 2025 keynote (Eric Xu) — Ascend 960, Q4 2027: 2 PFLOPS FP8 / 4 PFLOPS FP4 per chip, with double the 950's memory (288 GB), memory bandwidth, and interconnect; Atlas 960 SuperPoD: up to 15,488 chips, 30 EFLOPS FP8, 4,460 TB of memory. [https://www.huawei.com/en/news/2025/9/hc-xu-keynote-speech](https://www.huawei.com/en/news/2025/9/hc-xu-keynote-speech)

[26] MindVL (arXiv 2509.11662), Appendix A.1, Table 10 — the paper's own Ascend-vs-NVIDIA comparison, giving 910B1 and 910B2 across TF32/FP32/BF16/FP16/INT8/HBM/HBM-bandwidth: [https://arxiv.org/abs/2509.11662](https://arxiv.org/abs/2509.11662)

[27] "Towards Efficient Multi-Scale Deformable Attention on NPU" (arXiv 2505.14022), Table 1, "Ascend 910B2C Performance Specifications" — 353 TFLOPS Cube@FP16, 64 GB, 1,800 GB/s: [https://arxiv.org/abs/2505.14022](https://arxiv.org/abs/2505.14022)

[28] xDeepServe (Huawei): "The Ascend NPU in CloudMatrix384 does not natively support FP8 arithmetic"; DeepSeek models are deployed there via INT8 post-training quantization. [https://arxiv.org/html/2508.02520v5](https://arxiv.org/html/2508.02520v5)

[29] Huawei's three-year Ascend roadmap as reported by Huawei Central: 950PR Q1 2026, 950DT Q4 2026, 960 Q4 2027, 970 Q4 2028, with a stated 2.5× interconnect-bandwidth increase for the 950 generation: [https://www.huaweicentral.com/huawei-reveals-3-year-ascend-ai-chip-roadmap-950-coming-in-2026/](https://www.huaweicentral.com/huawei-reveals-3-year-ascend-ai-chip-roadmap-950-coming-in-2026/). The page carries the schedule, not the per-chip FLOPS figures, which come from the Connect keynote and its trade-press coverage.

[30] CFR's critique is framed in TPP (total processing performance, the export-control metric), on which the 950 series regresses vs the 910C: [https://www.cfr.org/articles/chinas-ai-chip-deficit-why-huawei-cant-catch-nvidia-and-us-export-controls-should-remain](https://www.cfr.org/articles/chinas-ai-chip-deficit-why-huawei-cant-catch-nvidia-and-us-export-controls-should-remain)

[31] [https://interconnected.blog/the-real-deepseek-moment-just-arrived/](https://interconnected.blog/the-real-deepseek-moment-just-arrived/) — DeepSeek's UE8M0 FP8 standard, co-designed for next-generation domestic silicon.

[32] CloudMatrix-Infer (Huawei) independently documents INT8 serving of DeepSeek-R1 on CM384 — 6,688 prefill / 1,943 decode tok/s per NPU. [https://arxiv.org/abs/2506.12708](https://arxiv.org/abs/2506.12708)

[33] Epoch AI, longest training runs — observed frontier runs cluster near three calendar months, and the analysis identifies ~9 months as the point past which longer runs become inefficient versus waiting for improved hardware and algorithms: [https://epoch.ai/data-insights/longest-training-run](https://epoch.ai/data-insights/longest-training-run)

[34] Serving is embarrassingly parallel across requests: each replica owns its requests end-to-end, and the only shared infrastructure is the request router and cache tier. There is no gradient to exchange and no step barrier — so replicas can sit in different datacenters, fail independently, and scale linearly, none of which a training cluster can do.

[35] SemiAnalysis, "Huawei AI CloudMatrix 384 — China's Answer to NVIDIA GB200 NVL72": [https://newsletter.semianalysis.com/p/huawei-ai-cloudmatrix-384-chinas-answer-to-nvidia-gb200-nvl72](https://newsletter.semianalysis.com/p/huawei-ai-cloudmatrix-384-chinas-answer-to-nvidia-gb200-nvl72)

[36] 上海证券报 (cnstock), 2026-07-17 — Huawei discloses 逾750套 (over 750) Ascend supernode commercial deployments: [https://www.cnstock.com/commonDetail/746743](https://www.cnstock.com/commonDetail/746743)

[37] Sustained vs decode-only throughput, and the 1.4–2.2× calibration band, are worked through in Appendix B.

[38] Simultaneous users solved for each model × system. We sweep the batch ladder and keep the largest batch that still 1. Meets SLA, and 2. Fits memory. Concurrent streams = system tokens/s ÷ per-user tokens/s. See Appendix B.

[39] Measured token mix from OpenRouter's provider-weighted data: DeepSeek V4 Flash runs 15.95 input tokens per billed output token, V4 Pro 26.8, R1 9.2. Reasoning tokens bill as output. See Appendix C.

[40] OpenRouter, "DeepSeek V4 adoption": ~20% of platform tokens by June 2026, top model since mid-May, from a 450T-token H1 sample. [https://openrouter.ai/blog/insights/deepseek-v4-adoption](https://openrouter.ai/blog/insights/deepseek-v4-adoption)

[41] OpenRouter (X), daily-token reporting for DeepSeek's top three models: [https://x.com/OpenRouter/status/2074150211300073578](https://x.com/OpenRouter/status/2074150211300073578). Login-walled; archived screenshot on file.

[42] PYMNTS, sourcing The Information: DeepSeek annualized run-rate approaching $400–500M, roughly double 2025. [https://www.pymnts.com/news/artificial-intelligence/2026/deepseek-revenue-nears-500-million-as-chinese-ai-startup-eyes-ipo](https://www.pymnts.com/news/artificial-intelligence/2026/deepseek-revenue-nears-500-million-as-chinese-ai-startup-eyes-ipo)

[43] DeepSeek published API pricing (per Mtok: V4 Flash $0.28 out / $0.14 in-miss; V4 Pro $0.87 out / $0.435 in-miss): [https://api-docs.deepseek.com/quick\_start/pricing/](https://api-docs.deepseek.com/quick_start/pricing/)

[44] Similarweb, chat.deepseek.com, June 2026: 335.4M visits over three months, 00:05:13 average visit, 3.00 pages per visit. Web only — excludes the mobile apps (57M+ downloads) and all API traffic. [https://www.similarweb.com/website/chat.deepseek.com/](https://www.similarweb.com/website/chat.deepseek.com/)

[45] InferenceX by SemiAnalysis, single-node DeepSeek V4 Pro dataset, fetched 2026-07-31. [https://inferencex.semianalysis.com/](https://inferencex.semianalysis.com/)

[46] SLAI T-Rex technical report — full-parameter post-training of DeepSeek V4 Pro on an Ascend CloudMatrix384 SuperPoD, 34.22% MFU, 2.93× over baseline: [https://arxiv.org/pdf/2607.20145](https://arxiv.org/pdf/2607.20145)

[47] Image via [@jenzhuscott](https://x.com/jenzhuscott): [https://x.com/jenzhuscott/status/2083516193743524342](https://x.com/jenzhuscott/status/2083516193743524342)

[48] LongCat 2.0 technical report: [https://longcat.chat/blog/longcat-2.0/](https://longcat.chat/blog/longcat-2.0/)

[49] LongCat 2.0 model card — 1.6T total, ~48B/token, LSA, 135B n-gram embedding, 3-step MTP, >35T tokens, "millions of accelerator-days… no rollbacks or irrecoverable loss spikes": [https://huggingface.co/meituan-longcat/LongCat-2.0](https://huggingface.co/meituan-longcat/LongCat-2.0)

[50] SCMP: first trillion-parameter model trained end-to-end on a ~50,000-card domestic cluster: [https://www.scmp.com/tech/tech-trends/article/3358854/china-debuts-biggest-ai-model-trained-local-chips-meituan-releases-longcat-20](https://www.scmp.com/tech/tech-trends/article/3358854/china-debuts-biggest-ai-model-trained-local-chips-meituan-releases-longcat-20)

[51] Artificial Analysis, LongCat 2.0: [https://artificialanalysis.ai/models/longcat-2-0](https://artificialanalysis.ai/models/longcat-2-0). The Flash-Lite/Nemotron comparison is our read of the surrounding leaderboard, not AA's phrasing.

[52] The standard 6N convention at the vendor-published activation count: 6 × 48B = 2.88 × 10¹¹ FLOPs/token, × 35T tokens ≈ 1.0 × 10²⁵. Recompute and attention overheads live in the MFU term, matching how the reference MFU figures are denominated. See Appendix A.

[53] The contrast matters for policy and safety researchers: counted by activated parameters, LongCat sits 10× under the 2023 Executive Order's 10²⁶-FLOP reporting threshold; counted by total parameters, 3.4× over it. When that threshold was set, dense transformers were the default frontier architecture. Mixtral 8x7B arrived weeks after the EO, and it was not until mid-2025 that MoEs were first measured surpassing dense transformers under strictly equal total parameters, compute, and data ([https://arxiv.org/pdf/2506.12119](https://arxiv.org/pdf/2506.12119); on threshold design see also [https://jolt.law.harvard.edu/digest/beyond-flops-shortcomings-of-flops-as-a-model-classification-metric-in-ai-regulation-1](https://jolt.law.harvard.edu/digest/beyond-flops-shortcomings-of-flops-as-a-model-classification-metric-in-ai-regulation-1)).

[54] Ascend 910B: 400 TF FP16/BF16 (910B1 bin), 64 GB HBM2e @ 1.6 TB/s, 310 W. Compute and memory per [https://flopper.io/gpu/huawei-ascend-910b](https://flopper.io/gpu/huawei-ascend-910b); the 310 W board power is Huawei's own archived specification[21], which supersedes the 400 W figure that aggregator sites carry. Note the fleet-energy figures in this piece do not use board power — they use the Atlas 800T A2's system power divided by its eight cards (0.725 kW/card), which includes CPUs, fans and NICs.

[55] Huawei booth display, "Scaled Trillion Parameter Inference / 万亿参数模型规模化推理部署," photographed July 2026 (author's photo). Technically an inference showcase, not a training disclosure — it confirms the two companies work closely on Ascend deployment; it doesn't by itself corroborate the training claim. And while plenty of other domestic accelerators could host a run this size (Alibaba's clusters among them), the display makes Ascend the obvious reading.

[56] Vendor claim circulated via X, login-walled: [https://x.com/Michaelzsguo/status/2071975363958260122](https://x.com/Michaelzsguo/status/2071975363958260122). No quantified baseline given.

[57] DeepSeek-V3 Technical Report — 2.664M H800-hours pretraining over 14.8T tokens; FP8 framework ("theoretically doubles the computational speed"; no end-to-end multiplier published): [https://arxiv.org/abs/2412.19437](https://arxiv.org/abs/2412.19437)

[58] Pangu Ultra MoE — 30.0% MFU on 6K Ascend NPUs, 718B model: [https://arxiv.org/pdf/2505.04519](https://arxiv.org/pdf/2505.04519). The Ascend goodput band is our assumption and is weakly sourced; see Appendix A.

[59] Pangu Ultra (135B dense, 8,192 NPUs): baseline ~43% MFU raised past 52% via kernel fusion, subsequence context parallelism, and caching: [https://arxiv.org/pdf/2504.07866](https://arxiv.org/pdf/2504.07866)

[60] The Llama 3 Herd of Models — 38–43% BF16 MFU, >90% effective training time, 419 unexpected interruptions over the 54-day 405B run, root-cause table: [https://arxiv.org/abs/2407.21783](https://arxiv.org/abs/2407.21783)

[61] Image via [@1owroller](https://x.com/1owroller): [https://x.com/1owroller/status/2077975698959208938](https://x.com/1owroller/status/2077975698959208938)

[62] Kratsios remarks of 2026-07-22: [https://x.com/mkratsios47/status/2079933645888880708](https://x.com/mkratsios47/status/2079933645888880708) (login-walled); via Reuters syndication: [https://m.investing.com/news/stock-market-news/moonshot-has-nvidia-chip-cluster-from-alibaba-computing-deal-bloomberg-news-reports-4828879](https://m.investing.com/news/stock-market-news/moonshot-has-nvidia-chip-cluster-from-alibaba-computing-deal-bloomberg-news-reports-4828879)

[63] Bloomberg, 2026-07-31 — Moonshot's Kimi runs on a ~20,000-chip NVIDIA (Hopper-generation) cluster via an Alibaba computing agreement, with Blackwell reportedly reachable through Southeast Asia: [https://www.bloomberg.com/news/articles/2026-07-31/moonshot-s-kimi-built-on-20-000-nvidia-chip-cluster-from-alibaba](https://www.bloomberg.com/news/articles/2026-07-31/moonshot-s-kimi-built-on-20-000-nvidia-chip-cluster-from-alibaba)

[64] Kimi K3 model card — 2.8T total / 104B activated, 16 of 896 experts, "69 KDA + 24 Gated MLA" across 93 layers, 1,048,576-token context; training tokens not disclosed (we assume 35T, matching LongCat's corpus). FLOPs use the 6N convention on the published activation count; we do not model Kimi Delta Attention directly, so treat the figure as an estimate. See Appendix A. [https://huggingface.co/moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3)

[65] EO 14110's reporting trigger (>10²⁶ operations): [https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence](https://www.federalregister.gov/documents/2023/11/01/2023-24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence)

[66] Epoch AI, notable AI models — "frontier" defined as top ten by training compute at release, "a threshold that grows over time"; the static 10²³ line applies only to the regulatory "large-scale" category: [https://epoch.ai/data/notable-ai-models](https://epoch.ai/data/notable-ai-models)

[67] Financial Times, August 2025 — R2 delayed after months of failed Ascend training attempts, Huawei engineers on-site, DeepSeek reverting to NVIDIA for the run: [https://www.ft.com/content/eb984646-6320-4bfe-a78d-a1da2274b092](https://www.ft.com/content/eb984646-6320-4bfe-a78d-a1da2274b092)

[68] Image via [@chamath](https://x.com/chamath): [https://x.com/chamath/status/2084242435215913234](https://x.com/chamath/status/2084242435215913234)

[69] H800 SXM5 per TechPowerUp: [https://www.techpowerup.com/gpu-specs/h800-sxm5.c3975](https://www.techpowerup.com/gpu-specs/h800-sxm5.c3975) — same GH100 die and 132 SM / 528 tensor-core config as H100 SXM5, clock-reduced to a 1755 MHz boost (vs 1980), confirmed by its FP32 (59.30 vs 66.91) and FP64 (29.65 vs 33.45) figures. Tensor rates derive as H100 SXM5 × 0.886 → 877 TF dense BF16. TPU's listed "FP16 (half) 237.2 TFLOPS" is the vector rate, not tensor.

[70] MegaScale — 55.2% MFU training a 175B dense model on 12,288 GPUs: [https://arxiv.org/abs/2402.15627](https://arxiv.org/abs/2402.15627)

[71] OpenRouter model pages, provider-weighted averages and token mixes (e.g. [https://openrouter.ai/deepseek/deepseek-r1-0528](https://openrouter.ai/deepseek/deepseek-r1-0528)). Pinned snapshot 2026-07-29, prices live-verified 2026-08-03; a weekly refresh daemon is planned.

[72] [https://compute.exchange/blogs/h100-gpu-price-2026](https://compute.exchange/blogs/h100-gpu-price-2026)

[73] [https://intuitionlabs.ai/articles/nvidia-ai-gpu-pricing-guide](https://intuitionlabs.ai/articles/nvidia-ai-gpu-pricing-guide)

[74] [https://viperatech.com/product/nvidia-dgx-h800-640gb-sxm5-2tb](https://viperatech.com/product/nvidia-dgx-h800-640gb-sxm5-2tb)

[75] Institute for Progress, "The H20 Problem": [https://ifp.org/the-h20-problem/](https://ifp.org/the-h20-problem/)

[76] [https://www.spheron.network/blog/gb300-nvl72-vs-gb200-nvl72-pricing-availability-2026/](https://www.spheron.network/blog/gb300-nvl72-vs-gb200-nvl72-pricing-availability-2026/); [https://tech-insider.org/nvidia-blackwell-gpu-pricing/](https://tech-insider.org/nvidia-blackwell-gpu-pricing/); [https://www.spheron.network/gpu-rental/b300/](https://www.spheron.network/gpu-rental/b300/)

[77] Zhihu market reporting, ¥120K/card (≈$16.5K): [https://www.zhihu.com/en/answer/50174999323](https://www.zhihu.com/en/answer/50174999323)

[78] China Research Collective — CloudMatrix 384 at ¥127–162M; note the figure is scoped as hardware + infrastructure + software + 3-year service: [https://chinaresearchcollective.substack.com/p/huawei-ascend-cloudmatrix-384-supernode](https://chinaresearchcollective.substack.com/p/huawei-ascend-cloudmatrix-384-supernode)

[79] Reuters, 2026-03-27 — Ascend 950 pricing and ByteDance/Alibaba order interest: [https://www.reuters.com/world/china/huaweis-new-ai-chip-find-favour-with-bytedance-alibaba-which-plan-place-orders-2026-03-27/](https://www.reuters.com/world/china/huaweis-new-ai-chip-find-favour-with-bytedance-alibaba-which-plan-place-orders-2026-03-27/)

[80] Personal SRE field report (the widely-circulated "Anyscale 10–30%" figure is repeated here secondhand; the author's own clusters averaged 8%): [https://rajyadavsredev.medium.com/about-a-year-ago-we-ran-gpu-utilization-reports-across-our-clusters-and-came-up-with-an-average-of-a743a708aab9](https://rajyadavsredev.medium.com/about-a-year-ago-we-ran-gpu-utilization-reports-across-our-clusters-and-came-up-with-an-average-of-a743a708aab9)

[81] [https://www.spheron.network/blog/ai-inference-power-electricity-cost-2026/](https://www.spheron.network/blog/ai-inference-power-electricity-cost-2026/)

## Appendix

### A. Training Arithmetic Methodology

We refer to the number of days spent on a training run as wall-clock days. These are derived from the following:

$$
\text{wall-clock days} = (\text{FPT} \space × \space T) \space / \space (C \space × \space TF_{dense}(p) \space × \space MFU_p \space × \space g \space × \space 86,400)
$$

Where:

- FPT is the model's training FLOPs per token: the standard transformer approximation of **6 × N\_active** — 2 FLOPs per active parameter on the forward pass, 4 on the backward[17]. We apply it exactly, with no architecture-specific inflation. This is a deliberate convention, not laziness: the MFU figures we calibrate against are themselves 6N-denominated (Llama-3 405B's "3.8×10²⁵ FLOPs" verifies as 6 × 405B × 15.6T; DeepSeek-V3's delivered rate verifies as 6 × 37B × 14.8T ÷ 2.664M H800-hours = ~343 TF/GPU, 39% of the 877 TF BF16 peak), so activation recompute, the quadratic attention term, and every other flavor of re-done or auxiliary work is **already inside** the published efficiency ratios. Put it in the numerator too and you charge it twice — an error worth +18–30% on wall-clock that an earlier vintage of this model made, caught in peer review against the Smol Playbook's sizing convention. At a 4,096-token training sequence, true attention FLOPs add only ~2% for these architectures — well inside the band width.
- T is the total number of pretraining tokens to process.
- N\_{active} is activated parameters per token, not total. In a mixture-of-experts model each token routes through a handful of experts, so only a fraction of the weights do arithmetic on any given token — activated parameters set the FLOPs bill, while total parameters set the memory bill. LongCat 2.0 activates ~48B of 1.6T total (97% of the model sits idle per token); a dense model like Llama 3.3 70B activates everything.
- TF\_dense(p) is the accelerator's dense throughput at the precision p the run trains in — the datapath the arithmetic actually executes on, never a sparsity figure. MFU\_p must then be denominated against that same peak, and the two accountings are algebraically interchangeable: BF16-basis-times-uplift (how our solver stores it, because every published calibration MFU is BF16-denominated) and dense-peak-at-p times MFU\_p (how this appendix states it) produce identical wall-clocks by construction. The realization question is where judgment enters. DeepSeek-V3 demonstrated FP8 pretraining at scale but publishes no end-to-end speedup figure[57]; for a **dense** FP8 run on NVIDIA we credit 65% realization of the datapath's nameplate 2× (equivalent to a 1.3× uplift on a BF16 basis) — consistent with GEMM-level measurements of 1.6–1.9× (Smol Playbook) shaved for the non-GEMM share of step time. FP8 **MoE** runs get the straight peak conversion with no further credit, because the NVIDIA-MoE band was measured on DeepSeek-V3's own FP8 run — its dividend is already inside the band. Ascend parts get no realization credit anywhere: the 910B/910C have no FP8 datapath at all[28], and the 950 series has no published at-scale FP8 training run to anchor on — when one lands, this framework prices it in one constant.
- MFU is the fraction of TF\_{eff} sustained during training, excluding downtime — and it is architecture-conditional, because MoE all-to-all traffic taxes utilization. Dense: Llama-3 405B reported 38–43% BF16 MFU[60]. DeepSeek-V3 cross-checks the arithmetic: the report logs 2.664M H800-hours of pretraining over 14.8T tokens, or ~343 effective TFLOPS per GPU; against the H800 SXM5's 877 TF dense BF16 peak[69] that is ~39% combined utilization, implying ~43% MFU at 90% goodput — the top of Llama-3's reported band, from an entirely independent run[57]. MegaScale's 55.2% (dense, 12K GPUs) is the known ceiling[70]. MoE: our base 30% sits deliberately below DeepSeek-V3's delivered ~39% — V3 is the efficiency exemplar, not the median — and inside the Smol Playbook's 20–30% reference for communication-bound MoE training. (An earlier draft quoted V3 at "34.6%"; that figure divided by the H100's 989 TF peak rather than the H800's 877 — a wrong-denominator artifact this convention exists to prevent.) For Ascend we anchor on Pangu Ultra MoE's reported ~30% MFU on ~6K NPUs — a vendor figure on a tuned in-house run, so it serves as our base case rather than optimistic-case evidence[58]. Bands used (pessimistic / base / optimistic): NVIDIA dense 25 / 38 / 45%; NVIDIA MoE 18 / 30 / 38%; Ascend 15 / 30 / 35% for both.
- g is our "goodput," the fraction of wall-clock time making forward progress after unexpected interruptions, failures, restarts, and checkpointing.Llama-3 405B reported ~90% effective training time across 419 unexpected interruptions in 54 days[60]. Bands used: NVIDIA 75 / 90 / 95%. We model a 40–75% goodput band for Ascend chips. This is an assumption, weakly sourced — it reflects DeepSeek V4 / R2 delay reporting concerning difficulty with training runs on Huawei chips[67] and general ecosystem maturity, and it is the single loosest input in the training math. Move it and the wall-clock moves with it.

Worked example, LongCat 2.0:

$$
FPT = 6 × 48×10⁹ active params = 2.88×10¹¹ FLOPs/token
$$

$$
Total = 2.88×10¹¹ FLOPs/token × 35×10¹² tokens ≈ 1.0×10²⁵ FLOPs
$$

At the Ascend base case (30% MFU × 60% goodput on a 400 TF BF16 910B), that solves to ≈1.6M accelerator-days - ~32 days of wall-clock across 50,000 chips - before the pessimistic/optimistic bands are applied. An earlier vintage of this model carried recompute in the numerator (+19% over 6N) and the superseded 41.6B activation count; the two corrections largely offset, so the headline FLOPs barely move — but the derivation is now consistent with how the reference MFUs are denominated, which the old one was not.

Knowing wall-clock days and the accelerator specs, we can proxy training run costs in dollars, gigawatt-hours, and/or accelerator-days.

Overall per-system costs are proxied via sensitized ranges of procurement prices for outright hardware (all-in costs where available; see Appendix C for the full table), a 10% opex-per-year ratio applied over a 4-year asset life with straight-line depreciation, and energy at system-power-per-card × chip-time - node overhead (CPUs, fans, NICs) included, facility-level PUE excluded.

### B. Inference Arithmetic

Our inference model is a first-principles decode roofline, solved per model × system pairing. Bottoms-up:

**One decode step.** To emit one token per request, the serving group must read the model's resident weights plus every concurrent request's KV cache out of HBM, perform ~2 FLOPs per active parameter of matrix math, and exchange expert-routing traffic over the scale-up fabric. Step time is whichever is slowest, plus a serialization floor:

$$
TPOT = max(t_compute, t_memory, t_interconnect) + t_serial
$$

- t\_compute = b × 2N\_active ÷ group compute peak, at the model's native precision where the chip supports it (INT8 datapath otherwise — DeepSeek's FP8 models quantize to INT8 on 910C-class parts).
- t\_memory = (replicated weights ÷ min(G, 64) + expert shard × f\_touch + b × KV\_bytes × ctx) ÷ HBM bandwidth. The expert-touch fraction f\_touch = 1 − (1 − k/n)^b: a batch of b tokens reads only the experts it routes to, saturating toward all n as batch grows — this is what bends per-user speed down with batch, and at production batch it recovers the reference texts' "expect almost all parameters active" rule.
- t\_interconnect = per-layer collective volume ÷ effective fabric bandwidth: dense tensor-parallel pays two all-reduces per layer (ring cost 4(G−1)·b·d\_model·bytes·L); MoE expert-parallel pays dispatch+combine of each token's hidden state to its top-k experts per layer (2·b·k·d\_model·bytes·L). When a group spans nodes, effective bandwidth blends scale-up and scale-out links by the fraction of pair-crossings that leave the domain.
- t\_serial is an additive per-step floor — 8.1 ms (node tier) / 1.8 ms (rack) / 6.8 ms (superpod), fitted against our solver's reference cubes — covering what a pure roofline omits: kernel launches serialized across 38–80 layers, collective latency, attention math. It is additive rather than multiplicative because these costs do not scale with the expert read; at low batch they dominate it.

Per-user speed u = 1/TPOT; per-card decode-only throughput = b × u.

**KV caching.** We normalize each model's HuggingFace config into its attention primitive and compute KV bytes per token directly: grouped-query attention stores 2 × g × L × d\_head values (Llama 3.3 70B: ~328 KB/token at BF16); DeepSeek's latent attention stores a shared compressed latent, 4.5 × L × d\_head with no factor of 2 (R1: ~70 KB/token at the BF16 reference cache); V3.2's sparse attention adds a lightning-indexer term on top of MLA (~78 KB/token); the V4 family's compressed-sparse hybrid amortizes to ~5 KB/token; LongCat sits GQA-like at ~49 KB/token. This single number drives most of the long-context differences in the article.

**Sharding and group sizing.** Routed experts shard across the serving group; attention, shared experts, and embeddings replicate per expert-parallel rank but shard under tensor- and pipeline-parallelism within reach (we cap that reach at 64 = TP8 × PP8). The minimum group that fits is a capacity answer — we report it as minimum deployment — but it is not the group anyone runs: we sweep feasible group sizes and solve at the throughput optimum, so a system with more memory per card shows its advantage as KV headroom rather than being penalized with a larger weight read. Remaining cards become independent replicas.

**Prefill and sustained throughput.** Prefill is charged explicitly at a 30% prefill MFU (the solver-implied node-tier rate; colocated production stacks report 35–55%, so this is conservative). Time-to-first-token is per request — one request's uncached prompt FLOPs across the whole group's compute — at the 50% cache-hit basis measured for production traffic, rescaled linearly by the cache slider. Sustained throughput then charges the group for prefilling everything it serves: over a cycle that prefills G×b requests and decodes 256 tokens each,

$$
sustained tok/s per card = b × 256 ÷ (G × b × TTFT + 256/u)
$$

Decode-only and sustained are reported separately everywhere; fleet capacity, economics, and every headline figure in this article use sustained.

**Concurrency.** For each pairing we sweep a batch ladder (1 → 8,192 per accelerator in ~1.5× steps) against a context grid (4K → 1M, clamped to each model's maximum) and select the largest batch that still fits memory and meets the tok/s/user SLA — the SLA-binding point, so every system is held to the same user-experience floor and throughput is the output. Concurrent streams = system tokens/s ÷ per-user tokens/s.

**Calibration.** The model reproduces Tensor Economics' published Llama serving arithmetic to within 2–5%, and runs 1.4–2.2× above measured production deployments (Aleph Alpha's EP72 measurements; Huawei's own CloudMatrix384 serving numbers). We therefore treat all absolute throughputs as a first-principles ceiling: ratios between chips and models are sound, and absolutes should be divided by 1.4–2.2× to approximate delivered systems. As an independent check against measurements that played no part in setting that band, we compared 44 SemiAnalysis InferenceX single-node runs (DeepSeek V4 Pro, FP4, on 8× B300 and 8× B200 across SGLang / TensorRT-LLM / vLLM at concurrencies 8 → 8,192): our throughput estimates run a median 1.8× above measured, p10–p90 1.4–2.2×, flat across the concurrency range — the same band, from an independent source. The full point-by-point comparison ships in Overclock's EMPIRICS view.
