Resources / References / 01
Four to One: Auditing the Leaked DeepSeek Notes
E² review · September 25, 2026
Separate the claim, its evidence and the check performed.
Claims ledger
Each record separates the claim’s type, its support and the check actually performed. “Reported fact” describes an attributable statement; it does not certify the underlying event. Read the rubric.
Symbol key
Symbols supplement the words; they do not replace them.
Reported source Assumption Calculation Measurement
Source checked Arithmetic reproduced Conflict Check still open
Support uses one to five filled marks, from conjectural to established within scope. This is an ordinal evidence scale, not a probability. A dash means the scale does not apply.
- F01
DeepSeek received about 16,000 Ascend 950 chips.
- Type
- Alleged leak
- Support
- Reported
- Review
- Not independently verified
The article attributes this allocation to reporting about leaked remarks. Neither the original meeting record nor delivery records were authenticated in this review. Repetition of the report is not independent corroboration.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: An authenticated primary account or delivery record identifying date, quantity and chip variant.
- F02
The allocation is 950DT hardware in two Atlas 950 SuperPoDs.
- Type
- Model assumption
- Support
- Not applicable
- Review
- Source checked
The author marks the topology as a guess. Huawei’s roadmap schedules 950DT and Atlas 950 for Q4 2026, after the article’s August publication. An early allocation is possible; advertised maximum size does not establish the delivered SKU or topology.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
- Huawei Connect 2025 keynote — Ascend 950 / 960 roadmap
- Huawei MWC 2026 announcement — Atlas 950 SuperPoD specifications
What would change this assessment: A deployment disclosure specifying SKU, system layout and operational date.
- F03
Huawei advertised an Atlas 950 system with up to 8,192 NPUs.
- Type
- Reported fact
- Support
- Supported
- Review
- Source checked
The vendor announcement supports this narrow statement about advertised scale. It does not measure workload throughput or verify DeepSeek’s allocation.
Evidence and revision conditions
- Huawei MWC 2026 announcement — Atlas 950 SuperPoD specifications
What would change this assessment: A revised vendor specification. Actual performance needs separate workload measurements.
- F04
The LongCat model card reports about 48B active parameters and more than 35T training tokens.
- Type
- Reported fact
- Support
- Supported
- Review
- Source checked
The first-party card supports the reported architecture and corpus scale. The 35T scenario uses a rounded lower input, not the exact disclosed training total. The underlying training logs were not independently measured.
Evidence and revision conditions
- LongCat 2.0 model card — Introduction; card retrieved September 25, 2026
What would change this assessment: Versioned architecture details or logs that revise the active count or exact token total.
- F05
At 48B active parameters and 35T tokens, the 6N training estimate is 1.008 × 10²⁵ FLOPs.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Arithmetic reproduced
6 × 48 × 10⁹ × 35 × 10¹² = 1.008 × 10²⁵. This is a conditional estimate of model FLOPs, not measured device work. Efficiency must use the same FLOP convention to avoid double-counting overhead.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
- How to Scale Your Model — Transformer training FLOP approximation
Depends on F04.
What would change this assessment: Changed inputs, or an explicit architecture-specific accounting with a consistently defined efficiency denominator.
- F06
The printed LongCat base scenario takes 32.41 days on 50,000 chips.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Arithmetic reproduced
At 400 trillion FLOP/s per chip, 30% MFU and 60% goodput, the result is 1.620 million accelerator-days. This does not identify the historical chip SKU or actual duration. A numerical TFLOP/s input requires a factor of 10¹² in the denominator.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: A different dense peak, efficiency denominator, training volume or measured run duration.
- F07
The LongCat caption’s 22–98-day range follows from fixed 30% MFU and 40–75% goodput.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Unresolved conflict
Those fixed-MFU inputs yield 25.93–48.61 days. The wider 22.22–97.22-day range requires varying MFU from 15% to 35% as well. The caption and Appendix A describe different sweeps; neither range is a confidence interval.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
Depends on F05.
What would change this assessment: Correct the caption to name both varying inputs, or replace its range with the fixed-MFU result.
- F08
Ascend goodput is 40–75% across the modeled training scenarios.
- Type
- Model assumption
- Support
- Not applicable
- Review
- Source checked
The author explicitly calls this weakly sourced. It is a sensitivity range, not a measured fleet statistic. Holding it and MFU constant when scaling from thousands to hundreds of thousands of chips is a further assumption.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: Measured productive wall-clock fractions by fleet size, topology, workload and restart policy.
- F09
Pangu Ultra MoE’s authors report 30.0% MFU on 6,000 Ascend NPUs.
- Type
- Reported measurement
- Support
- Supported
- Review
- Source checked
The abstract supports this first-party measurement report. It anchors one workload and scale; it does not independently validate the article’s broader Ascend fleet scenarios.
Evidence and revision conditions
- Pangu Ultra MoE report — Abstract: 30.0% MFU on 6K Ascend NPUs
What would change this assessment: Independent reproduction, or detailed counterevidence about the workload and MFU denominator.
- F10
Kimi K3’s model card reports 104B active parameters.
- Type
- Reported fact
- Support
- Supported
- Review
- Source checked
The model summary reports this architecture figure. Combined with an assumed 35T tokens, the 6N calculation yields 2.184 × 10²⁵ FLOPs. The assumed corpus does not become a reported training fact.
Evidence and revision conditions
- Kimi K3 model card — Model summary; card retrieved September 25, 2026
What would change this assessment: A revised model card or a disclosed training-token total.
- F11
Kimi K3 was most likely trained on foreign silicon.
- Type
- Inference
- Support
- Reported
- Review
- Not independently verified
The article cites reporting about a Hopper cluster. A plausible 47-day modeled schedule alone cannot establish which machines trained the model. The cited historical reporting was not authenticated in this audit.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
Depends on F10.
What would change this assessment: Primary training provenance, allocation records or a technical report naming the fleet.
- F12
50,000 GB300s take one quarter of the cited 99-day Ascend run.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Unresolved conflict
The body gives 53 days, so its time ratio is 53/99 = 0.535. One quarter would be 24.75 days. Confirm model version, precision, efficiency and GPU-versus-superchip units before choosing a replacement headline.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: A consistent configuration export that reconciles the summary and body.
- F13
The 33T Flash / 8.6T Pro headline denotes generated tokens per day.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Contradicted as labeled
At the printed 16:1 input:output mix, 1.9T Flash output implies 32.3T total; 0.5T Pro output implies 8.5T total. The larger numbers include input tokens. “Generated” is the wrong label under those inputs.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: Relabel total tokens and keep generated/output tokens separate throughout.
- F14
The Pro body’s roughly 14T total is the same 16:1 scenario as the opening.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Unresolved conflict
At 0.5T output, 14T total instead matches the 26.8:1 Pro mix in footnote 39: 0.5 × 27.8 = 13.9. Changing input mix also changes prefill work; multiplying a fixed output projection is only a reconciliation of labels, not a rerun.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
Depends on F13.
What would change this assessment: Name the workload mix for each result and recompute sustained throughput at that mix.
- F15
Overclock reproduces Tensor Economics calculations within 2–5%.
- Type
- Calculated comparison
- Support
- Reported
- Review
- Not reproduced
This is author-reported implementation agreement. Even if reproduced, agreement with another model is not independent measurement of deployment performance.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: Publish reference cases, inputs, expected values and executable comparison code.
- F16
Across 44 InferenceX runs, Overclock overpredicts throughput by median 1.81×, with p10–p90 of 1.39–2.18×.
- Type
- Comparison with independent measurements
- Support
- Reported
- Review
- Not reproduced
The author says these measurements were held out. The public historical query is retained with this audit, but the exact 44-row selection and paired Overclock predictions are unavailable here. We cannot recompute the median, quantiles or flatness by concurrency from aggregate statements.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
- InferenceX API documentation — Historical inference view, metric and configuration definitions
- Official InferenceX repository — Benchmark implementation and reproducibility materials
What would change this assessment: Publish result IDs, paired predictions, metric units, versions, exclusions and the tuning/holdout split.
- F17
The stated fleet capacities meet 20 output tokens per second per user.
- Type
- Calculated output
- Support
- Not applicable
- Review
- Not validated
The model selects its batch using predicted interactivity. The article also reports a 2.8× median overestimate of that metric. Dividing a projected 20 by 2.8 gives 7.14 as an illustrative sensitivity, not a pointwise correction. Aggregate throughput derating cannot validate the batch-selection constraint.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: Recalibrate latency, rerun the batch sweep, and validate the selected configurations against measured service targets.
- F18
Overclock’s per-user interactivity is 2.8× high at the median.
- Type
- Comparison with independent measurements
- Support
- Reported
- Review
- Not reproduced
Disclosing this residual is a strength. The figure still requires paired data and an explicit latency statistic. Its effect on the selected operating point matters more than whether the aggregate-throughput error looks stable.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
Depends on F16.
What would change this assessment: Publish the per-row interactivity residuals and evaluate held-out configurations after recalibration.
- F19
Cross-chip and cross-model ratios remain sound despite absolute throughput errors.
- Type
- Inference
- Support
- Conjectural
- Review
- Not established
Stable bias on B200/B300 does not establish matched bias on Ascend superpods or other models. If two biases independently occupy 1.4–2.2, a predicted ratio can be 0.64–1.57 times the true ratio. This is a sensitivity bound, not an empirical interval.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
Depends on F16.
What would change this assessment: Matched measurements across the compared systems, workload mixes and precisions, with ratio residuals reported.
- F20
The 16,000-chip fleet comfortably covers total DeepSeek demand.
- Type
- Inference
- Support
- Conjectural
- Review
- Not established
Flash capacity is compared with mixed-model demand. Router traffic is a subset; revenue-based estimates omit free usage. Price, cache, workload mix and peak load also matter. The printed figures do not establish enough capacity for total demand at the service target.
- F21
Nine months is a ceiling beyond which a training run is not viable.
- Type
- Inference
- Support
- Reported
- Review
- Unresolved interpretation
Epoch estimates an economic optimum near 8.6 months with a 6.1–14.2-month interval under its assumptions. It is not a physical ceiling. Restricted access to future hardware can change the opportunity cost of waiting.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
- Epoch: longest training runs — Estimated economically optimal training duration
What would change this assessment: An economic calculation using the actual lab’s hardware access, algorithmic progress and release incentives.
- F22
A predominantly domestic frontier-performance run is plausible within twelve months of publication.
- Type
- Forecast
- Support
- Conjectural
- Review
- Not yet resolvable
The implied deadline is August 6, 2027. The article gives no fixed benchmark suite, comparator set, domestic-content threshold or qualifying training stage. Without these, later events can be fitted to the forecast after the fact.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: Preregister resolution criteria and a probability, then score against the August 6, 2027 outcome.
- F23
Performance per dollar matters most when defining the frontier.
- Type
- Value judgment
- Support
- Not applicable
- Review
- Scope identified
This selects an objective. It can be appropriate for a user or operator, but it is not a universal empirical result. Capability, latency, reliability and affordability can lead different users to different choices.
Evidence and revision conditions
- Original article — Body and Appendices A–B; retrieved September 25, 2026
What would change this assessment: State the decision-maker and their objective; compare decisions across plausible priorities.
Review history: September 25, 2026 — first published ledger and arithmetic audit. Original article unchanged; corrections remain open.