Resources / References / 01

Four to One: Auditing the Leaked DeepSeek Notes

Ryan Cunningham · Published August 6, 2026

E² review · 2026-09-25

Review

Contents

Summary

Claims reviewed
23
Weighted support index
1.08 / 4
Arithmetic checked
6 / 10
Empirical comparisons replicated
0 / 2

Review data · Fallacy review data

Support Distribution

Not applicable · excluded
10
1 · Reported
6
2 · Supported
4
0 · Conjectural
3
3 · Corroborated
0
4 · Established within scope
0

Equal weights · 13 scored claims · 14 / 13 = 1.08

The 0–4 encoding is an editorial convention. Gaps between categories are not measured distances; changing the encoding changes the mean.

Claims Provenance

Article attribution 9

The claim is traceable to the article. Its exact proposition is not independently established here.

C01 · C11 · C15 · C16 · C18 · C19 · C20 · C21 · C22

Conditional calculation 7

Inputs and arithmetic are identified in the article. Availability and outcome of reproduction are counted separately below.

C05 · C06 · C07 · C12 · C13 · C14 · C17

Declared value judgment 1

An objective or preference, outside the empirical support ladder.

C23

Replicability

Not validated 1

The 44 paired calibration runs remain unavailable. Read the audit.

Claim Types

Calculated output
7 / 23
Inference
4 / 23
Reported fact
3 / 23
Model assumption
2 / 23
Comparison with independent measurements
2 / 23
Alleged leak
1 / 23
Reported measurement
1 / 23
Calculated comparison
1 / 23
Forecast
1 / 23
Value judgment
1 / 23

Review Outcomes

Unverified 9

Further evidence, clarification or validation is needed. Each claim records the required check.

Checks and claims
Verified within scope 6

The attribution or conditional arithmetic checks out; underlying events and inputs may remain unverified.

Checks and claims
Pending 1

A future outcome; verification requires a deadline and fixed resolution criteria.

Checks and claims

Fallacies

2 findings · 4 unresolved · full preserved text reviewed

Findings

Unresolved cases (4)
Passages and reasoning
L01 · Calibration across hardware Finding

Hasty generalization · C16 · C18 · C19

ratios between chips and models are sound

Similar errors on the tested systems do not establish cancellation on different architectures. Two biases within the same broad band can still distort their ratio.

Article passage · Hasty generalization criterion

Reasoning and repair
  1. The reported throughput bias is similar across the selected B200/B300 runs.
  2. Other deployment anchors produce a similar broad bias band.

Conclusion: Ratios across the modeled chips and models remain sound.

Counterreading: A planning heuristic could use common-mode error as an explicit assumption. The sentence presents ratio validity as a result, rather than an assumption.

Repair: Restrict the conclusion to tested comparisons, or publish matched cross-system ratio residuals.

L02 · Capacity versus demand Finding

Faulty comparison · C13 · C14 · C17 · C20

This is comfortably above the demand range, but not ridiculously so. Seems in the right ballpark.

The comparison does not establish fleet sufficiency: it uses an upper Flash capacity against partial mixed-model demand.

Article passage · Faulty comparison criterion

Reasoning and repair
  1. The Flash projection derates to an upper capacity near 23T total tokens/day.
  2. Router and revenue estimates cover a mixture of models and only part of total demand.

Conclusion: The fleet can comfortably serve the current demand.

Counterreading: The phrase “ballpark” permits a rough plausibility check. It does not make the stronger fleet-sufficiency conclusion follow from unmatched quantities.

Repair: Compare sustained capacity with a matched model/input mix, complete demand coverage and the same latency target.

L03 · Nine-month training ceiling Unresolved

Accident · C21 · C22

which exceeds the theoretical 9-month Epoch ceiling

The ceiling wording overstates the source, but the surrounding argument can be read as a qualified economic comparison. That ambiguity prevents a confirmed fallacy label.

Article passage · Accident criterion

Reasoning and repair
  1. Epoch estimates an economically preferred training duration under hardware and algorithmic progress.
  2. The proposed smaller fleet requires thirteen months.

Conclusion: A frontier-scale run on that fleet is less plausible.

Counterreading: A thirteen-month run can be less attractive without being impossible; the article explicitly calls the frontier-scale outcome less plausible.

Repair: Replace the universal ceiling language with the economic assumptions and test whether they hold for the lab.

Analysis: The estimate uses opportunity costs from hardware and algorithmic progress.

L04 · Generated versus total tokens Unresolved

Equivocation · C13 · C14 · C20

~33 Flash (~8.6 Pro) Ttok/day of generated tokens

The labeling error is established in Claims. Its role in the reasoning is not clear enough to call it confirmed equivocation.

Article passage · Equivocation criterion

Reasoning and repair
  1. The summary calls both output-only and input-plus-output quantities generated tokens.
  2. Later calculations distinguish total from output tokens.

Conclusion: The larger token figures support the fleet-capacity comparison.

Counterreading: The summary may contain a copy error. The later definitions let the reader recover the intended quantities.

Repair: Use total and output consistently, then reassess the actual capacity comparison.

L05 · The position attributed to critics Unresolved

Straw man

everything else is just commentary

There is a risk of refuting a stronger position than the cited delay argument. The broad target and omitted embedded posts prevent a definitive attribution finding.

Article passage · Straw man criterion

Reasoning and repair
  1. The article describes a school of analysis that treats process nodes as decisive.
  2. Systems-level compensation can offset some chip-level disadvantages.

Conclusion: The described outlook is undermined by system-level compensation.

Counterreading: Some critics may actually assert the extreme position. Others, including the cited Transformer argument, argue for delay rather than permanent prevention.

Repair: Name the exact proposition and source being rebutted, and test the strongest delay-and-cost version.

Opening argument: The author advocates slowing development, not establishing permanent physical impossibility.

L06 · Policy and economic damage Unresolved

False cause

those negative externalities - like our present memory crunch - create objectively worse economic conditions at home and abroad

The causal claim is under-supported in this edition. Missing premises alone do not establish a false-cause pattern.

Article passage · False cause criterion

Reasoning and repair
  1. The article associates a policy outlook with constraints and shortages.

Conclusion: That outlook causes economic harm through the cited externalities.

Counterreading: This may summarize an argument made in the missing supporting material, rather than infer causation from coincidence.

Repair: Supply the mechanism, comparison and relevant alternatives, then review the complete inference.