Resources / References / 01
Four to One: Auditing the Leaked DeepSeek Notes
E² review · 2026-09-25
Review
Summary
Support Distribution
Claims Provenance
- Article attribution 9
The claim is traceable to the article. Its exact proposition is not independently established here.
- Conditional calculation 7
Inputs and arithmetic are identified in the article. Availability and outcome of reproduction are counted separately below.
- First-party source checked 4
The exact report is traceable to an inspected vendor, model-card or research source. This verifies the attribution, not the underlying event independently.
- Declared value judgment 1
An objective or preference, outside the empirical support ladder.
Replicability
- Discrepancy identified 4
- Not reproduced 3
- Arithmetic matched 2
- Not validated 1
- C17 Latency and batching require a new validation run. Sensitivity arithmetic does not validate the service target.
Fallacies
Findings
- Hasty generalization 1
- L01 · Calibration across hardware
- Faulty comparison 1
- L02 · Capacity versus demand
Unresolved cases (4)
- Accident 1
- L03 · Nine-month training ceiling
- Equivocation 1
- L04 · Generated versus total tokens
- Straw man 1
- L05 · The position attributed to critics
- False cause 1
- L06 · Policy and economic damage
Passages and reasoning
L01 · Calibration across hardware
ratios between chips and models are sound
Similar errors on the tested systems do not establish cancellation on different architectures. Two biases within the same broad band can still distort their ratio.
Article passage · Hasty generalization criterion
Reasoning and repair
- The reported throughput bias is similar across the selected B200/B300 runs.
- Other deployment anchors produce a similar broad bias band.
Conclusion: Ratios across the modeled chips and models remain sound.
Counterreading: A planning heuristic could use common-mode error as an explicit assumption. The sentence presents ratio validity as a result, rather than an assumption.
Repair: Restrict the conclusion to tested comparisons, or publish matched cross-system ratio residuals.
L02 · Capacity versus demand
This is comfortably above the demand range, but not ridiculously so. Seems in the right ballpark.
The comparison does not establish fleet sufficiency: it uses an upper Flash capacity against partial mixed-model demand.
Article passage · Faulty comparison criterion
Reasoning and repair
- The Flash projection derates to an upper capacity near 23T total tokens/day.
- Router and revenue estimates cover a mixture of models and only part of total demand.
Conclusion: The fleet can comfortably serve the current demand.
Counterreading: The phrase “ballpark” permits a rough plausibility check. It does not make the stronger fleet-sufficiency conclusion follow from unmatched quantities.
Repair: Compare sustained capacity with a matched model/input mix, complete demand coverage and the same latency target.
L03 · Nine-month training ceiling
which exceeds the theoretical 9-month Epoch ceiling
The ceiling wording overstates the source, but the surrounding argument can be read as a qualified economic comparison. That ambiguity prevents a confirmed fallacy label.
Article passage · Accident criterion
Reasoning and repair
- Epoch estimates an economically preferred training duration under hardware and algorithmic progress.
- The proposed smaller fleet requires thirteen months.
Conclusion: A frontier-scale run on that fleet is less plausible.
Counterreading: A thirteen-month run can be less attractive without being impossible; the article explicitly calls the frontier-scale outcome less plausible.
Repair: Replace the universal ceiling language with the economic assumptions and test whether they hold for the lab.
Analysis: The estimate uses opportunity costs from hardware and algorithmic progress.
L04 · Generated versus total tokens
~33 Flash (~8.6 Pro) Ttok/day of generated tokens
The labeling error is established in Claims. Its role in the reasoning is not clear enough to call it confirmed equivocation.
Article passage · Equivocation criterion
Reasoning and repair
- The summary calls both output-only and input-plus-output quantities generated tokens.
- Later calculations distinguish total from output tokens.
Conclusion: The larger token figures support the fleet-capacity comparison.
Counterreading: The summary may contain a copy error. The later definitions let the reader recover the intended quantities.
Repair: Use total and output consistently, then reassess the actual capacity comparison.
L05 · The position attributed to critics
everything else is just commentary
There is a risk of refuting a stronger position than the cited delay argument. The broad target and omitted embedded posts prevent a definitive attribution finding.
Article passage · Straw man criterion
Reasoning and repair
- The article describes a school of analysis that treats process nodes as decisive.
- Systems-level compensation can offset some chip-level disadvantages.
Conclusion: The described outlook is undermined by system-level compensation.
Counterreading: Some critics may actually assert the extreme position. Others, including the cited Transformer argument, argue for delay rather than permanent prevention.
Repair: Name the exact proposition and source being rebutted, and test the strongest delay-and-cost version.
Opening argument: The author advocates slowing development, not establishing permanent physical impossibility.
L06 · Policy and economic damage
those negative externalities - like our present memory crunch - create objectively worse economic conditions at home and abroad
The causal claim is under-supported in this edition. Missing premises alone do not establish a false-cause pattern.
Article passage · False cause criterion
Reasoning and repair
- The article associates a policy outlook with constraints and shortages.
Conclusion: That outlook causes economic harm through the cited externalities.
Counterreading: This may summarize an argument made in the missing supporting material, rather than infer causation from coincidence.
Repair: Supply the mechanism, comparison and relevant alternatives, then review the complete inference.