On 20 January 2025, a Hangzhou lab most Western investors had never heard of released an AI model under a permissive MIT licence — and within a week had rewritten the market’s assumptions about what frontier AI costs.[1] DeepSeek-R1, a reasoning model that matched OpenAI’s o1 on several hard benchmarks, followed DeepSeek-V3: a 671-billion-parameter mixture-of-experts system whose technical report disclosed a final training run of 2.788 million H800 GPU-hours — roughly $5.6 million in rented compute.[2]
The number detonated. On 27 January, Nvidia fell 16.9% and shed close to $600 billion of market value in a single session — the largest one-day loss for any company in US history.[3] The narrative wrote itself: if a Chinese team could reach the frontier for the price of a mid-size house, the case for hundred-billion-dollar compute build-outs looked shakier.
What the headline number leaves out
The $5.6 million figure is real, but narrow. DeepSeek’s own paper is explicit that it covers only the final official training run and “excludes the costs associated with prior research and ablation experiments on architectures, algorithms, or data.”[2] Independent analysis by SemiAnalysis put DeepSeek’s cumulative hardware and capital spend well above $500 million.[4] The efficiency was genuine; the “trained for the price of a car” framing was not.
The story was never the training bill. It was that a frontier capability had suddenly become free to download.
What made DeepSeek matter is less the training bill than the distribution model. R1 shipped open-weight, with distilled versions from 1.5B to 70B parameters and an API priced roughly 90–95% below o1.[1] A capability that had been a paid, gated service became something any developer could download, inspect and self-host for free. That is not a cost story; it is a commoditisation story — and commoditisation of the model layer pushes value up the stack, toward applications and inference infrastructure.
The scale of the reaction told its own story. A $600 billion single-day move is not the market repricing one model; it is the market discovering how little it had understood about China’s trajectory. The correction was overdone in the other direction too — cheaper, more efficient models expand total AI usage rather than shrinking demand for compute, a dynamic economists call Jevons’ paradox. What endured from the episode was not a number but a template: a Chinese lab competing by giving capability away, distilling it into small models a laptop can run, and pricing the API to the floor. That template, not the training bill, is what recurs.
Why it matters for investors
DeepSeek is a preview of a structural feature of China’s AI sector, not a one-off. Its labs have strong incentives to compete on openness and price rather than defend closed-model margins they cannot easily hold. For anyone allocating to the AI stack, the lesson is to underwrite the layers that benefit when models get cheap and abundant — deployment, tooling, inference silicon and vertical applications — rather than assuming the model itself is the moat. The frontier is now a fast-moving, partly commoditised input. Priced correctly, that is an opportunity, not a threat.
References
- TechCrunch, “DeepSeek claims its ‘reasoning’ model beats OpenAI’s o1 on certain benchmarks,” Jan 2025. Read source ↗
- DeepSeek-AI, “DeepSeek-V3 Technical Report (arXiv:2412.19437),” Dec 2024. Read source ↗
- TechCrunch, “Nvidia drops $600BN off its market cap amid the rise of DeepSeek,” Jan 2025. Read source ↗
- NBC New York (citing SemiAnalysis), “DeepSeek’s hardware spend could be as high as $500 million, new report estimates,” Jan 2025. Read source ↗
Every source above is public and linked. This is a research brief from Singularity Dynamics; full briefings on this topic are available on request.