← Insights AI-frontier & modellen 17 August 2026 19 min Written with AI assistance

The Other AI Race

Why the race is not about the best model but about the chain beneath it.

Ruben Horbach Ruben Horbach Co-founder
Download as pdf

01

Why this, why now

In January 2025 a Chinese lab called DeepSeek released a model as good as the best American reasoning model, open, at a fraction of the cost, and wiped hundreds of billions off Nvidia's value in a day. That was the alarm.

Eighteen months later it is not an alarm any more, it is the market, and the market moved again while we were writing. We published the first edition of this dossier on 17 July 2026. That same day Moonshot released Kimi K3, a 2.8-trillion-parameter model it bills as the largest open-source model in the world. We missed it by hours, which is a reasonable summary of how this subject behaves.

This edition does two things the first one did not. It updates the adoption figures, which have moved a long way in a month. And it takes the headline number apart, because the most quoted statistic in this debate turns out to be measuring a workload shift as much as a national one, and because the same is true of the chip figures on the other side of the argument. Adjudicating that is more useful than repeating either number louder.

Our underlying claim survives that scrutiny. The question at home is who builds the single most capable model, and the American labs still lead that. The question that decides where the world's AI actually runs is who builds the layer everyone deploys, and that is increasingly answered in Hangzhou, Beijing and now Shenzhen.

Timeline

  • Jan 2025 · DeepSeek-R1 lands, matching a top US model at open weights and a fraction of the cost. A "DeepSeek selloff" hits Nvidia.
  • Apr 2025 · The US declares the H20 chips that trained DeepSeek noncompliant with export controls.
  • Nov 2025 · Chinese models reach roughly 15 percent of global usage on one widely quoted tracker.
  • 15 Jan 2026 · US export policy shifts H200 sales to China from presumption of denial to case-by-case, with a 25 percent revenue share.
  • Apr 2026 · DeepSeek V4, Kimi K2.6 and Qwen 3.6 ship inside four days, all open, all frontier-class. V4 is tuned for Huawei silicon.
  • Jun 2026 · Reporting confirms DeepSeek training some models on Huawei chips.
  • 17 Jul 2026 · Moonshot releases Kimi K3 at 2.8 trillion parameters.
  • 29 Jul 2026 · An AEI working paper estimates Huawei could meet half of China's compute demand by 2028.

02

Contents

1. The model you didn't notice

2. The Sputnik moment and what it proved

3. The cost-structure moat

4. Open weights as the default layer

5. What the 45 percent actually measures

6. The reframe: capability was never the moat

7. How far behind, exactly

8. Chips: the controls that backfired

9. Huawei at 50 percent and at 4 percent

10. Whose values sit in the default model

11. Europe caught in the middle

12. Where we could be wrong

13. What a European business picks, and why

14. Verification and sources

03

1. The model you didn't notice

Start with the number that reframes the conversation, and then spend section 5 being careful about it.

On OpenRouter, one of the largest marketplaces for model usage, Chinese providers now serve more than 45 percent of all traffic, against less than 2 percent a year ago. The single largest is not one of the labs a Western reader would name. It is Xiaomi, at around 21 percent of the marketplace by weekly tokens, against roughly 7.5 percent for OpenAI.

The Western view is still shaped by the frontier: which lab has the most capable model this quarter. That contest is real and the US still leads it, and we think it is the wrong contest for understanding where the world's AI runs. Four of the top five open-weight models are Chinese, and Alibaba's Qwen passed a billion downloads faster than any model family in history. When a developer anywhere reaches for a model they can run and change themselves, more often than not it is Chinese.

04

2. The Sputnik moment and what it proved

DeepSeek-R1, released in January 2025, made this legible. It matched the reasoning quality of the best American model of the day, shipped with open weights and a permissive licence, and had reportedly been trained for a few million dollars of compute. The market read it instantly, and Nvidia and a string of US tech stocks fell in what was dubbed the DeepSeek selloff.

The lasting lesson went past the one model. Andrew Ng, writing days after the release, named three things at once. China had closed the capability gap faster than most of the West believed. Open weights were commoditising the foundation-model layer, turning the expensive part into something a developer could reach for dollars. And raw scale was not the only path to progress, because the DeepSeek team, denied the best chips, had innovated on efficiency instead.

Ng also gave the concrete version of the price gap, which we keep because it dates the trend: OpenAI's o1 cost $60 per million output tokens, DeepSeek R1 cost $2.19. Every one of those three observations has only become truer since.

05

3. The cost-structure moat

The most important figure in this story is not a benchmark. It is a price.

A US frontier model through its API can run around $30 per million tokens. A comparable Chinese open model runs around $0.14. When DeepSeek released V4 in April 2026, one widely circulated summary put its flash tier at $0.28 per million tokens and called it 99 percent cheaper than the comparable Western model. That is a different order of magnitude, not a discount.

For a startup the gap is existential, and the arithmetic is short enough to do here. A company burning 100 million tokens a month pays roughly $300,000 on a US frontier API and roughly $1,400 on a Chinese open model. On a $5 million seed round, and holding everything else equal, that is the difference between around three months of runway and around eighteen. a16z has said that something like 80 percent of the startups pitching it build on Chinese open models, and once you have done the division that stops being surprising. As one investor put it, economics beats nationalism every time.

06

4. Open weights as the default layer

Price is half of it. The half we think matters more is that these models are open, and open changes the shape of the market.

A closed API is rented and cannot be altered. An open-weight model is held: it runs on hardware the buyer controls, it can be fine-tuned for one exact case, inspected, and kept inside the building along with the data it sees. For a large class of organisations, especially those with privacy, latency or sovereignty constraints, that moves the model from a nice-to-have to the obvious choice. The licences reinforce it. Zhipu's GLM-5 ships under MIT, Alibaba's Qwen flagship under an Apache-style commercial licence, Moonshot's Kimi as open weights.

This is how a default forms. Every company that survives the next funding winter will have built around whatever was cheapest and most flexible, chosen as a survival strategy rather than a China strategy. Once a product, its fine-tuned variants and a team's know-how sit on a particular model family, switching is expensive. The layer everyone builds on becomes sticky, and the layer everyone is building on now is open and Chinese.

07

5. What the 45 percent actually measures

Now the part the July edition should have done and did not.

That marketplace share is the most quoted number in this debate, and it is measuring two things at once. Alongside the shift towards Chinese providers, the workload on that marketplace changed shape: programming went from roughly 11 percent of usage at the start of 2025 to more than half by mid-2026. Coding is exactly the category where Chinese open models are strongest relative to price, and it is also the category where developers are most willing to swap providers, because the output is immediately testable. A model that writes bad code fails visibly in seconds, which makes cheap models safe to try in a way they are not for, say, medical summarisation.

So some unknown part of "Chinese providers reached 45 percent" is really "the marketplace became a coding marketplace, and Chinese models win on coding price". Both statements are true. They support quite different conclusions.

I spent an afternoon trying to split them and could not, and we want to say why, because the reason generalises. To separate a workload effect from a provider effect you need share by category over time: Chinese share within coding, and Chinese share within everything else, at two dates. The public dashboards give you the totals and the category totals, not the cross-tabulation, and without it any decomposition is a guess dressed as a calculation. What can be said is a bound rather than a figure. Coding went from about a ninth of the marketplace to more than half, and if Chinese models had held their non-coding share flat while winning only the new coding volume, that alone would move the aggregate by a large fraction of the observed rise. That is not a finding. It is a reason to distrust anyone who quotes the 45 without mentioning the composition, and until July that included us.

BFF chart · OpenRouter usage data via industry reporting, 2026.
BFF chart · OpenRouter usage data via industry reporting, 2026.

We attach two further cautions to the figure. OpenRouter is one marketplace, skewed towards developers and startups rather than the regulated enterprise buying through a hyperscaler, so it systematically under-samples the segment where Western incumbents are strongest. And a token is not a decision: a model doing high-volume, low-stakes work generates far more tokens than one doing a small number of consequential ones, so token share overstates the importance of cheap, chatty workloads by construction.

What survives all of that we still find substantial. Under 2 percent to over 45 in twelve months is not a composition effect on its own, whatever share of it the workload shift explains. The direction is not in question. We would quote the magnitude with the denominator attached, and in July we quoted it without.

08

6. The reframe: capability was never the moat

Step back and, as we read it, the strategic picture inverts.

Silicon Valley spent years assuming the moat in AI was model quality: build the smartest system and the market follows. What DeepSeek and its peers demonstrated is that the moat was cost structure. The American labs optimised for the ceiling, the most capable model, sold at a premium, funded by billions raised on the promise of general intelligence. The Chinese labs, boxed in by the chip embargo, optimised for efficiency first and distribution second, and gave the models away to become the default.

We should be precise about who is winning what. On the absolute frontier, the hardest reasoning and the newest capabilities, the US still leads. On the layer the world deploys at scale, the cheap, open, good-enough model, China is ahead and pulling further. Silicon Valley built its moat around a ceiling that most applications never touch, and left the floor to whoever could make intelligence cheapest.

09

7. How far behind, exactly

"Caught up" and "still behind" are both doing a lot of work in this debate without either side putting a number on it. Epoch AI has, and the number is more useful than the adjectives.

Measuring open models against state-of-the-art closed ones on a common capability index, Epoch reported in June 2026 that open models lag the closed frontier by about four months. Not four years, and not zero.

That single figure disciplines both camps. Four months is short enough that for the overwhelming majority of commercial work the distinction is invisible: almost nothing a mid-sized company deploys depends on capability that arrived last quarter rather than the one before. It is also long enough to matter where it matters, in the hardest research, the longest reasoning chains and the newest modalities, which is precisely where the American labs concentrate.

It also reframes the price comparison in section 3. Paying roughly 200 times more per token buys, on this measure, about four months of capability lead. Whether that is worth paying depends entirely on whether the work in question sits in the part of the distribution where the lead shows up. For most organisations, honestly assessed, it does not.

The forward view is where the disagreement concentrates, and I would treat all of it as scenario rather than forecast. One modelling effort published in July 2026 puts China's first model of a given frontier class at around February 2027, and reports that the interventions usually proposed to delay it, restricting remote access to compute and blocking both H200 sales and smuggled chips, move that date only modestly, because each measure splits its effect across several Chinese firms rather than landing on one. Anthropic's May 2026 paper reaches close to the opposite conclusion, arguing that a co-ordinated squeeze on compute and on distillation could lock in a 12-to-24-month American lead by 2028. Four months today, somewhere between two months and two years in 2028, depending which model you believe and who paid for it. That spread is the honest state of knowledge, and any dossier claiming better precision than that is selling something.

10

8. Chips: the controls that backfired

The American response to all this was export controls: deny China the best Nvidia chips and slow it down. The controls were real and they bit. Compute remains China's tightest constraint, and its labs still train on a patchwork of restricted and older hardware.

The second-order effect ran opposite to the intent, and a paper posted to arXiv in July 2026 argues the point in its title: US policies unintentionally accelerated China's open AI ecosystems. Denial did what subsidies could not. It handed Huawei a captive market and gave every Chinese lab a reason to design around Nvidia. DeepSeek's V4, tuned to train on Huawei's Ascend hardware, is the proof that a domestic, good-enough stack now exists, and reporting in June 2026 confirmed DeepSeek training some models on Huawei silicon rather than smuggled or grandfathered Nvidia parts.

Then the policy turned around, and what happened next is the thing we find most instructive in this section.

On 15 January 2026 the Bureau of Industry and Security moved H200 and comparable AMD parts from presumption of denial to case-by-case review for China, subject to a 25 percent tariff, a volume cap, third-party testing and know-your-customer requirements, with the US government taking a quarter of the sale revenue. Licences were approved for around ten Chinese firms.

Almost nothing shipped. Chinese authorities, according to contemporaneous reporting, have not permitted domestic AI companies to buy the chips.

We read that twice, because it inverts three years of assumptions. The constraint on Nvidia selling to China is no longer only Washington. Having spent a year being told they could not rely on American silicon, Chinese buyers and their government built a domestic alternative and then declined the American one when it was offered. A share of a market lost during an embargo does not automatically return when the embargo lifts, and the more general lesson is that the effects of a supply restriction outlive the restriction, because they change what the other side builds.

11

9. Huawei at 50 percent and at 4 percent

Here is the adjudication we wrote this dossier for, because two figures about the same company are circulating, both are approximately right, and they support opposite conclusions.

Huawei holds roughly half of China's AI-chip market, with Nvidia down to something like 8 percent, the mirror image of a few years earlier. An AEI working paper published on 29 July 2026 projects Huawei meeting about half of China's total compute demand as soon as 2028.

Huawei may produce around 4 percent of Nvidia's aggregate compute in 2026, falling to 2 percent in 2027 as Nvidia's own output scales. That figure comes from an Anthropic policy paper published in May 2026.

Both can be true because they have different denominators. The first measures Huawei's share of a national market that was largely closed to its main competitor. The second measures Huawei's output against global production. A company can dominate a walled garden and remain small against the world, and Huawei currently does both.

Which number matters depends on the question. If you are asking whether China can keep its own labs supplied, the domestic share is the relevant one and the answer is increasingly yes. If you are asking whether China can match American training runs at the frontier, global compute is the relevant denominator and the answer is not yet, and not soon on current output. Epoch made a version of this argument in April 2026 under the heading that China is not about to leap ahead of the West on compute, and on their denominator they are right.

The mistake, made constantly in both directions, is quoting one figure to settle the other's question. A commentator who says "Huawei has half the market" to argue that export controls failed at the frontier is using the wrong denominator. So is one who says "Huawei is 4 percent of Nvidia" to argue that Chinese labs will run short of chips at home.

I want to flag the provenance here rather than bury it in the appendix, because the sourcing is asymmetric in a way that matters. The domestic-share figures come from market trackers and a think-tank working paper. The global-compute figure comes from a policy paper published by Anthropic, an American frontier lab arguing for tighter controls on a competitor. That does not make the number wrong, and I have no better one. It does mean the two sides of this adjudication are not equally disinterested, and a reader should weigh them accordingly. The version of this dossier published in July carried the domestic figure alone, which flattered the argument it was making.

12

10. Whose values sit in the default model

There is a quieter consequence of a Chinese default that we think gets too little attention, and it has nothing to do with price.

A model reflects the data, the fine-tuning and the constraints of whoever built it. A model trained and aligned under Chinese rules carries Chinese defaults on the topics where that matters, from history to politics to what it will and will not discuss. For a marketing chatbot this is irrelevant. For a model embedded in a newsroom, a school, a court or a government office, it is not.

Ng put the strategic version bluntly at the start of all this: if the US keeps stifling open source, the world will end up running on models that reflect China's values far more than America's, simply because those are the models that are available and affordable. That is not a claim about any particular model's behaviour, and it should not be read as one. It is a claim about defaults, and defaults are chosen by availability far more often than by deliberation.

13

11. Europe caught in the middle

Europe sits uncomfortably in this race, building neither the frontier models nor the cheap open ones at the scale that matters. Its own champion, France's Mistral, is real and respected and small against the American and Chinese giants. Its instinct, understandably, is sovereignty.

That instinct produces revealing moments. The European Parliament dropped Google for the French search engine Qwant to cut its dependence on American technology, and observers pointed out that Qwant's results are substantially served by Microsoft's Bing. It is the lesson we keep returning to: swapping the interface is not sovereignty if the layer underneath is unchanged.

For AI the question is sharper. A European organisation choosing between an expensive American model and a cheap Chinese open one is choosing between two dependencies, neither European, each with its own strings. The European position we would defend in 2026 is not independence. It is informed, deliberate dependence, with the switching costs understood in advance rather than discovered later.

14

12. Where we could be wrong

The bullish reading of Chinese AI can be overstated, and we put four counterweights against our own case.

The frontier lead is real and it is measurable. Section 7 puts it at about four months on open-versus-closed capability. For most work that is nothing. For the hardest problems it is the whole game, and it is the segment with the highest margins.

Compute, on the global denominator, still favours the US heavily. Anthropic's May 2026 paper argues the US and its allies could lock in a 12-to-24-month frontier lead by 2028 by closing China's access to advanced compute and to copied model outputs, and it frames chips as the gatekeeper for training, deployment, revenue and iteration alike. We would note that this is a policy paper from a company with a direct commercial interest in the conclusion, which does not make it wrong and does mean it should not be read as a neutral assessment. The same paper characterises distillation of American model outputs as systematic industrial espionage, which is a contested framing rather than an agreed fact.

Benchmarks are not deployment. Leading a leaderboard is not the same as being the model an enterprise trusts with its most sensitive work, and Western incumbents still hold that trust in regulated industries. Security and provenance are genuine concerns with open weights from any source, and specific concerns attach to models whose training and alignment sit within a particular government's reach.

And our own headline number is softer than it looks. Section 5 is the counter-case to section 1, which is an uncomfortable structure for a dossier and the one we would rather publish.

We hold two things at once. On raw frontier capability and on global compute, the American position leads and that lead is worth something. On the economics and distribution of everyday AI, China has built a position that is hard to unwind, because it is made of cost, openness and a domestic supply chain rather than any single model that could be leapfrogged. A strategy that only watches the frontier will keep being surprised by the floor.

15

13. What a European business picks, and why

For a Dutch or European company the practical question is narrower: which model to build on, and how to avoid being trapped. Three moves, and they are what we would do.

We would match the model to the job rather than to the flag. For a great deal of everyday work, a cheap open model, run where you control it, is the right answer on cost and on data, and section 7 says the capability you give up is measured in months rather than generations. For the hardest reasoning, or for regulated work where provenance and support matter, a frontier Western model may earn its premium. That is a decision per workload, taken with the four-month figure in hand, rather than a company-wide posture.

Take the sovereignty question literally. The questions are where the weights run, who can see the data, what happens if the licence changes, and how long a migration would actually take. The Qwant example in section 11 is what happens when that question is answered at the level of the logo rather than the layer.

And we would stay portable, which is the whole lesson of this race. The model layer is commoditising and shifting fast enough that a dossier can be overtaken between drafting and publication, as this one was. We rate architecting to be swappable across models above betting on any provider or any country. The durable value sits in the data, the workflows and the people, not in the dependency.

Our move: choose the model per task, keep the architecture portable, and treat every model, American or Chinese, as a supplier you might one day need to replace rather than a foundation you are stuck with.

16

14. Verification and sources

This dossier draws on live web research and a personal archive of more than 15,000 sources. The notes below flag confidence and the material caveats.

ClaimConfidenceNote
Chinese providers now serve more than 45% of OpenRouter traffic, against under 2% a year earlier; Xiaomi ~21% of weekly tokens against OpenAI ~7.5%MediumOne marketplace, developer-skewed, measured in tokens rather than decisions or revenue. Section 5 sets out why the figure overstates a national shift: coding rose from ~11% of that marketplace's usage in early 2025 to over half by mid-2026, and Chinese models are strongest relative to price on coding. Direction well supported, magnitude not cleanly separable.
Kimi K3 released 17 July 2026 at 2.8 trillion parameters, billed as the largest open-source modelMedium-highMoonshot announcement and contemporaneous reporting. Parameter count is a vendor claim; "largest open-source model" is a claim about a fast-moving field and may already be stale.
Four of the top five open-weight models are Chinese; Qwen passed a billion downloads faster than any model familyHighOpen-weight leaderboards and Hugging Face data, 2026. Leaderboard composition changes monthly.
Open models lag state-of-the-art closed models by about four monthsMedium-highEpoch AI, June 2026, measured on a common capability index. A single-metric summary of a multi-dimensional gap; different capability axes would give different answers.
Cost: Chinese open models ~$0.14 per million tokens against ~$30 for a US frontier API; DeepSeek V4 flash reported at $0.28 per millionMedium-highProvider pricing and contemporaneous analysis, 2026. Exact ratios vary sharply by model, tier and context length. The $0.28 figure and the "99% cheaper" framing come from a widely shared social-media summary rather than a primary price list, and are labelled as such in the text.
A 100M-token-per-month workload costs roughly $300,000 on a US frontier API against roughly $1,400 on a Chinese open model; runway implicationHigh as arithmeticOur calculation from the prices above. It ignores fine-tuning, hosting and engineering costs, which narrow the gap for self-hosted open models.
OpenAI o1 at $60 per million output tokens against DeepSeek R1 at $2.19HighAndrew Ng, January 2025, contemporaneous with the R1 release. Historical, retained to date the trend.
a16z: ~80% of startups pitching it build on Chinese open modelsMediumPartner comments, widely reported, not a published survey.
Huawei ~50% of China's AI-chip market against Nvidia ~8%; AEI projects Huawei meeting ~half of China's compute demand by 2028Medium-highMarket-share estimates and an AEI working paper, 29 July 2026. Projections, not measurements.
Huawei may produce ~4% of Nvidia's aggregate compute in 2026 and ~2% in 2027MediumAnthropic policy paper, May 2026. A policy paper from a company with a commercial interest in the conclusion. Reported because it is the clearest statement of the global-denominator case, not because it is neutral. Section 9 adjudicates the two figures rather than choosing between them.
H200 policy shifted to case-by-case on 15 January 2026 with a 25% revenue share and volume cap; ~10 Chinese firms licensed; almost nothing shipped because Chinese authorities have not permitted purchasesMedium-highBIS final rule and contemporaneous reporting. The non-shipment is reported rather than officially confirmed by either government, and is the kind of claim that could change quickly.
DeepSeek V4 tuned for Huawei Ascend; DeepSeek training some models on Huawei chipsMedium-highThe Information, June 2026, and contemporaneous analysis.
US export policies unintentionally accelerated China's open AI ecosystemsMediumarXiv paper, July 2026. A preprint; we have not seen it peer-reviewed. Cited because it argues the section 8 thesis formally, not as settled evidence.
EU Parliament switched to Qwant, whose results are substantially Bing-backedHighVerified separately in BFF Signal work.
Chinese models at ~15% of global usage in November 2025Medium, supersededRetained in the timeline for shape only. Definitions of "usage" differ so much between trackers that the November figure and the current marketplace figure should not be treated as points on one series. The July edition of this dossier drew a line between them, and that was a mistake.

The framed figure in this dossier is taken from a publicly posted chart, credited to its original source in the caption. Charts labelled "BFF" are our own, drawn from the sources named beneath them.

Dossier as pdf

Download this dossier

The full dossier as a pdf, with every figure and the source list. Fill in your details and the download starts right away.

We use your details to give you this dossier and to contact you about it. More about that in our privacy statement.

Ruben Horbach

Ruben Horbach

Co-founder · Back From the Future

Ruben researches how organisations adopt AI meaningfully — not as technology, but as a change in work and people. He builds the agent infrastructure behind BFF and speaks about the near future of work.

Translate this to your situation?

Book a conversation — we're happy to think along about what this means for you.