China is Closing the AI Quality Gap.

Three Chinese AI labs put major models into the open in the space of a week.

Start with who they are, because they are not the same kind of company. DeepSeek is a research lab that spun out of the Chinese hedge fund High-Flyer, and it competes almost entirely on cost. Moonshot AI is a Beijing startup whose Kimi models chase the top of the capability rankings and which sells to consumers as well as developers. Alibaba is not a startup at all. It is a listed company with a large retail and cloud business, and its Qwen models are the AI arm of that business.

All three shipped inside seven days. Moonshot published the full weights for Kimi K3 on July 27. DeepSeek released V4-Flash on July 31. Alibaba launched Qwen3.8-Max on August 3, and its shares rose more than 5 percent that day.

The models are good. One caveat before the scores, though. These are different tests measuring different things, and some of the numbers come from the labs themselves rather than from independent reviewers.

Alibaba says Qwen3.8-Max scored 93.0 on PaperBench, a test of whether a model can reproduce a research paper's results from scratch. That puts it ahead of OpenAI's GPT-5.6 Sol at 90.5 and Anthropic's Claude Opus 4.8 at 80.3. That figure is Alibaba's own and has not been independently checked, so treat it as a claim rather than a result.

Kimi K3 scored 156 on Epoch AI's capability index, which combines many separate tests into a single score. That is a record for any openly published model.

DeepSeek's V4-Flash is a different kind of release. It is a budget model, and on Artificial Analysis's overall index it scores nine or more points below Claude Opus 5, Claude Fable 5 and GPT-5.6. What makes it notable is the price. It performs close to Claude Opus 4.8 on difficult coding work while charging $0.14 per million tokens of input and $0.28 of output. Artificial Analysis puts it at more than a hundred times cheaper than Claude Fable 5.

They built all three on less money and with worse hardware. US export rules restrict which Nvidia chips can be sold into China. The capability still rivals American models built with far more money behind them.

Now the honest counter. That Kimi score lands between two American models released in February and March of 2026, so the Chinese labs are running four to five months behind. That is closer than the seven-month average of the past three years, but it is still behind, and no Chinese model has held the top spot on that index at any point since 2023. Epoch also notes that openly published models tend to be tuned harder for benchmarks than closed ones, which means the real gap may be wider than the scores suggest.

So both things are true. The Chinese models are very close. They are still not first. What surprised me was not the capability. It was what investors are paying for it.

Valuations

Two terms first, because every figure below uses them.

To sanity-check what a company is worth, divide its valuation by the revenue it earns in a year. A company worth $100 million that earns $10 million trades at ten times revenue. The higher that number, the more the investor is paying for growth that has not arrived yet.

Run rate is the second. It means the current monthly pace multiplied by twelve. A company earning $2 billion last month has a $24 billion run rate, even if it has only been earning at that pace for a few weeks. It is the standard measure for companies growing this quickly, and I have used it on both sides.

Company

Latest raise

Valuation

Revenue run rate

Multiple

Anthropic

$65B, May 2026

$965B

$47B

21x

OpenAI

$122B, March 2026

$852B

$25B

34x

DeepSeek

$1.5B, in talks

$71B

$400M to $500M

about 158x

Moonshot

$3.5B, late July 2026

$35B

$300M

about 117x

Anthropic and OpenAI are the two most valuable private companies ever created, and both trade at multiples a normal growth investor would recognize. DeepSeek and Moonshot do not.

Put the two pairs side by side. Anthropic and OpenAI are worth $1.82 trillion together and earn $72 billion. DeepSeek and Moonshot are worth $106 billion together and earn about $750 million.

Investors are paying roughly five times more for each dollar the Chinese labs earn than for each dollar the American labs earn.

That is the opposite of the story most people tell. China is not the cheap way to own this. It is the expensive one.

Alibaba sits outside that comparison because it is a public company with a large retail business attached. Its AI products earn about $5.3 billion a year inside a company worth $293 billion.

Open Weights Against Closed

The reason for those multiples is strategy, and these three releases make it clear.

First, what a model's weights actually are. When a model is trained, the result is an enormous set of numbers that determines how it responds to anything you give it. Those numbers are the model. Everything else is packaging.

Publishing the weights means handing over that file. Anyone can download it, run it on their own computers, look inside it and build on it, without asking permission or paying anyone. Keeping the weights closed means the file never leaves the company. You send your request to their servers, they run the model, they send back the answer. You are renting the output rather than owning the tool.

That difference matters more than it sounds. If you hold the weights, you control the cost, you control where your data goes, and you control when the model changes. If you do not, then you are renting access that someone else controls.

Now the three labs. DeepSeek publishes its model under an MIT license, which means anyone can download it and run it on their own machines, free, forever, including commercially. Moonshot publishes Kimi K3's weights too, but under its own license with conditions that trigger above certain revenue levels, so it is open to download without being open source in the strict sense. Anthropic and OpenAI publish nothing at all. You reach their models by paying for access and no other way.

Alibaba spent the first half of 2026 moving toward the American approach, keeping its best models behind paid access. Qwen3.8-Max reverses that. It is the first Alibaba model at this level to open its weights, with the files due to publish within days.

Giving the model away is not generosity. It is a way to win users, and it is working. AI companies bill by the token, which is roughly a word or part of a word. Everything a model reads and writes gets counted and charged.

OpenRouter is a service that lets developers send work to whichever model they choose, open or closed, and it publishes what they actually pick. It does not capture ChatGPT's consumer traffic or the enterprise work routed straight through Microsoft and Amazon, so read it as a signal about what developers choose when every option is on the menu.

Between January and June of 2026, DeepSeek doubled its share of that traffic from 9 percent to 18 percent and became the single most used model provider on the platform. Chinese models passed American ones in total usage in early June. A year earlier, American models handled about three quarters of it.

Price is how DeepSeek did it, at both ends of its range. Its flagship, V4 Pro, charges $0.43 per million tokens of input and $0.87 per million of output. OpenAI's flagship, GPT-5.6 Sol, charges $5.00 and $30.00 for the same million tokens. DeepSeek is roughly twelve times cheaper going in and thirty-four times cheaper coming out.

Those OpenAI prices are already the discounted ones. It cut them on July 30, 2026, the day before DeepSeek shipped V4-Flash. The price war is not coming. It is here.

The Catch

Here is the part the triumphant version of this story leaves out. OpenRouter says directly that the jump in usage did not produce a matching jump in money spent. The Chinese labs took the usage and left the revenue behind, because taking the usage was the entire point of the price. Free weights and near-free pricing buy you users. They do not buy you profit, and so far they have not turned into any.

Which means DeepSeek's $71 billion and Moonshot's $35 billion are not prices on the businesses that exist today. They are prices on businesses that have not been built yet, set by investors buying into a Hong Kong listing window. That window has already reset once in 2026. Two Chinese model developers listed there lost more than 40 percent of their value in a fortnight when early investors were first allowed to sell.

One caveat I will not skip, because it cuts against my own point. DeepSeek gives its model away. Every company that downloads it and runs it in-house pays DeepSeek nothing, so its reported revenue understates how much of the world's AI work it is actually doing. That is not an accident. It is the strategy working exactly as designed, and it is also why $71 billion is a bet rather than a valuation.

So have they caught up?

  1. On quality, not yet but very close.

  2. On the models developers pick when every option is available, they passed.

  3. On revenue, they earn about one percent of what the two American labs earn.

Half the developer traffic. One percent of the revenue. Nobody has proven yet which of those two numbers matters more.

The closed labs have the revenue and cut prices last week to defend it. The open labs have the users and no way yet to charge them. Both are spending heavily to find out who was correct, and this week did not resolve it.

What China did this week is important. Two of the best models in the world were published for anyone to download, the third arrives within days, and the most expensive one got cheaper the day before. The fight has no winner yet. The people learning to build with these tools are winning anyway.

Bashar Aboudaoud
Managing Member, UpRound

Q3 2026 Live Webinar

Live Webinar: Thursday, September 17

You leave with three things. Where private valuations actually sit right now. Which sectors are repricing fastest. How I am reading prices into 2027.
20 minutes, then live Q&A. 20 slots only for UpRound readers.

Referral Program

Share with 10 people to receive a quarterly call with Bashar.

Keep Reading