The claim you have probably heard: Chinese AI models now match the American flagships at a fraction of the price. Independently measured, that is half true, and it is not yet testable for the model making the loudest headlines. On the Artificial Analysis Intelligence Index, an independent composite score built from reasoning, coding and agent-style tests across 186 models, Anthropic's Claude Fable 5 leads at 60. Alibaba's flagship Qwen3.7-Max scores 46, rank 28, below even Gemini 3.6 Flash, the speed-focused model that is currently the strongest Gemini available. The Chinese model that genuinely plays at the frontier is Moonshot AI's Kimi K3 at 57, and on July 27 Moonshot shipped its weights, so K3 is the one flagship here you can actually download and own, under a license with conditions. Alibaba's shipping Max flagship stays a rental API, exactly like the American ones. What you are really trading in 2026 is capability, real per-task cost, and whose jurisdiction risk becomes your downtime.
Qwen 3.8 exists, but you cannot measure it yet
If you searched that name, you are ahead of the leaderboards. On July 19, 2026, at the World AI Conference in Shanghai, Alibaba announced Qwen3.8: a claimed 2.4 trillion parameters, its first multimodal model above the trillion mark, with a purchasable Qwen3.8-Max-Preview on Alibaba's own platforms at 10% of standard rates, and open weights promised with no date attached. Alibaba's own framing is "second only to Fable 5." What does not exist yet: a model card, a license, a benchmark table, an active-parameter count, or any independent score. A vendor claim you cannot test is marketing, not measurement.
So the comparison you can act on today uses the flagship Alibaba actually ships and independent testers actually score: qwen3.7-max, released May 19, 2026, listed in Alibaba Cloud's Model Studio catalogue with qwen3.7-plus and qwen3.6-flash below it. This piece measures Qwen3.7-Max against Claude Fable 5, with the four other current flagships as context rows: OpenAI's GPT-5.6 Sol, xAI's Grok 4.5, Google's Gemini 3.6 Flash, and Moonshot AI's Kimi K3. When Qwen3.8 ships with real documentation and an independent score, the radar below gets a seventh polygon.
Does "same capability at a fraction of the price" hold up?
Capability: Qwen3.7-Max is not the rival, Kimi K3 is
On capability, no. Claude Fable 5 is rank 1 of 186 at 60. GPT-5.6 Sol sits at 59 when run at its highest reasoning setting, Kimi K3 at 57, xAI's Grok 4.5 at 54, Gemini 3.6 Flash at 50, and Qwen3.7-Max at 46. Alibaba's flagship is not Fable 5's capability rival. Kimi K3 is: only three points off the top, above every Google model available today.
Price: the 5x sticker gap is 2.7x per completed task
On price, the sticker misleads in the other direction. AI usage is billed in tokens, the word-chunks a model reads and writes; one token is roughly three quarters of an English word. Fable 5 costs $10 per million tokens in, $50 out; Qwen3.7-Max costs $2.50 and $7.50. The sheet says 4x to 7x. But models differ wildly in verbosity, meaning how many tokens they burn to finish an answer. Fable 5 produced 87 million output tokens on the independent evaluation against a 63 million average, and its text splitter produces roughly 30% more tokens on the same input than Claude models did before the Opus 4.7 generation. Measured per completed task, the gap compresses: Fable 5 costs $2.75 per task, Qwen3.7-Max $1.03. That is 2.7x, not 5x. Kimi K3 ($0.95) and GPT-5.6 Sol ($1.04) land at essentially the same real cost despite very different stickers. The outliers sit below them: Gemini 3.6 Flash at $0.50, and Grok 4.5 at $0.31, the cheapest measured task in this comparison, the product of a $2 in / $6 out sticker and unusually concise output (60 million evaluation tokens against the 63 million average).
Speed: the axis that splits the Chinese pair
Speed splits the Chinese pair. Qwen3.7-Max is the 4th-fastest model tracked, streaming about 200 tokens per second; Kimi K3 measures 35 tokens per second with a 4.5-second pause before the first word. Grok 4.5 sits in between: about 69 tokens per second, with an 11.6-second wait before the first token on xAI's API, which rules it out of interactive work even though it is the cheapest per task. So the honest identities are: Qwen3.7-Max is a speed-and-price model, Kimi K3 is the near-frontier option that makes your team wait, and Grok 4.5 is the batch-work bargain.

Six models, the questions that matter
| The question | Claude Fable 5 | GPT-5.6 Sol | Kimi K3 | Grok 4.5 | Gemini 3.6 Flash | Qwen3.7-Max |
|---|---|---|---|---|---|---|
| Who makes it, where? | Anthropic, US | OpenAI, US | Moonshot AI, China | xAI, US | Google, US | Alibaba, China |
| How capable, independently scored? | 60, rank 1 of 186 | 59, rank 2 | 57, rank 4 | 54, rank 9 | 50 | 46, rank 28 |
| Price per 1M tokens, in / out? | $10 / $50 | $5 / $30 | $3 / $15 | $2 / $6 | $1.50 / $7.50 | $2.50 / $7.50 |
| What does one real task cost? | $2.75 | $1.04 | $0.95 | $0.31 | $0.50 | $1.03 |
| How much can it read at once? | 1M tokens | 1.05M tokens | 1M tokens | 500K tokens | 1M tokens | 1M tokens |
| Can you download and own it? | No | No | Yes, under the K3 License | No | No | No |
| Who answers your regulator? | Anthropic; EU Code signatory; mandatory 30-day data retention | OpenAI; EU Code signatory | Entity and data location not published; not an EU signatory | xAI; signed only the safety and security chapter of the EU Code | Google; EU Code signatory | Entity not published; not an EU signatory; regions outside mainland China available if explicitly pinned |
| Content policy independently measured? | Refusal behaviour documented by vendor | No independent measurement | No independent measurement | No independent measurement | No independent measurement | No independent measurement |
Scores are the Artificial Analysis Intelligence Index; GPT-5.6 Sol's score is measured at its maximum reasoning setting and Grok 4.5's at its high setting. Per-task costs are that firm's measured cost to complete its evaluation tasks, which is why verbose models look worse than their stickers and terse ones look better.

Draw your own comparison
Tick the models you are weighing. Same data as the chart above; each axis is scaled to the best of the six, so farther out is better.
Runs in your browser. Sources: Artificial Analysis Intelligence Index v4.1 model pages and vendor pricing pages, late July 2026.
Who answers your regulator?
The leaderboard is the easy half. The harder question: when something goes wrong, which legal entity answers, and under whose rules did your data travel?
The EU signatory line cuts both ways
Anthropic, OpenAI and Google are all full signatories of the EU's voluntary Code of Practice for general-purpose AI, the compliance companion to the EU AI Act, whose general-purpose obligations have applied since August 2, 2025. xAI has signed only the Code's safety and security chapter, which leaves it to demonstrate transparency and copyright compliance by other means. Alibaba and Moonshot AI are not on the list at all.
Accountability cuts both ways, though. On June 12, 2026 the US government placed export controls on Claude Fable 5, and Anthropic suspended it globally for nearly three weeks before redeploying it on July 1. A US-accountable vendor also means exposure to US government leverage. And Fable 5 carries mandatory 30-day data retention with no zero-retention option, a real caveat inside the "accountable" column.
Where the data actually travels
The Chinese column is more layered than "your data goes to China." Alibaba's international Model Studio runs from regions outside the mainland (Singapore, US, Frankfurt and Tokyo), with request data stored in the selected region: a defensible data-residency story. The catch is a default. The "Global" inference scope can route requests through any available node, including nodes inside China. A team that never explicitly pins the region has a different data flow than it believes. Alibaba's Model Studio documentation also does not say which legal entity answers a regulator. Moonshot is thinner still: its international platform's pricing docs state neither the operating entity nor the data location. That is the weakest documented posture of the six.
The censorship question nobody has measured
On censorship and content policy, start with what nobody can tell you. No independent measurement exists for Qwen3.7-Max or Kimi K3. The US standards body NIST, through CAISI, has evaluated other Chinese models: DeepSeek models echoed four times as many inaccurate state narratives as US reference models, complied with 94% of malicious jailbreak requests versus 8%, and their agents were twelve times more likely to follow hijack instructions. CAISI has also measured Moonshot's earlier Kimi K2 Thinking, finding it heavily censored in Chinese, in line with the most censored PRC models it has tested. None of those findings transfer to K3 or to any Qwen model, which CAISI has not evaluated. Unmeasured is not the same as bad, and it is not the same as safe; in a risk assessment it must be written down as unmeasured.
The ownership option came back, at exactly one flagship
Update, 6 August 2026: When this piece first ran, no flagship in the comparison could be downloaded and owned. That changed on July 27, when Moonshot shipped Kimi K3's weights on time. The section below is updated to reflect it; the rest of the comparison stands.
Our earlier Kimi piece framed this choice as two currencies of trust: rent accountability from a Western vendor, or own control by downloading open weights, the model files you can run on your own servers. At the flagship tier that second currency is now real, but scarce. Moonshot shipped Kimi K3's weights on July 27 under a bespoke license, so K3 is the one model here you can actually own a copy of, with license fine print attached and a roughly 1.4 terabyte footprint that still needs datacenter hardware. The rest stay rentals: Alibaba's shipping Max flagship is API-only, and the open weights Alibaba promises for the just-announced Qwen3.8 still carry no date. Owning a copy narrows the sovereignty question but does not erase it: the license, the jurisdiction, and the censorship record travel with the weights to whatever host runs them. Choosing Chinese no longer buys sovereignty by default; with K3 it can now buy an owned copy, and with everything else it still buys a discount from a landlord who does not publish who answers your regulator.
Real ownership now reaches the flagship tier, but thinly: Kimi K3 is the one near-frontier model you can download, and its size keeps it on datacenter hardware, not your laptop. One tier down the choice is wider and lighter. Alibaba's open-weight Qwen3.5 122B scores 32 on the same index, about half of Fable 5's 60 and 14 points below Alibaba's own flagship, but it runs on modest hardware. So the honest sovereignty question is no longer "US or China" but "which tier of ownership fits my workload, and can I live with the license and jurisdiction attached to it?" And the accountable column carries its own government-leverage exposure, as the Fable 5 suspension showed. Both ends of the chart share one hidden axis: how much of your vendor's jurisdiction risk you inherit as your own downtime.



