K3 lists at half of GPT-5.6 Sol, and the saving is either 9.6% or 30.6%
Artificial Analysis publishes both figures, for the same metric and the same two models, on two of its own pages on the same day. Both are in the table. Neither is endorsed here.

On this page
Moonshot AI’s Kimi K3 lists at $15.00 per million output tokens. OpenAI’s GPT-5.6 Sol lists at $30.00. Exactly half, on both vendors’ own pricing pages, read on 6 August 2026.
So the saving is 50%, and it is not. Artificial Analysis runs both through its Intelligence Index and publishes what each actually costs to finish the work, and this is where I have to stop and report a problem rather than a number.
Its article gives $0.94 per task for K3 against $1.04 for Sol, a 9.6% saving. Its model page for the same model, read the same day, carries the same metric with different values: costPerIntelligenceIndexTask is 0.8552 for K3 and 1.2325 for Sol, which is a 30.6% saving. Same publisher, same metric, same afternoon, one figure three times the other.
Table: Cost per Intelligence Index task, as published by Artificial Analysis in two places on 6 August 2026.
| Source | Kimi K3 | GPT-5.6 Sol | Saving |
|---|---|---|---|
| Article prose | $0.94 | $1.04 | 9.6% |
| Model page dataset | $0.8552 | $1.2325 | 30.6% |
I cannot tell you which is right, and that is the finding. An earlier version of this page put 9.6% in its own title, having read one of those pages and not the other. What survives both readings is the direction and the mechanism: the saving is real, it is nowhere near the 50% the list prices imply, and the reason is token volume, which Moonshot documents on the same page that carries the price.
Where I was already standing
I had this model in a costed build before the measurement existed. In what an always-on AI agent actually costs to run I put moonshotai/kimi-k3 on a three-model allowlist at $3.00 in and $15.00 out, marked it escalation only, and wrote that the output column is the side of the ledger you control least, because you can budget your input and you cannot budget how much a model decides to write.
I derived that rule from the list price. It was the right rule and I could not have told you the mechanism, because I had a price and nobody had published a cost. The mechanism is now on the vendor’s page, one paragraph under the price table:
Always reasons and supports configuring its reasoning effort with the top-level
reasoning_effortrequest field (low/high/max, defaultmax).
Two words carry the whole gap. Always, so there is no cheap non-reasoning path. Default max, so the most expensive setting is the one you get by not choosing. Reasoning tokens bill as output tokens, at $15.00 per million.
How much more it emits, and where the arithmetic stops
Take the floor first, because it needs no assumptions about the workload at all.
Every one of K3’s unit prices is 60% or less of Sol’s: input $3.00 against $5.00, cached input $0.30 against $0.50, output $15.00 against $30.00. So for any mix of cached, uncached and output tokens whatsoever, a task that consumed identical token counts on both models would cost K3 at most 60% of Sol’s $1.04, which is $0.62. It costs $0.94. Divide: K3 is consuming at least 1.5 times the tokens, and that number falls out of four published prices and two published costs with no modelling in between.
The obvious next step is to turn that floor into an exact multiple, and it does not survive contact with the definitions. Artificial Analysis publishes a per-task token figure for Sol: 15k output tokens per Intelligence Index task at max reasoning effort. For K3 it publishes totals instead, 130M tokens across the index at a cost of $2,437.41. Dividing that total by the $0.94 per-task cost looks like it hands you a task count, and K3’s tokens per task would fall straight out of it. But the measurer defines that per-task number two different ways. The v4.1 announcement that introduced the metric says it takes “the total cost, total time, and total output tokens for a model to run the Intelligence Index and divide by the number of tasks across its evaluations”. The K3 model page calls the same metric a “Weighted average cost (USD) per Artificial Analysis Intelligence Index task”, and the index weights are nowhere near uniform: GDPval-AA v2 carries 20% of the score, AA-Omniscience Non-Hallucination 4%. Under the first reading the division is a task count. Under the second it is an unweighted total over a weighted mean, which counts nothing. One of the two readings voids the arithmetic, so the arithmetic does not ship and 1.5x is what I am willing to put my name to.
The measurer agrees on direction, in plain language. Artificial Analysis’s own summary calls K3 “notably slow and very verbose” and sets its 130M tokens against a 100M median among open-weight models of similar size, which is a narrower comparison set than the full index.
The one place the list price tells the truth
There is a tier where the headline discount is real and understated, and the post I am building on missed it.
Table: List prices per million tokens from each vendor’s own pricing documentation, read 6 August 2026. OpenAI’s page does not state the input-token threshold at which its long-context tier engages, so the last row applies above an unstated boundary.
| Per million tokens | Kimi K3 | GPT-5.6 Sol | K3 as % of Sol |
|---|---|---|---|
| Input, cached | $0.30 | $0.50 | 60.0% |
| Input, standard | $3.00 | $5.00 | 60.0% |
| Output, standard | $15.00 | $30.00 | 50.0% |
| Output, long context | $15.00 | $45.00 | 33.3% |
K3 charges one flat output rate all the way to its 1,048,576-token window. Sol charges a premium above a threshold. Above that line the gap widens from half to a third, and there is an architectural reason rather than a commercial one. vLLM’s engineering write-up describes Kimi Delta Attention as a linear-attention mechanism holding a fixed-size recurrent state instead of a growing KV cache, interleaved with periodic full-attention layers, and says directly that this is what makes a 1M-token context affordable. Flat long-context pricing is that design showing up on an invoice.
So the honest shape is the inverse of the headline. On short agentic turns, where reasoning dominates the bill, the 50% discount collapses to 9.6%. On very long single-pass context, where the recurrent state does the work, it improves to 33%.
The 2.7% figure measures a preference vote, and the ratio is undefined
Stanford’s 2026 AI Index says the top US model leads by 2.7% as of March 2026. That number is real and the popular restatement of it fails three ways.
It is an Arena rating, which the report itself describes as human voting on ratings “inspired by chess ratings”. It measures which answer people preferred, not what either model can do. The underlying pair, per press reads of the report, is Claude Opus 4.6 at 1,503 against ByteDance’s Dola-Seed-2.0-Preview at 1,464. A 39-point difference; 1503/1464 gives 2.66%.
Dividing two Elo ratings is also not a defined operation. Elo is an interval scale with an arbitrary zero, so the ratio moves when you shift the origin while nothing about the models changes at all.
Table: The same 39-point Arena gap expressed as a percentage, with every rating on the board shifted by a constant. Win rate is 1 / (1 + 10^(-39/400)) throughout.
| Shift applied to all ratings | Gap as a percentage | Head-to-head win rate |
|---|---|---|
| none | 2.66% | 55.6% |
| +500 | 1.99% | 55.6% |
| +1,000 | 1.58% | 55.6% |
The quantity that survives the shift is the expected score. The top US model was preferred in about 56 of every 100 blind comparisons, which is a far less comfortable number than 2.7% for anyone arguing the models are interchangeable. It is also five months stale for August 2026, predating Kimi K3 entirely.
What the report actually concludes is better than the statistic anyway: competitive pressure is shifting “toward cost, reliability, and domain-specific performance”. That is the argument, from the primary source.
Open weights, with the licence read
The weights are open and the licence is new, which is the part worth reading. Moonshot’s own back catalogue is not a guide to it. Kimi K2 and Kimi K2.6 each ship a file whose first line is Modified MIT License, and the K2 text says exactly what the modification is:
Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display “Kimi K2” on the user interface of such product or service.
K3 ships a bespoke Kimi K3 License instead. Its section 3 is that same branding clause with the model name changed. Its section 2 has no counterpart in the K2 file at all:
If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
Read the trigger carefully. It is the aggregate revenue of the licensee and its affiliates over any twelve months, not the revenue earned from serving K3, and 20 million dollars is a small company. An inference provider above that line needs a signed agreement with Moonshot before it charges for a single token, unless it is one of the certified inference partners that section 4(b) exempts. The claim that you can deploy the weights yourself and fine-tune them survives, but through section 4(a), which exempts internal use, defined as use that does not make the software, its outputs or its underlying capabilities available to third parties.
Self-hosting has a hardware floor too. vLLM’s launch post states the model “can barely fit in a single NVIDIA DGX B300 and requires a minimum of 16 NVIDIA B200/GB200 GPUs to serve on that hardware generation”. Hugging Face’s repository metadata for the model reports 2.78 trillion parameters, 2.72 trillion of them stored as packed 4-bit MXFP4, across 1.56 TB of safetensors. That is the download, before a byte of KV cache or recurrent state.
Which undercuts the data-residency argument for most of the named adopters, because they are not self-hosting. Airbnb uses Alibaba’s Qwen for customer service; DoorDash routes lower-level coding work to Kimi for, in its CTO’s words, “better quality [at] cheaper cost”. Both consume an API, and K3 reaches Azure customers through Fireworks on Microsoft Foundry. Routing through a US intermediary to a Chinese-trained model is a different data posture from running weights on your own metal, and US lawmakers requested information from DoorDash about exactly this on 31 July.
That distribution route also corrects the source post: Microsoft is not an adopter, it is the channel. Kimi K2.6 and K2.5 are sold directly through Microsoft Foundry. The American hyperscaler is monetising the Chinese price advantage as a distributor, which is a better fact than the one it replaces. Siemens should come out of a list about American buyers too; it is a German industrial group.
Three corrections to the surrounding numbers
The semiconductor selloff is real and smaller than it reads. The SOX fell 1.6% on 17 July to close 20.2% below its 22 June record, its steepest week since April 2025, and the same report notes the index was still up more than 60% on the year. A 20% drawdown from a 105% run is profit-taking with a bear-market label attached.
The $725 billion capex figure is four hyperscalers, not five, and “roughly” rather than “more than”. It is an analyst aggregation of company guidance whose own component table sums to well under its headline, with Microsoft appearing at two different values on one page.
And the trend runs both ways. K3’s $15.00 output price is 3.75 times its own predecessor’s $4.00. The lab leading the price attack quadrupled its own rate, which is the strongest available objection to the commoditisation thesis.
Detecting this in your own stack
The test is one query against your provider’s billing export, and it works for any model pair.
- Take one real task class, not a prompt. A whole agent turn including tool calls and retries.
- Pull total spend and total completions for that class over a week, and divide. That is your cost per task. Per-token price is an input to it, never a substitute for it.
- Split output tokens into reasoning and visible answer. If your provider reports reasoning tokens separately, the ratio is the story.
- Compute the floor: multiply the incumbent’s cost per task by the challenger’s worst unit-price ratio. If the challenger costs more than that, it is emitting more tokens, and you now know the minimum multiple.
- Re-run at each
reasoning_effortlevel before you conclude anything. A default ofmaxis a benchmark setting, not a production setting.
Designing so the bill cannot surprise you
Set reasoning_effort explicitly on every call, including the ones where you want the default, so the value is in your code rather than in the vendor’s. Cap max_tokens per task class, because an uncapped verbose model has no upper bound you control. Price migrations on measured cost per task and refuse to sign off on a per-token comparison. And route by tier: long single-pass context to the flat-rate model, short reasoning-heavy turns to whichever model finishes in fewer tokens.
What I have now that I did not
A floor calculation that converts two published prices and two published costs into a hard lower bound on token consumption, which I can run on any model pair the day it launches. The knowledge that reasoning_effort defaults to max on the model I had already put on an allowlist, which turns a rule I wrote on instinct into one I can defend. An Elo shift table that retires percentage comparisons of Arena ratings permanently. And the correction that reorganises the whole thesis: the cheap model is cheap per token and average per task, so the commoditisation of foundation models is arriving through a different door than the price list suggests.
Half price bought 9.6%. Measure the task.
Sources
Every source below was opened and checked on the date shown. Links open in this tab.
- American AI Has a Pricing Problem Mo RezaAli on X x.com Accessed 6 August 2026
- Flagship Model Kimi K3 Pricing Kimi API Platform Documentation, Moonshot AI platform.kimi.ai Accessed 6 August 2026
- Create Chat Completion Kimi API Platform Documentation, Moonshot AI platform.kimi.ai Accessed 6 August 2026
- Pricing OpenAI Platform Documentation developers.openai.com Accessed 6 August 2026
- Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index Artificial Analysis artificialanalysis.ai Accessed 6 August 2026
- Kimi K3 model page Artificial Analysis artificialanalysis.ai Accessed 6 August 2026
- GPT-5.6 has landed Artificial Analysis artificialanalysis.ai Accessed 6 August 2026
- Kimi K3 License Moonshot AI, Hugging Face huggingface.co Accessed 6 August 2026
- moonshotai/Kimi-K3 model card Moonshot AI, Hugging Face huggingface.co Accessed 6 August 2026
- LICENSE, moonshotai/Kimi-K2-Instruct (Modified MIT License) Moonshot AI, Hugging Face huggingface.co Accessed 6 August 2026
- LICENSE, moonshotai/Kimi-K2.6 (Modified MIT License) Moonshot AI, Hugging Face huggingface.co Accessed 6 August 2026
- Repository metadata for moonshotai/Kimi-K3, parameter counts by dtype and file sizes Hugging Face Hub API huggingface.co Accessed 6 August 2026
- Artificial Analysis Intelligence Index v4.1: a shift toward agentic workloads Artificial Analysis artificialanalysis.ai Accessed 6 August 2026
- Kimi K3 Is Here: Efficient Day-0 Support on vLLM vLLM Team and Inferact vllm.ai Accessed 6 August 2026
- 2026 AI Index Report Stanford Institute for Human-Centered Artificial Intelligence hai.stanford.edu Accessed 6 August 2026
- 2026 AI Index Report, Technical Performance Stanford Institute for Human-Centered Artificial Intelligence hai.stanford.edu Accessed 6 August 2026
- Stanford's AI Index finds China has nearly closed the performance gap with the US despite spending 23 times less The Next Web thenextweb.com Accessed 6 August 2026
- Stanford AI Index 2026: the numbers that matter Digital Applied www.digitalapplied.com Accessed 6 August 2026
- Chinese AI models are challenging US rivals on cost Fortune fortune.com Accessed 6 August 2026
- US companies are realizing Chinese AI models are way cheaper Futurism futurism.com Accessed 6 August 2026
- Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI The Decoder the-decoder.com Accessed 6 August 2026
- Philadelphia Semiconductor Index Falls Into Bear Market After 20.2% Slide From Record High Bloomingbit en.bloomingbit.io Accessed 6 August 2026
- Semiconductor Stocks Plunge Into Bear Market As China's New Kimi K3 AI Model Rattles Wall Street HNGN www.hngn.com Accessed 6 August 2026
- Models sold directly by Azure Microsoft Learn, Azure AI Foundry learn.microsoft.com Accessed 6 August 2026
- Moonshot AI's Kimi K3 reaches enterprise users through Fireworks on Microsoft Foundry TechNode technode.com Accessed 6 August 2026
- U.S. lawmakers request information from DoorDash on use of Chinese AI models CNBC www.cnbc.com Accessed 6 August 2026
- AI Spending Tracker 2026 Value Add VC valueaddvc.com Accessed 6 August 2026