MOONSHOT'S KIMI K3 MATCHES TOP US MODELS AT FRACTION OF THE PRICE
The 2.8 trillion-parameter model beat every Western system except Fable 5 in early blind testing, and it is already available.
by editor6 min readcomments soon

Chinese startup Moonshot AI released Kimi K3 on Thursday, a 2.8 trillion-parameter model that early benchmarks place alongside the best American systems at dramatically lower cost. Developers using the independent Arena evaluator preferred K3 over every leading US model for front-end coding, including Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol. In Arena's broader text ranking, K3 outranked the standard version of Anthropic's Opus 4.8 and tied Sol.
The reaction inside US labs is not subtle. Raffi Krikorian, a former OpenAI executive now at Mozilla, said, "Right now, it's a U.S. versus China question," and described US labs as "clearly worried". Their CEOs have been lobbying Washington against open-weight models, and K3 is a direct reason: it is the first open model in the roughly 3-trillion-parameter range, and its weights are scheduled for release on July 27.
The timing is deliberate. K3 lands just before the 2026 World Artificial Intelligence Conference in Shanghai, where President Xi Jinping is expected to outline Beijing's AI priorities. Meanwhile, DeepSeek, Moonshot's domestic rival, is expected to release an updated model soon. China is not testing the frontier any more. It is competing on it.
WHAT K3 ACTUALLY DOES
The model uses a mixture-of-experts architecture that activates only 16 of 896 total experts at a time, a standard efficiency pattern that lets the full 2.8 trillion parameters exist without requiring 2.8 trillion parameters of compute per inference. It processes text and images within a 1-million-token context window, which is sufficient to analyse entire codebases in a single pass.
K3 is designed for long-running software development with minimal human oversight. It can coordinate terminal tools, navigate large repositories, and sustain focus across dozens of sequential work steps. Moonshot paired the model with a new attention architecture called Kimi Delta Attention, which enables up to 6.3x faster decoding for million-token contexts. The system also uses what the company calls "Attention residuals," which reportedly boost training efficiency by about 25 per cent while adding less than 2 per cent extra compute overhead.
THE BENCHMARK PICTURE
Kim's own evaluations ran across 35 tests. K3 took first place about seven times and landed second or third in most of the rest. Fable 5 won the most individual tests. In nearly every benchmark, K3 beat Opus 4.8, GPT-5.5, and GLM-5.2 by a wide margin. The results were achieved at maximum or high thinking intensity, and three different agent systems were used depending on the test: KimiCode, Claude Code, or Codex. The conditions were not identical across all tests, which matters for comparing results.
Artificial Analysis published the first independent evaluation, giving K3 a score of 57 on its Intelligence Index, on par with Opus 4.8 and GPT-5.5 but behind Fable 5 and GPT-5.6 Sol. On agentic tasks, K3 reached an Elo rating of 1,668 on GDPval v2, up from K2.6's 1,190. It beats GLM-5.2, GPT-5.5, and Claude Opus 4.8 on that benchmark but falls short of Claude Fable 5's 1,760. K3 took the top spot on AutomationBench-AA with a score of 5 per cent. On AA-Briefcase, it reached an overall Elo of 1,547, up 732 points from K2.6. Artificial Analysis called the model "well-rounded" and noted that only Claude Fable 5 scores higher on AA-Briefcase.
GPT-5.6 Sol still leads on presentation quality, but K3's accuracy rate on the AA-Omniscience Index improved from per cent to per cent, pushing the overall score from +6 to +18. The tradeoff is real: K3's hallucination rate climbed from per cent to per cent—more capability, more confident wrong answers.
THE PRICING STORY
K3 is much cheaper than the top Western models while sitting in the upper midrange. One million input tokens cost $0.30 with a cache hit and $3.00 with a cache hitout. One million output tokens cost $15.00. That is above K2.6's pricing ($0.16 cached, $0.95 uncached, $4.00 output) but still aggressive compared to the frontier. Anthropic's new Sonnet 5 costs $3 per million input tokens and $15 per million output tokens, but delivers lower performance than K3.
Artificial Analysis pegs K3 at $0.94 per task on the Intelligence Index, close to GPT-5.6 Sol at $1.04 and about half the price of Opus 4.8 at $1.80. It is above open-weight peers like GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04), but those models score lower. K3 also uses fewer tokens than its predecessor: about 132 million output tokens to complete all nine evaluations, down from 166 million for K2.6, a percentage reduction while scoring 13 points higher.
WHAT IS REAL AND WHAT IS NOT
K3 has been available for only hours. Early benchmarks and viral demonstrations may overstate how reliably it performs across real-world work. The model is accessible now through Kimi.com, mobile apps for iOS, Android, and HarmonyOS, the Kimi Work desktop client, and Kimi Code. On OpenRouter, it is listed under the identifier "moonshotai/kimi-k3" and served only through Moonshot. The open weights are expected by the end of July, and until then, developers cannot independently inspect, modify, or run the models. A planned platform called Kimi Hosted Agent will provide isolated environments and runtimes for long-running tasks.
The caveats matter. Moonshot's own benchmarks used maximum thinking intensity and different agent systems, which means direct comparisons with other models tested under different conditions are inexact. The increase in hallucination rate is a genuine concern for production deployment, especially given K3's target use case of unsupervised coding work. A model that confidently suggests incorrect API calls across 40 steps of an automated pipeline needs careful guardrails.
WHAT IT MEANS FOR THE COMING MONTHS
America's lead in advanced AI is shrinking by the month, and K3 is the strongest signal yet that the gap is no longer measured in years. An open-weight Chinese model that matches or beats proprietary US systems on multiple benchmarks at a lower cost, with a million-token context window and multimodal capabilities, is not a research demonstration. It is a product that shipping developers can begin using today.
The lobbying by US CEOs against open-weight models becomes harder to justify when the best open model comes from China and their own teams cannot match its performance at comparable cost. The argument that open models pose safety risks is strained when the alternative is a closed model that US companies charge premium prices for and a Chinese competitor gives away for fractions of the cost.
K3 does not end the AI competition. It is one data point, launched hours ago, with real flaws and legitimate questions about benchmark conditions. But it is the kind of data point that changes conversations in boardrooms and government offices. The US labs are not wrong to be worried. They have time to respond, but less of it than they had a week ago.
what did you make of it?
more from ai
ai
TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLION
The additional investment will build at least four more 2nm fabs and advanced packaging, bringing the company's total US commitment to $265 billion.
ai
META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AI
The opt-in feature flags self-harm references in chatbot conversations, with human review before any notification is sent.
ai
ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AI
The new toolset, launching July 28, turns text prompts into playable experiences and puts game creation on iPhone and iPad.
ai
ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILURE
A robotaxi drove into an active fire scene obscured by smoke. NHTSA called emergency scenes not edge cases.





