Kimi K3 Benchmarks: How Close Is It to Fable 5 and GPT-5.6 Sol?
Kimi K3 explained in full. Moonshot AI's 2.8 trillion parameter open-weight model compared against Claude Fable 5 and GPT-5.6 Sol across every major benchmark, plus pricing, architecture, and what independent testing actually confirms.

On July 16, 2026, Moonshot AI released a model that made Silicon Valley pay attention for reasons beyond the usual launch day hype. Kimi K3 is now the largest open-weight AI model ever released, and it's landing within a few points of Claude Fable 5 and GPT-5.6 Sol on some of the hardest benchmarks in the industry. This is the complete picture: what K3 actually is, how it stacks up against the two most advanced proprietary models on the market, and which of the headline claims independent testing has actually confirmed so far.
What Is Kimi K3?
Kimi K3 is Moonshot AI's newest large language model, developed by the Beijing based startup behind the Kimi chatbot and the earlier Kimi K2 family. It began rolling out on July 16, 2026, first inside Kimi Code and the Kimi app, arriving a day ahead of schedule after a promotional page on Moonshot's own developer platform leaked the release early.
The model is built as a Mixture-of-Experts system with roughly 2.8 trillion total parameters, more than twice the size of DeepSeek's 1.6 trillion parameter V4 Pro, making K3 the largest open-weight model released to date by a meaningful margin. It ships with a 1,048,576 token context window and native multimodal understanding across text and images. Two variants launched together: K3 Max, tuned for chat and general agent tasks, and K3 Swarm Max, built for large-scale parallel processing workloads.
One detail worth flagging clearly: K3 is available today through Moonshot's API and inside Kimi's own products, but the open weights themselves are not yet public. Moonshot has committed to releasing them by July 27, 2026, meaning independent researchers can't yet inspect, modify, or self-host the model.
Architecture: What's Actually Confirmed
Moonshot describes K3 as using KDA (a gated linear attention variant) alongside Attention Residuals for computational efficiency, according to details published through OpenRouter's model listing. Separately, some early technical writeups describe the architecture as "Stable LatentMoE," reportedly activating only 16 experts out of 896 at inference time, which would explain how a 2.8 trillion parameter model can run at competitive latency. That specific figure hasn't yet been confirmed in an official Moonshot model card, so treat it as an early report rather than a verified spec until Moonshot's full technical documentation lands alongside the July 27 weights release.
Benchmarks: The Complete Breakdown
This is where K3's launch became a genuine story rather than routine model release noise. Moonshot's own released benchmark table places K3 directly against Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 across more than 30 tests. A few important caveats before the numbers: Moonshot's own testing mixed different coding harnesses (Kimi Code for K3, Claude Code for Anthropic's models, Codex for OpenAI's models), meaning conditions weren't fully identical across every comparison. Independent evaluator Artificial Analysis has since published its own separate testing, which is the more reliable reference point where available.
Coding and Agentic Benchmarks (Moonshot's released figures)
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 |
|---|---|---|---|---|
| SWE Marathon (long-horizon coding) | 42.0 (leads) | 35.0 | 39.0 | 40.0 |
| Program Bench | 77.8 (leads) | 76.8 | 77.6 | Not listed |
| Terminal Bench 2.1 | 88.3 | 84.6 | 88.8 (leads) | 84.6 |
| DeepSWE | 67.5 | 70.0 | 73.0 (leads) | Not listed |
| FrontierSWE | 81.2 | 86.6 (leads) | 71.3 | Not listed |
| Kimi Code Bench 2.0 (Moonshot's internal benchmark) | 72.9 | 76.9 (leads) | 64.8 | Not listed |
| Automation Bench | 30.8 (leads) | 29.1 | 29.7 | Not listed |
| SpreadsheetBench 2 | 34.8 (leads) | 34.7 | 32.4 | Not listed |
| BrowseComp (long-context search) | 91.2 (leads) | 88.0 | 90.4 | Not listed |
| JobBench | 52.9 | 57.4 (leads) | 46.5 | Not listed |
Across roughly 35 tests Moonshot published in total, K3 took first place in about seven of them, and landed second or third in most of the rest, consistently beating Opus 4.8, GPT-5.5, and GLM-5.2 by wide margins. Fable 5 still won the largest single share of individual benchmarks, and GPT-5.6 Sol kept a narrow lead on the specific Terminal Bench 2.1 test that both companies have used as a headline metric this cycle. K3 also reportedly performed competitively with Fable 5 (in its fallback configuration) and substantially outperformed Opus 4.8, GPT-5.6 Sol, and GPT-5.5 specifically on GPU kernel optimization tasks, a technical measure of how efficiently a model can improve hardware throughput.
Independent Verification: Artificial Analysis Results
Artificial Analysis, the industry's most cited independent benchmarking group, ran its own separate evaluation and published results within hours of K3's launch. On its Intelligence Index v4.1, K3 scored 57.1, placing it fourth overall, behind Claude Fable 5 (59.9 with Opus 4.8 fallback), GPT-5.6 Sol (58.9 at max reasoning), and just 0.54 points behind Sol's higher xhigh reasoning setting, while landing ahead of Claude Opus 4.8 (56) and GPT-5.5.
On more granular independent measures, K3 scored 76.24 on Artificial Analysis's Coding Index and 50.07 on its Agentic Index, running at roughly 62 output tokens per second with a 1.99 second time to first token. On AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reached an Elo of 1,547, a jump of 732 points over the previous Kimi K2.6 model, with only Fable 5 scoring higher. On GDPval-AA, a real-world professional task evaluation, K3's Elo climbed to 1,668, up sharply from K2.6's 1,190.
One number worth sitting with: on the AA-Omniscience Index, K3's accuracy improved from 33 percent to 46 percent generation over generation, a genuinely strong gain. But its hallucination rate climbed at the same time, from 39 percent to 51 percent, meaning the model is both getting more questions right and fabricating more answers than before. That's a real tradeoff worth understanding if you're deploying K3 anywhere factual accuracy matters more than raw task completion.
Community and Preference Testing
Blind testing on LMArena, where human evaluators compare unlabeled model outputs, put K3 at #1 on the platform's Frontend Code Arena, a 17-place jump from K2.6's #18 ranking, taking first place in six of seven frontend development categories and landing second specifically in Gaming, behind Fable 5. Design Arena, a separate community preference platform, also ranked K3 first overall in its results. Both are live preference boards with genuinely wide confidence intervals on young vote counts, so treat them as directional signal rather than settled fact.
Kimi K3 vs Claude Fable 5: The Direct Comparison
Across the benchmark data available, the relationship between K3 and Fable 5 is consistent: Fable 5 remains ahead on the composite Intelligence Index and wins the largest share of individual benchmark categories, particularly FrontierSWE, Kimi Code Bench 2.0, and JobBench, where its lead is wide and clear. K3 closes the gap or takes the lead on several other specific tests, including SWE Marathon, Automation Bench, SpreadsheetBench 2, and BrowseComp, and edges ahead on Frontend Code Arena's blind preference testing. On pricing, the gap is dramatic. Fable 5 costs $50 per million output tokens. K3 costs $15, less than a third of that, while landing within single digits of Fable 5 on the composite intelligence score. Moonshot itself described K3 as performing "competitively" with Fable 5 rather than claiming outright superiority, and that framing holds up against the independent data.
Kimi K3 vs GPT-5.6 Sol: The Direct Comparison
This matchup is closer and more genuinely split. GPT-5.6 Sol keeps a narrow lead on Terminal Bench 2.1 (88.8 versus 88.3, a half-point gap) and leads clearly on DeepSWE (73.0 versus 67.5). K3 pulls ahead on Program Bench, SWE Marathon, FrontierSWE by a wide margin (81.2 versus Sol's 71.3), Automation Bench, SpreadsheetBench 2, BrowseComp, and decisively on Kimi Code Bench 2.0 relative to Sol's score there. On the independent Artificial Analysis Intelligence Index, Sol holds a real but narrow lead, 58.9 to 59 versus K3's 57.1. Pricing tells a similar story to the Fable 5 comparison: Sol costs $30 per million output tokens, twice what K3 charges, for a composite intelligence advantage of under two points.
Pricing: The End of Cheap Chinese AI
K3's pricing is arguably as significant a story as its benchmarks. At $3 per million input tokens and $15 per million output tokens, it's priced closer to Anthropic's Claude Sonnet tier than to the deep discount pricing Chinese labs have built their reputation on. For comparison, Moonshot's own previous model, Kimi K2.6, charged just $0.95 input and $4 output. K3's output pricing is also notably higher than competing Chinese models: z.ai's GLM-5.2 charges $4.40 per million output tokens, and DeepSeek V4 charges $0.87. K3 is now the most expensive model ever released by a Chinese AI lab.
That said, the comparison that matters most is against the models K3 is actually competing with. At $15 output versus Fable 5's $50 and GPT-5.6 Sol's $30, K3 remains meaningfully cheaper than either proprietary frontier model it's benchmarking against, even as it abandons the ultra-budget positioning of earlier Kimi releases. Analysts have read this as a broader signal: Chinese labs are increasingly confident enough in their model quality to charge closer to parity with US labs rather than competing purely on price.
The Reality Check Independent Analysts Are Flagging
Enthusiasm around K3's launch has been genuinely intense, with at least one widely followed benchmark commentator suggesting the release makes Opus-tier models obsolete and predicting frontier labs will need to respond with far larger models. It's worth treating that kind of launch-day reaction with real skepticism, and several independent analysts have said so directly.
The specific concerns worth understanding: Moonshot's benchmark table mixed different coding harnesses across different models, meaning the comparisons weren't run under fully identical conditions. Several of K3's claimed results weren't yet visible on the public leaderboards they reference at time of launch. Most importantly, the promised open weights aren't available yet, so no outside party can independently verify Moonshot's full performance claims by running the model themselves. One detailed independent writeup put it plainly: K3 deserves an immediate controlled pilot for coding agents, research, and long-context work, but it doesn't yet justify a blanket replacement of Fable 5 or GPT-5.6 Sol, and any regulated self-hosting decision should wait until Moonshot publishes the actual checkpoint, license, and technical report.
The Ecosystem Around K3
Moonshot didn't just ship a model in isolation. Kimi Code, the company's open-source coding agent that competes directly with Claude Code and Gemini CLI, received two updates the same day as K3's launch, versions 0.25.0 and 0.26.0, adding expanded subagent tooling, background task management, todo lists, plan mode, skill invocation, and nested agents, effectively turning it into a multi-layered autonomous coding system. The Kimi Code CLI has accumulated over 3,100 stars on GitHub and integrates with VS Code, Cursor, and Zed.
Moonshot's models already have real adoption among Western AI companies. Cursor used Kimi to help build Composer 2, its AI coding agent. DoorDash's CTO has publicly said the company delegates lower-level engineering work to Kimi K2.6. Thinking Machines used Kimi K2.5 to generate early post-training data for its Inkling model, released just a day before K3. That existing adoption context matters when evaluating how seriously to take K3's benchmark claims, this isn't an unknown lab's first release, it's a company with a track record of shipping models that real engineering teams have already put into production.
On the business side, Moonshot raised $2 billion in funding in May 2026 at a valuation exceeding $20 billion, with annual recurring revenue reportedly exceeding $200 million, a genuine comeback story for a company whose market position had eroded before this release cycle.
Should You Actually Use It
Use Kimi K3 today if you want to pilot it for coding agents, long-horizon automation, or long-context research work, ideally through a controlled test against your own tasks rather than trusting benchmark rank alone. It's a reasonable candidate to add to a routing setup alongside Fable 5 or GPT-5.6 Sol for the specific categories where it independently tests well, agentic automation, spreadsheet work, and long-context search among them.
Hold off on treating K3 as a wholesale replacement for either proprietary frontier model until two things happen: the open weights land on July 27 as promised, allowing independent, reproducible verification of Moonshot's claims, and enough real-world usage accumulates to confirm the benchmark gains translate into production reliability, particularly given the rising hallucination rate flagged in Artificial Analysis's own testing.
Bottom Line
Kimi K3 is a genuinely significant release, not because it definitively beats Fable 5 or GPT-5.6 Sol, it doesn't, but because it gets close enough to force a real conversation about what a $15 open-weight model can now do next to $30 and $50 proprietary systems. The honest summary of the data: Fable 5 remains the strongest model overall and wins the most individual benchmarks. GPT-5.6 Sol keeps narrow, specific leads on tasks like Terminal Bench and DeepSWE. K3 doesn't beat either consistently, but it's now close enough, and cheap enough, that ignoring it entirely would be a mistake. The next two weeks, when the actual weights ship and independent labs get to run their own tests, will tell you a lot more than launch day benchmarks alone.
