Kimi K3 vs Claude Fable 5: Complete Benchmark and Pricing Comparison (2026)
Kimi K3 vs Claude Fable 5, fact-checked. Every benchmark, the real pricing gap, independent verification, and what Moonshot's claim of matching Fable 5 actually holds up to.

Moonshot AI said its new model performs "competitively" with Claude Fable 5, the most capable model Anthropic has shipped. That's a serious claim from a company most of the West had barely heard of eighteen months ago. This is a full fact check of that claim, benchmark by benchmark, plus everything else worth knowing before you decide whether K3 belongs anywhere near your production stack.
The Claim
On July 16, 2026, Moonshot AI released Kimi K3, a 2.8 trillion parameter open-weight model, and positioned it directly against Claude Fable 5 in its own launch materials. The company's stated framing: K3 performs competitively with Fable 5 and substantially outperforms Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5. That's a specific, checkable claim, not vague marketing language, and it's worth testing against the actual data rather than taking either the hype or the skepticism at face value.
Meet the Two Models
Claude Fable 5 launched June 9, 2026 as Anthropic's first generally available Mythos class model, sitting above Opus in Anthropic's lineup and sharing its underlying architecture with the more restricted Claude Mythos 5. It's built specifically for long, ambitious, asynchronous agentic work, capable of running inside a harness like Claude Code for extended, multi-hour sessions. It's closed weight and accessible only through Anthropic's own API, cloud partners, and subscription plans.
Kimi K3 is Moonshot AI's newest flagship, a Mixture-of-Experts model with roughly 2.8 trillion total parameters, more than twice the size of DeepSeek's V4 Pro, making it the largest open-weight model released to date. It ships with a 1,048,576 token context window and native image understanding. It's available today through Moonshot's API and the Kimi app, with full open weights promised by July 27, 2026.
The Evidence: Benchmark by Benchmark
Moonshot's own released comparison table puts K3 against Fable 5 across more than 30 tests. A methodology note worth keeping in mind throughout: Moonshot ran K3 through Kimi Code, while Fable 5 was tested through Claude Code, meaning the two weren't evaluated under fully identical harness conditions. Where independent, harness-neutral data exists, it's flagged separately below.
| Benchmark | Kimi K3 | Claude Fable 5 | Who leads |
|---|---|---|---|
| SWE Marathon (long-horizon coding) | 42.0 | 35.0 | K3, by a clear margin |
| Program Bench | 77.8 | 76.8 | K3, narrowly |
| Terminal Bench 2.1 | 88.3 | 84.6 | K3, by 3.7 points |
| DeepSWE | 67.5 | 70.0 | Fable 5, narrowly |
| FrontierSWE | 81.2 | 86.6 | Fable 5, clearly |
| Kimi Code Bench 2.0 (Moonshot's own benchmark) | 72.9 | 76.9 | Fable 5, clearly |
| Automation Bench | 30.8 | 29.1 | K3, narrowly |
| SpreadsheetBench 2 | 34.8 | 34.7 | Essentially tied |
| BrowseComp (long-context search) | 91.2 | 88.0 | K3, narrowly |
| JobBench | 52.9 | 57.4 | Fable 5, clearly |
| GPU kernel optimization | Competitive | Baseline (with fallback) | Roughly matched |
| Frontend Code Arena (LMArena, blind preference) | #1, first in 6 of 7 categories | #2 overall, led in Gaming | K3 |
The pattern that emerges from Moonshot's own numbers is genuinely mixed, not a clean win for either side. K3 takes real, sometimes decisive leads on long-horizon coding endurance (SWE Marathon), terminal-based agentic work, and long-context search. Fable 5 holds clear, wider leads specifically on FrontierSWE, Moonshot's own internal coding benchmark, and JobBench, a broader professional task evaluation. That last detail matters: Fable 5 winning clearly on Kimi Code Bench, a benchmark Moonshot itself built, is a stronger signal than it might first appear, since it isn't a case of home-field advantage working in K3's favor.
The Independent Check: Artificial Analysis
This is the part of the claim that matters most, because it comes from a third party with no stake in either company's marketing. Artificial Analysis ran its own separate evaluation and published results within hours of K3's launch.
On the Intelligence Index v4.1, the industry's most cited composite benchmark, Fable 5 scored 59.9 (with Opus 4.8 fallback active) and K3 scored 57.1, a gap of 2.8 points. That's real, but it's a notably narrower gap than the raw parameter count difference might suggest, and it's close enough that Moonshot's "competitive" framing holds up better here than an outright victory claim would have. On Artificial Analysis's more granular measures, K3 posted a Coding Index score of 76.24 and an Agentic Index score of 50.07, while running at roughly 62 output tokens per second.
On AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reached an Elo of 1,547, a genuinely massive jump of 732 points over the previous Kimi K2.6 model, with only Fable 5 scoring higher. Artificial Analysis's own summary described K3 as well-rounded, with rubric scoring and analytical quality close to Fable 5's level, though it noted GPT-5.6 Sol still leads specifically on presentation quality in that evaluation, a category outside this particular comparison.
The number that complicates the story. On AA-Omniscience, K3's accuracy improved from 33 percent to 46 percent generation over generation, a real and meaningful gain. But its hallucination rate climbed at the same time, from 39 percent to 51 percent. In plain terms, K3 is answering more questions correctly and fabricating more answers, at the same time. That's not a detail Moonshot's own launch materials emphasized, and it's directly relevant if you're considering K3 for anything where factual reliability matters as much as task completion speed.
The Evidence That's Still Missing
A serious fact check has to name what it can't yet confirm. Three things are worth flagging clearly:
The weights aren't public yet. Moonshot has promised open weights by July 27, 2026, but as of K3's launch, no outside lab can independently load the model and run its own reproducible tests. Every benchmark discussed above, including Artificial Analysis's, was run against Moonshot's hosted API, not an independently verifiable local checkpoint.
The harnesses weren't consistent. Moonshot's own benchmark table used Kimi Code for K3 and Claude Code for Fable 5. Different agentic harnesses can meaningfully affect how a model performs on tool-use heavy benchmarks, so some portion of K3's apparent lead on tasks like SWE Marathon and Automation Bench could reflect harness advantage rather than pure model capability. This cuts both ways, it's also possible Fable 5 would score similarly or better under Kimi Code, but nobody has published that test yet.
Some claimed results weren't independently visible at launch. At least one detailed technical review noted that several of K3's cited leaderboard placements weren't yet showing up on the actual public boards being referenced, at the moment of launch. That's a normal lag in how these things get published, but it means some of the specific numbers in Moonshot's table hadn't been externally cross-checked at the time of release.
None of this means K3's results are fabricated. Artificial Analysis's independent testing broadly corroborates Moonshot's general positioning, K3 is genuinely close to Fable 5, just not proven equal or superior across the board yet.
Pricing: Where the Comparison Tilts Hard Toward K3
| Kimi K3 | Claude Fable 5 | |
|---|---|---|
| Input (per 1M tokens) | $3.00 | $10.00 |
| Output (per 1M tokens) | $15.00 | $50.00 |
| Context window | 1,048,576 tokens | 1,000,000+ tokens |
| Open weights | Promised July 27, 2026 | Closed, proprietary |
This is where the claim becomes genuinely compelling regardless of how the benchmark debate settles. K3 costs less than a third of Fable 5 on both input and output tokens, while landing within roughly three points of it on the composite Intelligence Index. If you're running production workloads at any meaningful volume, that price gap compounds fast, and it's arguably a bigger practical factor than a few benchmark points in either direction.
What This Means for Moonshot's Bigger Story
K3's pricing is also notable in isolation. At $15 per million output tokens, it's more expensive than Moonshot's own previous model K2.6 ($4) and pricier than competing Chinese models like GLM-5.2 ($4.40) and DeepSeek V4 ($0.87). Analysts have read this as a signal that Chinese labs are gaining enough confidence in model quality to price closer to US frontier labs rather than competing purely on discount pricing. Even at that higher price point, K3 still undercuts both Fable 5 and GPT-5.6 Sol substantially, so the pricing story is really about narrowing the gap from below, not matching Western pricing outright.
The Verdict
Moonshot's specific claim, that K3 performs competitively with Fable 5, holds up reasonably well against the independent evidence. It's not an exaggeration, but it's also not a victory. Fable 5 remains the stronger model on the composite intelligence measure and wins clearly on several individual benchmarks, including the one benchmark Moonshot itself built. K3 takes real, verified leads on long-horizon coding endurance and long-context search, and comes close enough everywhere else that the gap is now measured in single digits rather than tiers.
The more interesting story isn't really "does K3 beat Fable 5," it doesn't, consistently. It's that an open-weight model priced at less than a third of Fable 5's cost is now close enough to force a genuine cost-versus-capability decision for teams that would have defaulted to a proprietary frontier model without a second thought a year ago. Whether that decision favors K3 depends heavily on two things this comparison can't yet fully resolve: how K3 performs once its actual weights are public and independently testable, and whether its rising hallucination rate becomes a real problem in your specific use case or stays a footnote.
So Which One Should You Actually Use
Choose Claude Fable 5 if you need the strongest available model on complex, high-stakes engineering work, particularly anything resembling Moonshot's own coding benchmark or broad professional task evaluation, and cost isn't the primary constraint.
Choose Kimi K3 if you want frontier-adjacent performance at less than a third of the price, your workload leans toward long-horizon agentic tasks or long-context search where K3 independently tests well, and you're comfortable running a controlled pilot before committing to production, ideally waiting for or verifying against the July 27 open weights release.
Wait and watch if you're making an infrastructure-level decision that's hard to reverse. Give it two to three weeks past the weights release for independent labs to publish reproducible, harness-neutral benchmarks before treating K3 as a settled alternative to Fable 5.
