The July 2026 Frontier Scorecard: Fable 5 vs GPT-5.6 Sol vs Kimi K3
Five weeks, three frontier releases. If you just want the verdict without the benchmark theatre, it fits in a single table.
| Claude Fable 5 | GPT-5.6 Sol | Kimi K3 | |
|---|---|---|---|
| Maker | Anthropic | OpenAI | Moonshot AI |
| Price (in/out, per M tokens) | $10 / $50 | $5 / $30 | $3 / $15 |
| Overall index (AA) | leader on GDPval, Briefcase | 59 | 57 |
| Long-horizon agents | Best (Briefcase 1,587) | 3rd (1,495) | 2nd (1,527) |
| Frontend code | 2nd | n/a | #1 (Arena 1,679) |
| Hard reasoning (GPQA etc.) | strong | Best | good |
| Open weights | No | No | Yes |
| Context | long | long | 1M tokens |
The table undersells how lopsided the pricing has got. On cost per million output tokens, the number that scales with every workflow you run, the whole story is in the gap.
How to actually choose
Choose Fable 5 when the agent runs long and unattended. It holds the top of the long-horizon agentic benchmarks, and in our experience that's exactly where cheaper models quietly come apart. You pay 3x for the privilege. On a workflow that touches your invoices, that's money well spent.
Choose Sol when the job is hard reasoning in a short window. Maths, science, the genuinely nasty debugging. It wins six of nine head-to-head benchmarks against K3, and OpenAI's tiering (Sol, Terra, Luna) lets you drop down to cheaper tiers for the lighter stuff.
Choose K3 when volume is the problem. Browsing, tool calls, structured extraction, frontend code. At $3/$15 it's the first open-weight model you can defend in a procurement meeting on capability rather than just price, and the clearest sign yet of how Chinese labs are competing on less compute. We go deeper in our Kimi K3 breakdown.
The uncomfortable truth
On most of these benchmarks the gap between first and third is smaller than the gap between a good prompt and a bad one. Model choice matters far less than workflow design, the retries, the validation, the human checkpoints in the right places. We've watched a mid-tier model with a well-built harness walk all over a frontier model that someone bolted onto a form.
Which is the polite version of this. The model is about 20% of the outcome. The build is the other 80%. Pick a row from the table, then put your energy where it actually counts.
AI, Pushed to Work.
Want this kind of thinking applied to your business? The audit takes 30 minutes.
Book a 30-min audit