← All posts
Professional insight·July 19, 2026·6 min read

The July 2026 Frontier Scorecard: Fable 5 vs GPT-5.6 Sol vs Kimi K3

ModelsBuyer's guide

Five weeks, three frontier releases. If you just want the verdict without the benchmark theatre, it fits in a single table.

Claude Fable 5 GPT-5.6 Sol Kimi K3
Maker Anthropic OpenAI Moonshot AI
Price (in/out, per M tokens) $10 / $50 $5 / $30 $3 / $15
Overall index (AA) leader on GDPval, Briefcase 59 57
Long-horizon agents Best (Briefcase 1,587) 3rd (1,495) 2nd (1,527)
Frontend code 2nd n/a #1 (Arena 1,679)
Hard reasoning (GPQA etc.) strong Best good
Open weights No No Yes
Context long long 1M tokens

The table undersells how lopsided the pricing has got. On cost per million output tokens, the number that scales with every workflow you run, the whole story is in the gap.

Cost per 1M output tokensLower is better
1 Claude Fable 5 $50 2 GPT-5.6 Sol $30 3 Kimi K3 $15
Chart: PushBox. Data: published API pricing (OpenRouter), July 2026.

How to actually choose

Choose Fable 5 when the agent runs long and unattended. It holds the top of the long-horizon agentic benchmarks, and in our experience that's exactly where cheaper models quietly come apart. You pay 3x for the privilege. On a workflow that touches your invoices, that's money well spent.

Choose Sol when the job is hard reasoning in a short window. Maths, science, the genuinely nasty debugging. It wins six of nine head-to-head benchmarks against K3, and OpenAI's tiering (Sol, Terra, Luna) lets you drop down to cheaper tiers for the lighter stuff.

Choose K3 when volume is the problem. Browsing, tool calls, structured extraction, frontend code. At $3/$15 it's the first open-weight model you can defend in a procurement meeting on capability rather than just price, and the clearest sign yet of how Chinese labs are competing on less compute. We go deeper in our Kimi K3 breakdown.

The uncomfortable truth

On most of these benchmarks the gap between first and third is smaller than the gap between a good prompt and a bad one. Model choice matters far less than workflow design, the retries, the validation, the human checkpoints in the right places. We've watched a mid-tier model with a well-built harness walk all over a frontier model that someone bolted onto a form.

Which is the polite version of this. The model is about 20% of the outcome. The build is the other 80%. Pick a row from the table, then put your energy where it actually counts.

AI, Pushed to Work.

Want this kind of thinking applied to your business? The audit takes 30 minutes.

Book a 30-min audit

Keep reading