The right AI for your work is a matter of record.
Four labs, dozens of models, one that actually fits how you work. We keep the benchmarks current, read the fine print, and match you to it in three minutes.
Today’s standings, asterisks included.
Independent numbers where they exist, vendor numbers labeled where they don’t. This table changes when the facts do.
| Model | Intelligence¹ | Output $/1M | The verdict |
|---|---|---|---|
| Claude Opus 5 | 63#1 | 25.00 | Anthropic outranks its own halo model: the $25 flagship now tops the index. |
| Claude Fable 5 | 62#2 | 50.00 | Still the deep-coding default for the hardest long-horizon work. |
| GPT-5.6 Sol | 61#3 | — | Brilliant, and gated to ~20 approved partners. The API gets Terra and Luna. |
| Grok 4.6NEW | 61#3 | 6.00 | Ties Sol at a fraction of frontier pricing, about $0.84 per agent task. |
| Kimi K3 | 60#5 | 15.00 | The open-weights flag, one point off the frontier tie. |
| Gemini 3.7 FlashNEW | 56 | 3.75² | The agent workhorse finally got smarter, and half price. The sale ends December 31. |
¹ Artificial Analysis Intelligence Index, measured independently. ² Introductory rate through Dec 31, 2026; doubles to $7.50 on Jan 1, 2027.
Prices are list API rates per 1M output tokens. Dashes mean we won’t print a number we haven’t checked.
Field notes, filed under oath.
Every model launch, separated into what’s verified and what’s marketing. No press-release stenography.
Seedance 2.5 Makes 30-Second Movies With Sound. The 4K Claim Is Marketing.
ByteDance's new video model generates 30 seconds of picture and audio in one pass, takes 50 reference files, and edits existing footage by timestamp.
→Nº 55 · AUG 15Gemini 3.7 Flash Is Smarter and Half Price. Read the Fine Print on Both.
Google shipped Gemini 3.7 Flash on August 13, just 23 days after 3.6, and this time the independent numbers actually moved: 56 on the Artificial Analysis index, up 4 points.
→Nº 54 · AUG 15Grok Bot Wants Your Passwords. Grok 4.6 Might Deserve Them.
xAI shipped Grok Bot on August 11: always-on AI teammates that sign into your apps with your own credentials and work while you sleep.
→Nº 53 · JUL 24Kimi K3: The Largest Open Model Ever Does Not Exist Yet. It Is Due Monday.
Moonshot AI's 2.8 trillion parameter Kimi K3 beat Claude at frontend code on an independent leaderboard, and the press already calls it the largest open-weight model in history.
→Written by T.W. Ghost · corrections welcome, receipts required
Three tracks. Pick your altitude.
Ordered by how much you already know, not by how much we can sell you.
PRO TRACKS · 38 industry courses, from AI for Marketing to AI for Day Traders.