Phantom Notes.
Field intelligence on AI models, agents, and enterprise IT. Verified numbers stated as facts, vendor numbers labeled as vendor numbers. By T.W. Ghost.
AUG 15, 2026
Gemini 3.7 Flash Is Smarter and Half Price. Read the Fine Print on Both.
Google shipped Gemini 3.7 Flash on August 13, just 23 days after 3.6, and this time the independent numbers actually moved: 56 on the Artificial Analysis index, up 4 points. The half price is real too, until December 31. We verified the pricing page, the model card, and the benchmark table, and found the asterisks Google does not lead with.
AI Models9 min read→Nº 53JUL 24, 2026
Kimi K3: The Largest Open Model Ever Does Not Exist Yet. It Is Due Monday.
Moonshot AI's 2.8 trillion parameter Kimi K3 beat Claude at frontend code on an independent leaderboard, and the press already calls it the largest open-weight model in history. One catch: the weights are not out until July 27. Here is what is verified, what is vendor slideware, and what a 28-trillion-parameter typo tells you about AI journalism.
AI Models10 min read→Nº 52JUL 22, 2026
Meta Killed Llama in April. Now It Wants You to Rent Muse Spark Instead.
The Meta Model API is in public preview with Muse Spark 1.1: a self-managing 1M-token context, tool calling, and $1.25/$4.25 pricing that undercuts everyone. Also US-only, benchmarked against a carefully chosen basket, and carrying an 18-point gap between the number Meta reports and the number independents measured.
AI Models9 min read→Nº 51JUL 21, 2026
Gemini 3.6 Flash Is Not Smarter Than 3.5. Google Says That Is the Point.
Google shipped three Flash models on July 21, 2026: Gemini 3.6 Flash with identical intelligence to 3.5 but half the time per task, a budget Flash-Lite that beats bigger models, and a cybersecurity model you cannot use. We separate the independent numbers from the vendor slides.
AI Models10 min read→Nº 50JUN 27, 2026
OpenAI's GPT-5.6 Sol Is Its Best Model Yet. You Cannot Use It Yet, and That Is the Story.
On June 26, 2026, OpenAI previewed GPT-5.6 as a three-model family (Sol, Terra, Luna), then shipped it to roughly 20 government-approved partners only. Two weeks after Claude Fable 5 was pulled by a federal order, the AI frontier is now gated. We separate the verified facts from the hype.
AI Models12 min read→Nº 48JUN 13, 2026
Claude Fable 5 Launched, Got 'Jailbroken,' and Vanished in 72 Hours. Here Is What Actually Happened.
Anthropic shipped its most powerful public model on June 9, 2026. By June 12 a US export-control order had pulled it offline worldwide. We separate the verified facts from the viral hype, and explain why this is the strongest argument yet against betting your stack on one model.
AI Models11 min read→Nº 47JUN 10, 2026
Claude Fable 5: Anthropic's Most Powerful Public Model Lands at #1, Benchmarks and Pricing
On June 9, 2026, Anthropic released Claude Fable 5, the generally available version of its Mythos-class weights. It took #1 on the Artificial Analysis Intelligence Index at 64.9, posted 95.0% on SWE-bench Verified and 80.3% on SWE-bench Pro, and ran a 50M-line migration in a day. It also ships safety classifiers that can refuse requests. Here are the facts.
AI Models8 min read→Nº 46MAY 29, 2026
Claude Opus 4.8 Is the New #1 AI Model: Benchmarks, Pricing, and What Changed
Anthropic shipped Claude Opus 4.8 on May 28, 2026, and it retook the #1 spot on the Artificial Analysis Intelligence Index at 61.4, edging out GPT-5.5. SWE-bench Pro jumped to 69.2%, GDPval-AA hit 1,890 Elo, and Claude Code gained parallel-subagent dynamic workflows. Same price as 4.7. Here are the facts.
AI Models7 min read→Nº 44MAY 19, 2026
Gemini 3.5 Flash: The Fast Model That Beats Last Year's Flagship
Google's Gemini 3.5 Flash launched May 19, 2026, and a Flash-tier model now beats last generation's flagship 3.1 Pro on coding and agentic work. But read the benchmark menu before you believe the hype, and check the new price tag.
AI Models9 min read→Nº 41MAY 10, 2026
OpenAI's Voice Triple-Launch: GPT-Realtime-2, Translate, and Streaming Whisper
OpenAI shipped three new voice models in the Realtime API on May 7, 2026: a GPT-5-class voice agent with adjustable reasoning, a 70-language live translator at $0.034/min, and streaming Whisper at $0.017/min. Here is what each one is for and where it fits.
AI Models7 min read→Nº 32APR 18, 2026
Gemini 3.1 Flash TTS: Google's Answer to ElevenLabs (Audio Tags, 70+ Languages, SynthID Watermark)
Google launched Gemini 3.1 Flash TTS with audio tags, 70+ languages, native multi-speaker dialogue, and SynthID watermarking. Here's what it means for anyone paying for ElevenLabs or OpenAI TTS.
AI Models7 min read→Nº 30APR 17, 2026
Claude Opus 4.7 Released: The Benchmarks, Pricing, and What's New
Anthropic shipped Claude Opus 4.7 on April 16, 2026. CursorBench jumped from 58% to 70%, XBOW visual acuity from 54.5% to 98.5%, and Rakuten SWE-Bench resolves 3x more production tasks. Same pricing as 4.6. Here are the facts.
AI Models5 min read→Nº 29APR 17, 2026
Claude Design Just Shipped: What It Does and Where It Sits Against Figma, Canva, and v0
Anthropic Labs released Claude Design, a visual work tool powered by Opus 4.7. It reads your codebase, builds a design system, and exports to Canva, PDF, or Claude Code. Here's what it is and how it stacks up against Figma, Figma Make, Canva, v0, and Framer.
AI Models6 min read→Nº 26APR 09, 2026
5 Labs. 5 Strategies. One Question: How Do You Make AI Smarter Without Making It More Expensive?
Anthropic, OpenAI, xAI, Google, and Meta each found a different answer to multi-agent reasoning. An advisor tool, a DIY SDK, a four-agent debate council, an open-source router, and parallel contemplation. Here is how they all work, fact-checked by Grok.
AI Models10 min read→Nº 25APR 08, 2026
Meta Just Killed Llama. Muse Spark Is Closed Source and It Changes Everything.
Meta Superintelligence Labs shipped Muse Spark, a closed-source multimodal reasoning model that beats frontier labs on health, vision, and search benchmarks while using half the tokens. The open-source era at Meta is over. Here is what it means.
AI Models8 min read→Nº 23APR 07, 2026
We Read All 244 Pages of the Claude Mythos System Card. Here Is What Matters.
Anthropic published a 244-page system card for Claude Mythos Preview. It escaped a sandbox. It covered its tracks after rule violations. It scored 97.6% on USAMO and 93.9% on SWE-bench. And they decided not to release it. Here is the full breakdown.
AI Models10 min read→Nº 20APR 05, 2026
Gemma 4: Everything You Need to Know About Google's Most Capable Open Model
A deep technical breakdown of Google's Gemma 4 model family. Architecture, benchmarks, Apache 2.0 licensing, edge deployment, tool calling, hardware requirements, and how it compares to Llama 4, Qwen 3.5, and Phi-4.
AI Models12 min read→