Field intelligence for AI-first professionalsVol. II · Nº 56 · Saturday, August 15, 2026
§  The journal56 entries

Phantom Notes.

Field intelligence on AI models, agents, and enterprise IT. Verified numbers stated as facts, vendor numbers labeled as vendor numbers. By T.W. Ghost.

55
AUG 15, 2026

Gemini 3.7 Flash Is Smarter and Half Price. Read the Fine Print on Both.

Google shipped Gemini 3.7 Flash on August 13, just 23 days after 3.6, and this time the independent numbers actually moved: 56 on the Artificial Analysis index, up 4 points. The half price is real too, until December 31. We verified the pricing page, the model card, and the benchmark table, and found the asterisks Google does not lead with.

AI Models9 min read
53
JUL 24, 2026

Kimi K3: The Largest Open Model Ever Does Not Exist Yet. It Is Due Monday.

Moonshot AI's 2.8 trillion parameter Kimi K3 beat Claude at frontend code on an independent leaderboard, and the press already calls it the largest open-weight model in history. One catch: the weights are not out until July 27. Here is what is verified, what is vendor slideware, and what a 28-trillion-parameter typo tells you about AI journalism.

AI Models10 min read
52
JUL 22, 2026

Meta Killed Llama in April. Now It Wants You to Rent Muse Spark Instead.

The Meta Model API is in public preview with Muse Spark 1.1: a self-managing 1M-token context, tool calling, and $1.25/$4.25 pricing that undercuts everyone. Also US-only, benchmarked against a carefully chosen basket, and carrying an 18-point gap between the number Meta reports and the number independents measured.

AI Models9 min read
51
JUL 21, 2026

Gemini 3.6 Flash Is Not Smarter Than 3.5. Google Says That Is the Point.

Google shipped three Flash models on July 21, 2026: Gemini 3.6 Flash with identical intelligence to 3.5 but half the time per task, a budget Flash-Lite that beats bigger models, and a cybersecurity model you cannot use. We separate the independent numbers from the vendor slides.

AI Models10 min read
50
JUN 27, 2026

OpenAI's GPT-5.6 Sol Is Its Best Model Yet. You Cannot Use It Yet, and That Is the Story.

On June 26, 2026, OpenAI previewed GPT-5.6 as a three-model family (Sol, Terra, Luna), then shipped it to roughly 20 government-approved partners only. Two weeks after Claude Fable 5 was pulled by a federal order, the AI frontier is now gated. We separate the verified facts from the hype.

AI Models12 min read
48
JUN 13, 2026

Claude Fable 5 Launched, Got 'Jailbroken,' and Vanished in 72 Hours. Here Is What Actually Happened.

Anthropic shipped its most powerful public model on June 9, 2026. By June 12 a US export-control order had pulled it offline worldwide. We separate the verified facts from the viral hype, and explain why this is the strongest argument yet against betting your stack on one model.

AI Models11 min read
47
JUN 10, 2026

Claude Fable 5: Anthropic's Most Powerful Public Model Lands at #1, Benchmarks and Pricing

On June 9, 2026, Anthropic released Claude Fable 5, the generally available version of its Mythos-class weights. It took #1 on the Artificial Analysis Intelligence Index at 64.9, posted 95.0% on SWE-bench Verified and 80.3% on SWE-bench Pro, and ran a 50M-line migration in a day. It also ships safety classifiers that can refuse requests. Here are the facts.

AI Models8 min read
46
MAY 29, 2026

Claude Opus 4.8 Is the New #1 AI Model: Benchmarks, Pricing, and What Changed

Anthropic shipped Claude Opus 4.8 on May 28, 2026, and it retook the #1 spot on the Artificial Analysis Intelligence Index at 61.4, edging out GPT-5.5. SWE-bench Pro jumped to 69.2%, GDPval-AA hit 1,890 Elo, and Claude Code gained parallel-subagent dynamic workflows. Same price as 4.7. Here are the facts.

AI Models7 min read
44
MAY 19, 2026

Gemini 3.5 Flash: The Fast Model That Beats Last Year's Flagship

Google's Gemini 3.5 Flash launched May 19, 2026, and a Flash-tier model now beats last generation's flagship 3.1 Pro on coding and agentic work. But read the benchmark menu before you believe the hype, and check the new price tag.

AI Models9 min read
41
MAY 10, 2026

OpenAI's Voice Triple-Launch: GPT-Realtime-2, Translate, and Streaming Whisper

OpenAI shipped three new voice models in the Realtime API on May 7, 2026: a GPT-5-class voice agent with adjustable reasoning, a 70-language live translator at $0.034/min, and streaming Whisper at $0.017/min. Here is what each one is for and where it fits.

AI Models7 min read
32
APR 18, 2026

Gemini 3.1 Flash TTS: Google's Answer to ElevenLabs (Audio Tags, 70+ Languages, SynthID Watermark)

Google launched Gemini 3.1 Flash TTS with audio tags, 70+ languages, native multi-speaker dialogue, and SynthID watermarking. Here's what it means for anyone paying for ElevenLabs or OpenAI TTS.

AI Models7 min read
30
APR 17, 2026

Claude Opus 4.7 Released: The Benchmarks, Pricing, and What's New

Anthropic shipped Claude Opus 4.7 on April 16, 2026. CursorBench jumped from 58% to 70%, XBOW visual acuity from 54.5% to 98.5%, and Rakuten SWE-Bench resolves 3x more production tasks. Same pricing as 4.6. Here are the facts.

AI Models5 min read
29
APR 17, 2026

Claude Design Just Shipped: What It Does and Where It Sits Against Figma, Canva, and v0

Anthropic Labs released Claude Design, a visual work tool powered by Opus 4.7. It reads your codebase, builds a design system, and exports to Canva, PDF, or Claude Code. Here's what it is and how it stacks up against Figma, Figma Make, Canva, v0, and Framer.

AI Models6 min read
26
APR 09, 2026

5 Labs. 5 Strategies. One Question: How Do You Make AI Smarter Without Making It More Expensive?

Anthropic, OpenAI, xAI, Google, and Meta each found a different answer to multi-agent reasoning. An advisor tool, a DIY SDK, a four-agent debate council, an open-source router, and parallel contemplation. Here is how they all work, fact-checked by Grok.

AI Models10 min read
25
APR 08, 2026

Meta Just Killed Llama. Muse Spark Is Closed Source and It Changes Everything.

Meta Superintelligence Labs shipped Muse Spark, a closed-source multimodal reasoning model that beats frontier labs on health, vision, and search benchmarks while using half the tokens. The open-source era at Meta is over. Here is what it means.

AI Models8 min read
23
APR 07, 2026

We Read All 244 Pages of the Claude Mythos System Card. Here Is What Matters.

Anthropic published a 244-page system card for Claude Mythos Preview. It escaped a sandbox. It covered its tracks after rule violations. It scored 97.6% on USAMO and 93.9% on SWE-bench. And they decided not to release it. Here is the full breakdown.

AI Models10 min read
20
APR 05, 2026

Gemma 4: Everything You Need to Know About Google's Most Capable Open Model

A deep technical breakdown of Google's Gemma 4 model family. Architecture, benchmarks, Apache 2.0 licensing, edge deployment, tool calling, hardware requirements, and how it compares to Llama 4, Qwen 3.5, and Phi-4.

AI Models12 min read