Field intelligence for AI-first professionalsVol. II · Nº 56 · Saturday, August 15, 2026
§  The comparison

Four labs, on the record.

ChatGPT, Claude, Gemini, and Grok, scored across nine dimensions, with the honest caveats the marketing pages leave out. The table is the record. The quiz is the verdict for your specific work.

§  The scorecard
DimensionChatGPTClaudeGeminiGrok
Reasoning9.51099.5
Code generation8.51098.5
Writing quality8107.57
Speed869.510
Multimodal88108
Context window1010109
Ecosystem1010107.5
Free tier78107
Privacy61065

Editorial ratings on a 10-point scale, revised at every major launch. Green marks the leader per row. Ties share the lead.

01

ChatGPT

OpenAI (ChatGPT)

Deepest ecosystem; its best model is gated, its budget model is nearly free

OpenAI's ChatGPT remains the most widely-used AI assistant in the world. On the current Artificial Analysis snapshot (August 15, 2026), OpenAI's best model, GPT-5.6 Sol, scores 61, tied with Grok 4.6 and behind Claude Opus 5 (63) and Claude Fable 5 (62); Sol remains gated to roughly 20 approved partners, while the API gets GPT-5.6 Terra ($2/$12 per M) and the budget GPT-5.6 Luna, cut 80% to $0.20/$1.20 on July 30, 2026. Terra is no consolation prize: on Google's own Gemini 3.7 Flash model card it leads the agentic evals, topping DeepSWE v1.1 (69.6%), Terminal-bench, and OSWorld-2.0. It dominates shell automation (Terminal-Bench 2.0: 82.7%, +13 over Opus 4.7) and advanced math (FrontierMath Tier 4: 35.4%). API pricing doubled vs 5.4 to $5/$30 per million tokens; a new GPT-5.5 Pro variant for longer reasoning sits at $30/$180. Codex now has a 1M-token context window with optional fast-mode at 2.5x cost. Honest caveats: GPT-5.5 still loses SWE-Bench Pro to Claude Opus 4.8 (58.6% vs 69.2%), loses MCP-Atlas tool use to both Opus (82.2%) and Gemini (83.6%), and posts an 86% hallucination rate on AA-Omniscience. The ecosystem advantage remains unmatched: DALL-E, Codex, Atlas browser, 60+ connectors, Memory, Projects, GPT Store, and Microsoft 365 Copilot integration. Sora video app/API is being discontinued (web/app April 26, 2026; API September 24, 2026).

Strengths
  • GPT-5.6 Sol at 61 on the AA Intelligence Index, tied with Grok 4.6 (but gated to ~20 partners)
  • GPT-5.6 Terra leads DeepSWE (69.6%), Terminal-bench, and OSWorld on Google's own comparison table
  • 1M-token context window now standard in Codex (not Pro-only)
  • Broadest ecosystem: DALL-E, Codex, Atlas browser, 60+ connectors
  • Microsoft 365 Copilot integration and GPT Store distribution
Best for

Shell automation, advanced math and research, broadest ecosystem, agentic task completion across multiple tools

Ecosystem

Microsoft 365 Copilot, GPT Store, Codex, Atlas browser, 60+ app connectors

Pricing

Free (with ads in US); Go $8/mo; Plus $20/mo; Pro $100/mo or $200/mo; Business $25/user. API: GPT-5.6 Terra $2/$12, GPT-5.6 Luna $0.20/$1.20, GPT-5.5 $5/$30, GPT-5.5 Pro $30/$180 per M tokens

Learn ChatGPTVisit ChatGPT
02

Claude

Claude (Anthropic)

The top two slots on the Intelligence Index, strongest coder

Anthropic's Claude Fable 5 (June 9, 2026) is the most capable model the company has ever made generally available, and Anthropic now holds the top two slots on the current Artificial Analysis snapshot (August 15, 2026): Claude Opus 5 at 63 and Fable 5 at 62, both measured at max reasoning, ahead of GPT-5.6 Sol and Grok 4.6 at 61 each. It is the production, safeguarded version of the same weights as the restricted Claude Mythos 5. On coding it posts 95.0% on SWE-bench Verified and 80.3% on SWE-bench Pro, beating Opus 4.8 (69.2%), GPT-5.5 (58.6%), and Gemini (54.2%), and on GDPval-AA, the benchmark for real economic-value work, Anthropic still leads with Opus 5 on top (Grok 4.6, at 1753 Elo, is second). It is built for long-horizon autonomy, working for days at a time in an agent harness and testing its own output, plus state-of-the-art vision for diagrams, charts, and tables inside PDFs. New wrinkle for developers: Fable 5 ships safety classifiers that can decline a request (returned as stop_reason "refusal", with server, client, or manual fallback to Opus 4.8), reroute in under 5% of sessions, and require 30-day data retention. Fable 5 prices at $10/$50 per million tokens with a 1M-token context and 128K output. Sitting alongside it, the cheaper, faster Claude Opus 5 ($5/$25) is the everyday flagship, and on the current index it actually edges Fable 5 by a point at half the price. Honest caveats: Fable runs slower per turn, costs double Opus 5, its classifiers have refused some innocuous prompts near security and biology topics, and on AA's Briefcase agent eval Grok 4.6 completes tasks in about half the turns and a quarter of the input tokens of Opus 5 at max reasoning.

Strengths
  • Holds #1 and #2 on the AA Intelligence Index: Opus 5 (63) and Fable 5 (62)
  • Best production coding: 95.0% SWE-bench Verified, 80.3% SWE-bench Pro (11+ points clear of the field)
  • Leads real economic-value work: Opus 5 tops GDPval-AA, with Grok 4.6 (1753 Elo) second
  • Built for multi-day autonomy: ran a 50M-line codebase migration in a day in early testing
  • State-of-the-art vision for diagrams, charts, and tables nested in files and PDFs
  • Two-tier lineup: Fable 5 for the hardest work, Opus 5 as the cheaper, faster default
  • 1M-token context, parallel-subagent workflows in Claude Code, 1,000+ Agent Skills
Best for

Long-horizon agentic coding, multi-day autonomous projects, hard knowledge work, large codebases, document-heavy research, and visual design

Ecosystem

Claude Code, Cowork, Design, Routines, Agent Skills (1,000+), Agent SDK, Projects, MCP, Claude Platform on AWS, Amazon Bedrock, Vertex AI, Microsoft Foundry

Pricing

Free tier; Pro $17-20/mo; Max from $100/mo (5x) up to $200/mo (20x); Team $20-125/seat; Enterprise $20/seat + usage. API: Fable 5 $10/$50, Opus 5 $5/$25 per M tokens

Learn ClaudeVisit Claude
03

Gemini

Gemini (Google)

The agent workhorse: smarter, faster, half price until December 31

Google's Gemini 3.7 Flash (August 13, 2026) is the current headline model: a Flash-tier workhorse for coding and agents that, unlike July's 3.6 release, actually got smarter. Artificial Analysis measured 56 on its Intelligence Index at high thinking, up 4 points from 3.6 Flash, at 340.1 tokens/sec output speed and a $0.58 blended price per million tokens. Google reports big agentic gains over 3.6 Flash: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 48.6%, Terminal-bench 2.1 85.8%, OSWorld-2.0 47.9%, and WebDev Arena 1588 Elo. Launch pricing is $0.75/$3.75 per million tokens, but read the fine print: that is an introductory rate through December 31, 2026, doubling to $1.50/$7.50 on January 1, 2027, and 3.6 Flash was repriced identically the same day. 1M-token context, 64K output, multimodal input (text, image, video, audio, PDF), and three thinking levels (low, medium, high; the old minimal level now returns an API error). Honest caveats: time to first token is 9.83 seconds at high thinking, painful for live chat; thinking tokens bill at the output rate; and on Google's own model-card comparison GPT-5.6 Terra still leads the hardest agentic evals. Deep integration across Google Workspace, NotebookLM, Antigravity, and Gemini Spark, Google's 24/7 personal agent for AI Pro/Ultra subscribers (not available in the EEA, UK, or Switzerland).

Strengths
  • Independently measured gain: AA Intelligence Index 56 vs 52 for 3.6 Flash
  • AA-measured 340.1 tokens/sec output at a $0.58 blended price per M tokens
  • Google reports DeepSWE 65.3%, Terminal-bench 2.1 85.8%, FrontierCode 43.6%, all big jumps over 3.6
  • Native multimodal input: text, image, video, audio, PDF, with 1M-token context
  • Intro pricing $0.75/$3.75 per M through Dec 31, 2026, cheaper than Claude Haiku 4.5 on both sides
Best for

Agentic and coding workloads at scale, multimodal tasks, Google Workspace integration, high-volume document processing

Ecosystem

Google Docs, Sheets, Gmail, Vids, NotebookLM, Canvas, Gems, Veo 3.1, Imagen, Antigravity, Gemini Spark, Enterprise Agent Platform

Pricing

Free tier; AI Pro $19.99/mo; AI Ultra from $99.99/mo, top tier $199.99/mo. API: Gemini 3.7 Flash $0.75/$3.75 per M intro through Dec 31, 2026, then $1.50/$7.50; Gemini 3.1 Pro $2/$12 under 200K context

Learn GeminiVisit Gemini
04

Grok

Grok (xAI)

Frontier intelligence at $2/$6, the best per-task economics measured

xAI's Grok 4.6 (August 12, 2026), released under the SpaceXAI banner (xAI merged into SpaceX in February 2026, was rebranded SpaceXAI in July, and the $60B Cursor acquisition closed August 14), put the company on the intelligence frontier for the first time: Artificial Analysis independently scores it 61, tied with GPT-5.6 Sol and one point behind Claude Fable 5 (62), with only Claude Opus 5 (63) clearly ahead. The sharper story is agent economics: AA measures roughly $0.84 per task on its Briefcase agent eval, completing tasks in about half the turns and a quarter of the input tokens of Opus 5 at max reasoning, Kimi K3 cost at slightly higher measured intelligence. API pricing holds at $2/$6 per million tokens (cached input $0.50, rates double past 200K input) with a 500K context and a fast variant at twice the price. It shipped one day after Grok Bot, xAI's always-on AI teammates that work on a persistent cloud computer and sign into your apps with your own credentials, bundled with SuperGrok Heavy and Cursor Ultra/Teams Premium. Honest caveats: Terminal-Bench v3.0 is a weak 26% on xAI's own launch numbers; the 500K context is half of Grok 4.3's 1M; Grok Bot's security model shares one computer and every login across all your Bots (xAI's docs warn "Do not use separate Bots as a security boundary"); and the reported 1.5T parameter count is secondary reporting, never published by xAI. Grok 4.3 stays in the lineup as the budget 1M-context option at $1.25/$2.50.

Strengths
  • Frontier tie: 61 on the AA Intelligence Index, level with GPT-5.6 Sol, one point behind Claude Fable 5
  • Best measured agent economics: ~$0.84 per AA-Briefcase task, half the turns of Opus 5 at max reasoning
  • $2/$6 per M tokens with $0.50 cached input, a fraction of frontier list prices
  • Real-time X/Twitter data, default model in Grok Build, day-one default in Cursor
  • Grok Bot: always-on agent teammates on a persistent cloud computer (SuperGrok Heavy and Cursor plans)
Best for

Cost-efficient agentic work at frontier intelligence, real-time information, long-running autonomous jobs, high-volume tool-calling

Ecosystem

X/Twitter integration, Cursor (SpaceX-owned since Aug 14, 2026), Grok Build, Grok Bot, DeepSearch, Voice API, Real-time Search API, Grok Imagine video, xAI API (OpenAI/Anthropic SDK compatible), OpenRouter, part of SpaceXAI

Pricing

Free tier; SuperGrok $30/mo; Grok Business $30/seat; Heavy $300/mo list ($99/mo six-month promo since May); Enterprise custom. API: grok-4.6 $2/$6 per M tokens (cached input $0.50, rates double past 200K input); grok-4.3 $1.25/$2.50 with 1M context. Grok Bot bundled with Heavy, Cursor Ultra $200/mo, Cursor Teams Premium $120/seat (xAI's figure)

Learn GrokVisit Grok

The record shows four answers. Yours takes three minutes.

Take the quiz