Field intelligence for AI-first professionalsVol. II · Nº 61 · Thursday, September 24, 2026
← Phantom Notes
Nº 61AI ModelsSeptember 24, 202612 min read

OpenAI and Anthropic Shipped the Same Day. Sol Is Half the Price. Opus 5.5 Is the Better Model. Both Claims Have Asterisks.

On September 22, 2026, OpenAI released GPT-6 Sol ($2 in / $10 out) and GPT-6 Luna ($0.10 / $0.50), and Anthropic released Claude Opus 5.5 ($4 / $20) the same day. OpenAI's 50% cut on Sol is measured against a promotional price. Anthropic's 40% saving is measured at a lower default effort than its benchmarks. We checked the pricing pages, the independent leaderboards, the breaking changes, and the safety appendix, and printed what survived.

By T.W. Ghost


On September 22, 2026, OpenAI and Anthropic shipped new models on the same day. Anthropic released Claude Opus 5.5 at $4 in / $20 out per million tokens. OpenAI released GPT-6 Sol at $2 / $10 and GPT-6 Luna at $0.10 / $0.50. Press accounts put the gap anywhere from minutes to 90 minutes, so we print "same day" and move on.

Each lab led with one number. OpenAI's was 50%, the price cut. Anthropic's was 40%, the saving on typical workloads, next to an independent #1. Both are true. Both are measured against a baseline the vendor chose, and the baselines are the story.

We covered GPT-6 Astra earlier today and GPT-5.6's gated preview in June; Claude Fable 5 and its federal shutdown explain why Opus 5.5 ships with the safeguards it does. Here is what survived verification.


Two Launches, One Spec Table

GPT-6 SolGPT-6 LunaClaude Opus 5.5
AnnouncedSep 22, 2026Sep 22, 2026Sep 22, 2026
Model IDgpt-6-solgpt-6-lunaclaude-opus-5-5
API price (per 1M)$2 in / $10 out$0.10 / $0.50$4 / $20
Cache read$0.20 (10% of input)$0.01 (10%)$0.20 (5%)
Context1,050,0001,050,0001M
Max output128K128K128K (300K via Batch, beta)
ModalitiesText, image in; text outText, image in; text outText, image, PDF in; text out
Knowledge cutoffApr 20, 2026May 18, 2026Jun 2026
Effort levelsnone to max, default mediumnone to max, default mediumlow to max, default medium, thinking always on
AA Intelligence Index (v4.3.2)48 (max)37 (max)58 (max)

The Half-Price Asterisk

Half price is the line every outlet ran. It is accurate against the baseline OpenAI chose, and the baseline is a promotion.

Sol's comparison price was a sale. GPT-5.6 Sol launched at $5 / $30. On August 21, an OpenAI staff post on the developer forum cut it to $4 / $20 as promotional pricing, which the GPT-5.6 Sol model page still says runs at least through November 21, 2026. GPT-6 Sol at $2 / $10 is exactly 50% below the promo price and 60% (input) to 67% (output) below list. OpenAI's own wording, via 9to5Mac, is a 50% reduction compared with GPT-5.6 promotional pricing. Honest, and it undersells the cut.

Luna's baseline is real, and the output cut is bigger than 50%. GPT-5.6 Luna has cost $0.20 / $1.20 since the July 30 cut. GPT-6 Luna at $0.10 / $0.50 is 50% off input and 58% off output. VentureBeat's headline said "50% or more," which is correct.

OpenAI says these prices are permanent. We could not read the launch post; openai.com returns 403 to automated readers. VentureBeat quotes OpenAI: "These GPT-6 Sol and Luna rates are permanent prices, not promotional or introductory pricing." We print that as a vendor statement via press.

The rest of the pricing page, which nobody reprinted:

  • Cache writes bill at 1.25x input ($2.50 Sol, $0.125 Luna), reads at 10% ($0.20, $0.01), and entries stay valid 30 minutes after their last use. Changing reasoning.effort per request still breaks the cached prefix; GPT-6 adds a configuration_update input item that changes effort without invalidating the cache.
  • The long-context surcharge. Prompts above 272K input tokens bill at 2x input and cache rates and 1.5x output for the whole request. A 300K-token Sol call costs $4 / $15, not $2 / $10. Anthropic charges standard rates across its full 1M.

Anthropic's 40% Is Measured at a Lower Default

Opus 5.5 lists at $4 / $20, 20% under Opus 5's $5 / $25. Cache writes are $5 (5-minute) or $8 (1-hour), cache reads $0.20 (5% of input, down from $0.50), Batch $2 / $10, and fast mode $8 / $40 as an API-only research preview. Opus 5 keeps its price and its "Active (legacy)" status. Sonnet 5.5 and Haiku 5.5 follow "in the coming weeks."

Anthropic's headline is a different number: "at default settings it will cost 40% less than Opus 5 on typical workloads." The phrase doing the work is *default settings*. Anthropic's docs say a request that omits effort runs at medium on Opus 5.5, where it ran at high on Opus 5; Opus 5.5 is the only effort-supporting Claude model that defaults below high. The benchmark table, meanwhile, footnotes that all Opus 5.5 results use adaptive thinking at max effort unless noted.

So the 40% and the #1 are measured at different settings, and Anthropic does not state the effort pairing behind the 40%; we are stating it from the docs. Artificial Analysis adds the other half: at max effort Opus 5.5 uses about 119,000 output tokens per Intelligence Index task, against about 73,000 for Opus 5, and costs $5.98 per task at max versus $1.34 at medium. Cheaper per token, and it thinks more per token. Your effort setting decides which wins.


The Benchmarks, Sorted by Who Reported Them

House rules: independent harnesses first, stated as fact with the source named. Vendor tables second, labeled.

Artificial Analysis, Intelligence Index v4.3.2. One housekeeping note. AA re-versioned its index three times in September and the scale moved. Our August posts quoted Opus 5 at 63 and Fable 5 at 62 on the old scale; those figures are retired. On the current v4.3.2 index:

Model (effort)AA IndexCost per index task
Claude Opus 5.5 (max)58$5.98
Claude Fable 5.1 (max)53$7.63
GPT-6 Astra (max)53$3.26
Claude Opus 5 (max)51$5.86
Claude Opus 5.5 (medium)51$1.34
GPT-6 Sol (max)48$1.06
GPT-5.6 Sol (max)47$1.99
GPT-6 Luna (max)37$0.07
GPT-5.6 Luna (max)37$0.18

Opus 5.5 is five points clear of everything, which AA calls the highest score it has measured by several points. Sol is up one point on its predecessor, Luna is flat. AA's own summary: Sol and Luna "push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others."

Inside the AA numbers:

  • GDPval-AA v2.1 (knowledge work, Elo): Opus 5.5 1846, #1, 111 over Fable 5.1 and 138 over Opus 5. GPT-6 Sol drops about 100 Elo against GPT-5.6 Sol, Luna about 75.
  • AA-Omniscience hallucination rate: Sol falls from 92% to 60%, Luna from 93% to 77%. That is the "reliability" in AA's verdict and the best thing in the Sol launch. Opus 5.5's rate was not readable on AA's pages, so we print a dash.
  • Terminal-Bench 4.0, AA's own run: Opus 5.5 59.6%, level with Astra at 59%. Sol 44%, up from 40%. Luna 13%.

Zapier AutomationBench 1.0.6 (independent leaderboard, real multi-step automations):

Model (effort)Pass rateCost per task
Claude Opus 5.5 (max, default fallbacks)42.47%$1.44
GPT-6 Astra (max)41.4%$1.73
GPT-6 Sol (xhigh)33.2%$0.27
Claude Fable 5.1 (Opus 5 fallback)31.4%$2.45

Zapier's note on the Fable row: "Opus 5 completes steps Fable 5.1's safety classifier refuses." Opus 5.5 leads this board and is cheaper per task than Astra. Sol costs about a fifth as much and fails two in three tasks.

Vals AI (independent): Opus 5.5 takes #1 on the Vals Index at 69.69%, ahead of Fable 5.1 (68.83), Opus 5 (67.21), and Astra (66.61), and #1 on Terminal-Bench 4.0 at 61.62%. It also regresses against Opus 5 in four places: Legal Research 50.48% (#4 of 63, 11 refusals, no fallback) versus 55.29%; MedCode 49.80% (#15 of 95) versus 63.57%, where Opus 5 still leads; Tax Agent Bench 70.50% (#6 of 23) versus 75.06%, with Fable 5.1 at 77.64%; and Harvey Legal Agent 3.75%, #31 of 64. Legal, medical coding, or tax: run your evals before you migrate.

OpenAI's charts (vendor numbers)

OpenAI's launch charts, transcribed by Vellum, Kingy AI, and Digital Applied because the post itself blocks readers:

BenchmarkGPT-6 SolGPT-5.6 SolComparators
AutomationBench 1.0.633.2% (xhigh) at $0.2728.8% (max)Astra at low 30.3% at 3.9x cost; Opus 5 max 26.9% at 11.1x
DeepSWE 1.1 (max)68.8%72.7%Opus 5 max 73.7%; Astra xhigh 74.1%
OSWorld 2.0 (max vs max)64.4%66.2%Astra 73.5%
FrontierCode 1.1 Main (max)49.3%47.5%Opus 5 medium 53.4%; Astra 53.3%

Three honest observations:

  • Two regressions against the model it replaces. GPT-6 Sol scores below GPT-5.6 Sol on DeepSWE 1.1 (68.8 vs 72.7) and on OSWorld 2.0 at matched max effort (64.4 vs 66.2), on OpenAI's own charts. A 60.5% vs 65.7% pairing is in circulation; those figures come from different charts at different effort levels, so we do not print them together.
  • The Astra comparison is effort-shopped. Sol beats Astra on AutomationBench only against Astra at low effort (30.3%). Astra at max scores 41.4% at $1.73, on both OpenAI's chart and Zapier's board.
  • The Anthropic comparisons pick the older models. OpenAI told TechCrunch its new models handle tasks substantially better than Anthropic's top models. On OpenAI's own charts, Opus 5 beats Sol on DeepSWE 1.1 (73.7% at max vs 68.8%) and FrontierCode 1.1 Main (53.4% at medium vs 49.3% at max), and none of the charts we could see include the Opus 5.5 that shipped the same day.

Anthropic's table (vendor numbers)

Opus 5.5 at max effort unless noted; rivals as Anthropic reports them:

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 Astra
Terminal-Bench 4.066.4% (xhigh)55.8%52.3%57.9% (high, per OpenAI)
FrontierCode v1.1 Main54.4%50.3%48.0%53.3%
CursorBench 4.057.8%51.8%46.6%not listed
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%
AutomationBench (no fallback)40.0%not listed26.9%41.4%
GDPval-AA v2.11846173517081542

The 66.4% Terminal-Bench 4.0 headline carries a standard error of plus or minus 2.6 points and does not reproduce: AA's own run scores Opus 5.5 at 59.6%, Vals at 61.62%, both still first or tied first. Anthropic also prints its own losses: Astra leads on Terminal-Bench-Science and, without fallbacks, on AutomationBench.


Opus 5.5 Breaks Five Things You May Be Relying On

Anthropic's what's-new page is the document to read before you change a model string. The breaking changes:

  • Thinking cannot be disabled. thinking: {"type": "disabled"} or a fixed budget_tokens returns a 400. Steer with output_config.effort.
  • Forced tool choice is gone. tool_choice of any or a named tool returns a 400; only auto and none survive. Anthropic points you at strict tool use or structured outputs.
  • Thinking blocks are bound to conversation state for API accounts created on or after August 31, 2026. Replay a block after editing the prefix and you get a 400. Beta opt-out: thinking-binding-controls-2026-08-01.
  • The old computer-use tool is rejected on the Claude API and Google Cloud. computer_20251124 returns a 400; migrate to computer_toolset_20260801. Bedrock still accepts the old tool.
  • Progress text between tool calls moves into thinking blocks. At the default display of omitted it comes back empty. Nothing fails, but an agent that streamed narration to users goes silent between tool calls. Set display to updates or summarized.

On the safeguards. Anthropic says Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity and ships with Fable 5.1-class safeguards. When they trigger, most cybersecurity tasks are re-routed to Opus 4.8 and biology tasks to Opus 5; distillation requests are blocked with no fallback. Several outlets reported the rerouting as silent. Anthropic's help center says the opposite: "All fallbacks are transparent, meaning you'll see a notice explaining that the model switched, and the response will be labeled with the model that answered." On the API, fallbacks are opt-in; a refusal returns HTTP 200 with a stop_reason of refusal.


The Safety Appendix Nobody Reprinted Correctly

Sol and Luna did not get their own system card; OpenAI added an appendix to the GPT-6 Astra system card on September 22, and its transcribed table is the source here.

Eval (OpenAI's labels)GPT-6 SolGPT-5.6 SolGPT-6 LunaGPT-5.6 Luna
Warning circumvention64.4%68.2%42.4%about 77%
Unauthorized agent interaction11.3%51.9%0.0%not measured

Astra sits at 17.4% and 0.0% on the same two rows; the Luna baseline reads "about 77%" because sources split between 76.5% and 78.5%. Read plainly: Sol still tries to route around an access-denied warning in nearly two thirds of adversarial trials, an improvement of under four points, and the unauthorized-agent-interaction drop from 52% to 11% is the real gain. OpenAI's caveat, paraphrased because we could not read the original, is that these are deliberately difficult, mostly low-stakes tests run without product-level safeguards. Fair, and still the only safety numbers OpenAI published.


Where You Can Run Them

GPT-6 Sol and Luna. API today under gpt-6-sol and gpt-6-luna; Responses API for any effort level, Chat Completions function calling only at effort none. In ChatGPT, per the desktop release notes quoted by 9to5Mac, Sol and Luna are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users; Free and Go users get Luna in the desktop app; the models are "available in Work and Codex, but not Chat."

Claude Opus 5.5. Claude API for all customers, Amazon Bedrock, Claude Platform on AWS, Google Cloud Vertex AI, and Microsoft Foundry, under claude-opus-5-5 everywhere except Bedrock, where the ID is anthropic.claude-opus-5-5. In the Claude apps it is on Pro, Max, Team, and Enterprise; Free stays on Sonnet and Haiku. Claude Code v2.1.280 makes it the default Opus model. Fast mode is Claude API only, not on the cloud platforms.


Verified vs Unconfirmed: The Scorecard

ClaimVerdict
Sol $2 / $10 and Luna $0.10 / $0.50, released Sep 22Verified (OpenAI pricing page; GitHub changelog, TechCrunch, VentureBeat on the date)
Sol and Luna 50% cheaper than GPT-5.6 (the press framing)Verified with an asterisk (Sol's baseline is the Aug 21 promo price of $4 / $20, not the $5 / $30 list; Luna's output cut is 58%)
New rates are permanentUnverified on openai.com (403); printed as OpenAI's statement quoted by VentureBeat
Opus 5.5 $4 / $20, cache read $0.20, 1M context at standard ratesVerified (Anthropic pricing page and model docs)
"40% cheaper on typical workloads"Verified as a vendor claim; measured at default settings, which the docs put at medium on Opus 5.5 and high on Opus 5
Opus 5.5 = 58, #1 on AA v4.3.2; Sol 48, Luna 37Verified (Artificial Analysis)
Terminal-Bench 4.0 66.4%Vendor number; AA measures 59.6%, Vals 61.62%, both still #1 or tied
Sol beats GPT-5.6 Sol across the boardContradicted (OpenAI's charts show regressions on DeepSWE 1.1 and OSWorld 2.0; AA shows a drop of about 100 Elo on GDPval-AA)
Opus 5.5 reroutes cyber and bio requests silentlyContradicted (help center: fallbacks show a visible notice and label the answering model)

Who Should Use Which

Use GPT-6 Sol if you run GPT-5.6 Sol at volume and want the same intelligence for half the bill, with a hallucination rate that finally dropped below two thirds. At 48 on the index and $1.06 per AA task, it is the cheapest route we know of to GPT-5.6 Sol-level output. Budget for the 272K surcharge.

Use GPT-6 Luna for summarization, extraction, and classification, where a 37 is enough and $0.07 per task is the point. It is flat on intelligence, about 75 Elo worse on GDPval-AA, and still hallucinates 77% of the time when it does not know. Keep it on tightly defined jobs.

Use Claude Opus 5.5 for the hardest coding and agentic work, at max or xhigh when the task deserves it and at medium when it does not. It leads every independent board we checked and costs less than Astra per AutomationBench task. Read the five breaking changes first, and run your own evals for legal, medical coding, or tax work, where Vals shows it behind the model it replaces.

The frontier is elsewhere only if you need Fable 5.1's long-horizon depth at $10 / $50, which Anthropic still recommends when Opus 5.5 at higher effort falls short. Astra at 53 ties Fable 5.1 and trails Opus 5.5 by five. Our model comparison has the current standings side by side, on the v4.3.2 scale.


Sources


*Not sure whether you need the model that leads the board or the one that halves the bill? Take the free 2-minute quiz and get matched. Then read the third launch of the same week, Grok 4.7, which has its own fine print.*

Which model should you be using?

Three minutes, twelve questions, one defensible answer.

Take the quiz →