Launch-day benchmark merge · Sep 22, 2026

Opus 5.5 ×
GPT‑6 Sol

The three benchmarks shared by both launch posts—re-charted so the two new models finally appear together.

40.0%Opus 5.5 · AutomationBench maxvs GPT‑6 Sol’s best reported 33.2% at xhigh.
54.6%Opus 5.5 · FrontierCode mediumThe highest new-model point; GPT‑6 Sol peaks at 49.3%.
Not cleanly comparableOSWorld 2.0Both posts use partial reward, but the reported sets and conditions differ.

Merged cost curves

Each point is an effort setting, from low through max. Cost is logarithmic. Click a legend item to isolate the curves you care about; hover or focus points for exact values.

Business workflows

AutomationBench 1.0.6

40.0% vs 33.2%Opus 5.5 max · GPT‑6 Sol xhigh

Merge rule: OpenAI’s chart supplies GPT‑6 Sol, Astra, Opus 5, Fable 5.1, and GPT‑5.6 Sol. Anthropic’s chart supplies Opus 5.5. Vendor-native task-cost estimates are preserved.

Agentic coding

FrontierCode 1.1 Main

54.6% vs 49.3%Opus 5.5 medium · GPT‑6 Sol max

One wrinkle: Opus 5.5 peaks at medium effort (54.64%) rather than max. The chart keeps every reported effort point rather than smoothing that away.

Computer use

OSWorld 2.0 · reported partial reward

Two source-native viewsPlaced together, not treated as one leaderboard

OpenAI post · offline set, v2026.08.08

Anthropic post · launch table

Do not read the 81.8% and 60.5% as a direct model gap. The posts do not establish matching task sets, harnesses, or effort/cost conditions for these two new-model figures.

Launch posts

The exact chart points above are transcribed from each publisher’s embedded launch-page data, not estimated from pixels.

Method: for the two compatible cost curves, the existing OpenAI series are used as the base and the Opus 5.5 series is added from Anthropic’s equivalent chart. Overlapping older-model series were not averaged. OSWorld is separated by source because the launch posts do not document a common evaluation setup for the two new-model scores.