Business workflows
AutomationBench 1.0.6
40.0% vs 33.2%Opus 5.5 max · GPT‑6 Sol xhigh
Merge rule: OpenAI’s chart supplies GPT‑6 Sol, Astra, Opus 5, Fable 5.1, and GPT‑5.6 Sol. Anthropic’s chart supplies Opus 5.5. Vendor-native task-cost estimates are preserved.
Computer use
OSWorld 2.0 · reported partial reward
Two source-native viewsPlaced together, not treated as one leaderboard
OpenAI post · offline set, v2026.08.08
Anthropic post · launch table
Do not read the 81.8% and 60.5% as a direct model gap. The posts do not establish matching task sets, harnesses, or effort/cost conditions for these two new-model figures.