Anthropic, OpenAI Tie Atop AI Composite Index as Google Leads Preference Ranking
Claude Fable 5.1 and GPT-6 Astra each scored 53 on Intelligence Index v4.3, while GPT-6 Astra led several individual tests.

Anthropic and OpenAI shared the top score on a major composite AI index, while Google led a separate preference ranking, highlighting how benchmark design shapes the frontier race.
Claude Fable 5.1 and OpenAI’s GPT-6 Astra each scored 53 on Intelligence Index v4.3, the joint-highest result. Claude Opus 5 scored 51, Claude Fable 5 scored 50, Muse Spark 1.3 scored 48 and GPT-5.6 Sol scored 47.
Google led the ranking based on preferred answers, preventing any single model developer from claiming an overall lead across the evaluations. Nathan Benaich said, “The frontier is now a three-lab race between Anthropic, OpenAI, and Google.”
The ninth annual State of AI Report was published Oct. 8. Published annually since 2018, the review covers AI research, commercial deployment, capital allocation, geopolitics, policy and safety. This year’s edition also examines AI-assisted system development, robotics, falling inference costs and cybersecurity risks.
GPT-6 Astra led several individual tests. It scored 59.1% on Terminal-Bench 4.0, compared with 52.0% for Claude Fable 5.1 and 49.0% for Claude Opus 5. On AutomationBench-AA, GPT-6 Astra completed every objective without a guardrail violation in 41.6% of workflows, versus 32.1% for Claude Fable 5.1.
Version 4.3 replaced τ³-Banking with AutomationBench-AA and upgraded Terminal-Bench from version 2.1 to 4.0. Private questions or answers account for 45% of the index, up from 40% in the previous version. Agents make up 30% of the weighting, coding 20%, general tasks 30% and scientific reasoning 20%.
Claude led 26% of Anthropic’s measured model research-and-development work in August, compared with less than 1% in February. The measurement covers work conducted inside Anthropic under human supervision.
Across five benchmarks, the cost of reaching a fixed score fell 47% per quarter since 2023, equivalent to roughly a 13-fold annual decline. The report’s prior-year prediction scorecard recorded two correct predictions, five partial outcomes and three misses.
Because version 4.3 introduced new evaluations and changed the share of private-test-set data, its scores are not directly comparable with earlier index versions.