2 min read
Add as a preferred source on Google

Epoch AI Benchmark Scores Frontier Models At Up To 65%

The 11-task evaluation found systems handled structured work better than assignments requiring subjective judgment, audience awareness and research design.

Blank task cards arranged beside a satellite image / TokenPost.ai
Blank task cards arranged beside a satellite image / TokenPost.ai

Epoch AI’s open-ended work benchmark gave Claude Fable 5.1 and GPT-6 Astra the highest average scores at 65%, while showing that frontier systems remain less reliable on tasks requiring judgment and research design.

The evaluation, published Oct. 8, covered 11 assignments across graphic design, data insight generation, data explorer generation, AI data-center research and research design. A score of 100% represented work that met the internal standard expected from an Epoch employee.

Grok 4.6 scored 59%, followed by Qwen 3.8 Max at 53%, Kimi K3 at 52% and Gemini 3.8 Flash at 42%.

The systems generally performed better on well-defined work, including coding and computational analysis. They had more difficulty inferring implicit standards, such as matching Epoch’s visual style and determining which subjects its audience would find relevant.

The research-design portion required each system to propose a project, conduct a pilot experiment, document the results and recommend follow-up work. The models often identified promising research directions but struggled to create experiments that effectively measured the questions they posed.

One GPT-6 Astra pilot illustrates the problem. After the evaluation imposed a 4,096-token output limit, the system treated 61 of 280 truncated responses as valid attempts.

The systems were tested at their highest available reasoning settings with corresponding agent harnesses. Each output was manually scored by one human grader using internal rubrics.

The benchmark used tools and reference materials including Figma, Google Drive, satellite imagery and internal-style examples. It was designed to measure end-to-end work that requires gathering context and applying subjective judgment, rather than tasks with easily verifiable answers.

One graphic-design task was removed after Fable 5.1 produced an Epoch-quality result. The assignment was considered relatively straightforward and is now largely automated internally.

The findings apply to Epoch’s own standards and work. They indicate that the tested systems are closer to automating structured subtasks than open-ended assignments requiring audience awareness, taste, interpretation and experimental judgment.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…