CRYPTONEWSFREE ← Back to Live Stream
Crypto Briefing • October 8th 2026, 5:41 PM

Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

Frontier AI Still Can't Automate Epoch's Work, Reports Show

Key Summary

Epoch AI's latest Automation Reports found that frontier AI models are not yet capable of fully automating tasks, even on clearly defined ones. While models performed well on defined subtasks, they struggled with creative and judgment-based work, indicating that human oversight is still necessary.

Please see our real time news feed on our Home Page

Introduction

Epoch AI spent much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch's own work?

The Benchmark

The new Epoch Automation Reports introduced on October 8, 2026, pull tasks from the organization's internal research. The evaluation covers 11 distinct tasks spread across five categories. Human graders score each model's output against the same internal quality rubrics Epoch uses for its own work.

The Results

Two models share the top spot. Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra each posted an average score of approximately 65% across the tasks. Grok 4.6 followed at 59%, while Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%. Gemini 3.8 Flash rounded out the listed results at 42%.

The Challenges

The models performed well on defined subtasks such as coding and computational analysis. However, they struggled with creative and judgment-based work, with recurring problems in style adherence and content targeting. Every model in the evaluation lagged on outputs that depend on implicit conventions – the house style, the expected framing, the sense of what belongs in a report and what does not.

What This Means

For companies deciding where to deploy AI, the findings draw a useful line between structured and open-ended work. Tasks with a clear right answer, like code or computation, look increasingly within reach for the top models. However, the benchmark reflects 11 tasks drawn from a single organization's research, so the results describe how well models handle Epoch's work specifically, not every knowledge job.
#Crypto#US#AI#Automation

Latest Related Headlines