Epoch AI 評估六款模型自動化其工作的能力
Can AI automate Epoch?
Epoch AI 以 11 項源自自身工作的任務評估六款模型,發現 Fable 5.1 和 GPT-6 Astra 在定義清晰的任務上較可靠,但仍難以處理開放式工作的判斷要求。開放權重模型表現落後,連封閉權重模型能可靠完成的明確任務也較難應付。每款模型每項任務只測試一次,評分亦包含主觀判斷;報告指出,結果不代表 AI 在所有工作上的能力。
We gave six models real work tasks from Epoch, such as graphic generation and research design, to test how close they are to automating our work. While frontier models such as Claude Fable 5.1 and GPT-6 Astra perform reliably on well-defined tasks, they fail in the more open-ended aspects that prevent full automation.
Epoch Automation Reports captures these shortcomings. We curate a diverse set of tasks directly from our work at Epoch, provide the models with the necessary context, use the highest available reasoning settings, and have them attempt the tasks without further intervention. We then manually grade the outputs against our own employee standards.
These evaluations aim to capture gaps in typical benchmarks, which concentrate on easily verifiable tasks that capture only a subset of real work. These existing benchmarks may also overstate progress if a model has been optimized to score well on them.
By combining the systematic approach of benchmarks with the holistic impressions that come from testing models on real work, we expose capability gaps that existing benchmarks miss and inform broader questions about AI adoption.
Key findings:
- Although Fable 5.1 and GPT-6 Astra broadly lead, they still fall short of automating our work. Open-weight models lag further behind, struggling even on the well-defined tasks that frontier models handle reliably.
- Models struggle to pick up Epoch’s standards, even with ample reference material. They miss implicit conventions such as our visual style and the kinds of topics our audience cares about, or produce more information-dense outputs than our typical work.
- Models can identify promising research directions, but lack judgment to successfully follow through. They struggle to design experiments that measure what they claim and treat flawed results as key findings.
- Although several tasks involve open-ended exploration with a large idea space, models often converge on similar ideas.
Methodology
Tasks
Our task suite is based on what people do within Epoch. As of launch, this includes 11 tasks across five diverse categories.
Graphic Design. Create an Epoch-style graphic that illustrates a given description.
Tasks:
- Create an Epoch-style diagram given a description and rough sketch of universal basic proposals.
- Create an Epoch-style graphic to illustrate how model progress may be understated with a fixed token budget.
- Create an Epoch-style conceptual diagram of the parity point: the point at which humans become more cost-effective than AIs.
Data Insight Generation. Create an Epoch-style Data Insight based on a provided direction.
Data Insights are short articles we regularly publish, which break down complex AI trends into focused snapshots.
Tasks:
- Create an Epoch-style Data Insight using Epoch’s July 2026 polling data.
- Create an Epoch-style Data Insight on AI usage in physics research with data from arXiv.
- Create an Epoch-style Data Insight using any data source.
Data Explorer Generation. Create an Epoch-style Data Explorer based on a provided direction.
Data Explorers present key datasets tracking historical and contemporary progress in AI. They are designed so that viewers can easily browse and understand the data.
Tasks:
- Create an Epoch Data Explorer on AI mentions in the Federal Reserve Board Beige Book over time, given the website repository and specific features to include.
- Create an Epoch Data Explorer on AI mentions in the Federal Reserve Board Beige Book over time, given the website repository.
AI Data Center Research. Research and analyze a given AI data center to add to Epoch’s AI Data Centers database, given a detailed guide.
Tasks:
- EcoDataCenter 1
- Tesla Cortex 1
Research Design. Propose a novel research project and run a pilot experiment to support further experimentation.
Task:
- Propose a research project and run a pilot experiment based on Epoch’s “9 big questions benchmarks can help answer”.
Setup
We test models on readily available systems on their highest reasoning settings:
| Model | Harness | Reasoning Effort |
|---|---|---|
| GPT-6 Astra | Codex | Ultra |
| Claude Fable 5.1 | Claude Code | Ultracode |
| Grok 4.6 | Grok Build | xhigh |
| Gemini 3.8 Flash | Antigravity | High |
| Kimi K3 | Kimi Code | Max |
| Qwen 3.8 Max | Qwen Code | Max |
For each task, we provide models with the relevant resources and full permissions within their own workspace and Google account. For example, for graphic design tasks, we provide models with a Figma integration and a Figma file with as much relevant information as possible (previous designs, reusable components, design guidelines). In general, we follow the principle of providing as much context to the model as we would to a new hire.
Evaluation
Each task category has a rubric that was developed with Epoch employees who have expertise in that task. We further developed the rubrics through an initial qualitative evaluation of these tasks, which identified common failure patterns and turned them into evaluation criteria.
Each rubric includes a mix of objective and subjective criteria. For example, the Data Insight Generation task includes:
- Objective: “Pre-chart and post-chart paragraphs each under 150 words”
- Subjective: “Title is clearly communicated”
Each output is manually reviewed and graded by a single human grader. We currently use this approach for grading expedience, and may add multiple graders in the future.
Our key findings come from qualitative observations, as we believe they best capture gaps in open-ended tasks. These observations come not only from the model outputs, but also from the trajectories, which help us understand what may have led to them. We use the rubric alongside these observations to compare models and track progress; however, these numbers alone do not capture the full picture and should be considered alongside the qualitative observations.
Fable 5.1 and GPT-6 Astra are broadly tied in the lead
Of the six models we evaluated, Fable 5.1 and GPT-6 Astra achieve the highest aggregate scores and lead on most individual task categories as well.

Their lead comes mostly from reliability on the well-defined parts of our tasks, such as coding and computational analysis, where their outputs were consistently accurate. Computer use was also far more capable than we expected. Earlier this year, frontier closed-weight models failed at tasks such as porting an article to Substack. When we tried this again with newer models, they succeeded.
We even removed a graphic design task from our suite since Fable 5.1 produced an Epoch-quality output. This design task was more straightforward than others: given a Matplotlib chart and our website repository of instructions and reusable chart components, evaluated models were asked to restyle the chart to match Epoch’s house style. We now largely automate this step in our own work.
Even so, these frontier closed-weight models fall short of fully automating our work due to gaps in the more open-ended parts. One common pattern was failing to pick up Epoch’s standards, despite being given ample reference material. These models consistently missed implicit conventions that humans would pick up, such as our visual style or what our audience cares about. Their research judgment was also weak: they struggled to design informative experiments and often presented results from flawed setups as key findings.
Open-weight models struggle where closed-weight models are reliable
Open-weight models lag further behind. They struggle not only on the open-ended parts of tasks, but also on the well-defined parts that closed-weight models handle reliably.
We can see this through one of our Data Insight Generation tasks, where models are asked to produce a data insight using arXiv data on physics papers. Kimi K3 based its entire insight on a data filtering error:

Kimi K3 output on “arXiv” task
Although AI mentions in physics papers have increased since ChatGPT, they have not increased as sharply as presented here. Kimi K3 computed these numbers across all arXiv papers, not just physics papers, and never checked whether its filter worked. None of the frontier closed-weight models made factual errors in their data insights.
These factual inaccuracies show up in other task outputs as well, such as Kimi K3’s wealth redistribution diagram where returns flow from people into capital assets:

Kimi K3 output on “Wealth Redistribution” task
The prompt describes people earning returns from capital assets, and the generated caption refers to “assets that pay a return”, but the visual does not follow this described flow. Despite Kimi K3 repeatedly working on this part of the diagram, its fixation on minor design details caused it to miss this larger communication error. In contrast, all frontier closed-weight model graphics were factually accurate.
This capability gap is larger than existing benchmark scores suggest. For example, Kimi K3 scores 158 on the Epoch Capabilities Index (ECI), which combines results from many AI benchmarks into a single measure of general capability. This is roughly tied with closed-weight models like Grok 4.6, despite Kimi K3 struggling on basic tasks that Grok 4.6 handles more reliably. Realistic open-ended tasks like these may reveal weaknesses that existing benchmarks miss.
Models fail to learn generalizable patterns from examples
Since we are testing for the ability to complete Epoch-quality work end-to-end, it is important for models to be able to pick up and understand Epoch’s standards. We provide references for task inputs to facilitate this process.
In the following task, we ask models to generate a graphic and provide sufficient reference material to convey Epoch’s simple house style and general design principles. We provide a link to our Gradient Updates page, which includes reference graphics, and a link to our Figma file, containing several months of design work, our color palette, design guidelines, templates, and reusable components.
I’m working on a gradient updates article.
Create a diagram to illustrate this: “Models have been improving over time in both performance and a fixed token budget but also their ability to productively use large token budgets. This means that constant-elicitation understates recent progress relative to max elicitation.“
Previous gradient updates articles can be found here: https://epoch.ai/gradient-updates
Here is the figma file we use for graphics generation (treat this as read-only): [Figma link]
Prioritize graphic design principles when creating your diagram. Your final output should be an epochified diagram.
Prompt for “Model Performance Chart” task

Screenshot from Figma file
The underlying idea is simple, so the diagram should be too. An Epoch designer given the task illustrated it with this simple chart.

Epoch designer output for “Model Performance Chart”
Fable 5.1 produced the following instead. It is far more information-heavy than our typical graphics and, to our eye, the additional information serves no explanatory purpose.

Fable 5.1 on “Model Performance Chart” task
Fable 5.1 borrows from a particularly complex reference example it found, shown below. This example uses the “A” and “B” badges to label the two concepts being compared, and they earn their place because each concept shows up in more than one part of the diagram. However, Fable 5.1 did not learn this standard and simply used these badges for one-off labeling.

Epoch diagram from Fable 5.1’s reference search
This failure to learn Epoch’s standards is also demonstrated by GPT-6 Astra’s attempt to choose an interesting topic that would appeal to our audience.
We asked the model to write an Epoch-style Data Insight, a short article on an AI trend. We gave the model an article focus (AI usage in physics research) and a link to all our published Data Insights, which show a clear pattern of appealing to a general audience interested in AI.
For comparison, we published a Data Insight on AI usage in math research shortly after running the models on this task. This Data Insight also uses arXiv paper data (from math) and focuses on a broad trend that is easy for a general audience to understand.

Actual Epoch “arXiv” Data Insight
However, GPT-6 Astra chose a topic that is too niche for such an audience:

GPT-6 Astra on “arXiv” task
This topic is specific to the physics domain rather than a general AI trend that Epoch’s audience would find interesting. Although the task involves physics arXiv data, GPT-6 Astra should have enough evidence in its references to determine a more fitting topic.
Models generate promising research directions, but struggle to design experiments that are actually informative
We task models with proposing a research project based on Epoch’s “9 big questions benchmarks can help answer”. This process involves choosing a novel research question, designing and running a pilot experiment, writing up the pilot results, and then proposing follow-up experiments.
The most important part of this process is experiment design. It’s not difficult to identify a valuable general direction for an experiment. For instance, an experiment that reveals exactly which jobs AI will automate next year would clearly be worth running. The challenge is in figuring out how to execute this experiment in a way that actually measures what we are looking for, which is what models struggle with.
For example, GPT-6 Astra chose the following research question, which could genuinely be a promising direction:
When an AI agent fails to improve with practice, has it failed to collect useful evidence, or failed to understand evidence it already has?
Its proposed experiment is also directionally good, as it involves realistic tasks that would help make the results more generalizable:

GPT-6 Astra on “Benchmarking (9 Questions)” task
However, it’s unclear how this experiment would be executed in practice. For example, the description does not specify what “matched externally selected experience” refers to or how “simulator-derived or restricted-policy reference bounds” would be used, even though both are critical parts of the experiment. A research proposal is only valuable if its experiments can be carried out and actually measure what they claim to.
Models present results driven by flaws in their own experiment setup as key findings
Alongside designing vague experiments, models also make mistakes setting up the experiments that they do run. Such mistakes are common in research, so we do not penalize models for them. What matters is how they handle the flawed results in informing next steps. We find that models sometimes recognize these mistakes, but treat the flawed results as key findings that drive the rest of the proposal.
For example, consider GPT-6 Astra’s research proposal. It asks whether AI agents fail to improve because they don’t collect useful evidence or use existing evidence. In the pilot experiment, it tested whether models could figure out the hidden rule behind a simple puzzle by trying a limited number of inputs and seeing their outputs. GPT-6 Astra compared how well they did when choosing their own inputs versus when given inputs designed to reveal the rule.
The setup flaw was that it gave models too low a token budget. The models Astra tested in its experiment could output up to 4096 tokens while choosing the input, which cut off 61 of 280 responses before they could name an input at all. These cut-off responses still counted as valid attempts, so the tested models had to guess the hidden rule with less information than they were supposed to get. GPT-6 Astra recognized this error and increased the token budget, which led to a significant improvement across the tested models. These results show that the 4096-token budget is insufficient to elicit a complete response, which is not relevant to the core research question.
In its write-up, GPT-6 Astra mentions that the low token budget lowered scores and that raising it brought them back up. However, it claims that these results reveal “sensitivity to the acquisition budget”, where acquisition budget refers to the token budget models get for choosing inputs:

GPT-6 Astra on “Benchmarking (9 Questions)” task
This framing is misleading because it implies models had a fair chance to produce a complete response and failed, when in reality they were unable to produce any answer. The performance differences therefore reflect incomplete responses rather than meaningful differences in reasoning effort during acquisition.
Fable 5.1 produces outputs with extremely high information density
In general, we want output to be presented in a simple and intuitive manner. For instance, a research report’s executive summary should be as concise as possible while preserving the main points. But models tend to introduce unnecessary complexity, especially Fable 5.1.
For example, consider the summary section of Fable 5.1’s research proposal and pilot experiment write-up.

Fable 5.1 on “Benchmarking (9 Questions)” task
This general pattern shows up not only in writing but also in other tasks, like graphic design. For example, Fable 5.1 leans towards more verbose captions, which contrasts with our core design principle of communicating clearly with as little detail as possible.

Fable 5.1 on “Wealth Redistribution” task
The reasoning for this is revealed in Fable 5.1’s rounds of self-critique when it pushes for more descriptive captions:
Descriptions are thin paraphrases: UBS ‘Everyone gets free public services’ drops ‘covering all basic needs, including food and housing’; UBC ‘Everyone gets capital that earns a return’ drops ‘their own’
Feedback from Fable 5.1 subagent
Models converge on similar ideas, despite a large idea space
On open-ended tasks with many possible directions, we may expect models to produce a diverse range of ideas. However, we find that models often converge on similar ideas.
For example, when we tasked models with creating a Data Insight using Epoch’s polling data, Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 all converged on the same topic:

Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 on “Polling” task
The idea space is somewhat more limited here, since models have to work within the polling dataset and avoid topics that have already been published. Yet there are still enough candidates that we did not expect three out of six models to pick the same one.
We also observe this with Kimi K3 and Qwen 3.8 Max on our research design task where they could generate any idea out of nine broad directions:

Kimi K3 and Qwen 3.8 Max on “Benchmarking (9 Questions)” task
Limitations
We acknowledge the following limitations of this benchmark:
Small sample size. We have a limited task suite, and we run each model only once per task. Given that manual grading is time-intensive, we prioritize in-depth qualitative analysis over more runs. Scores are therefore noisy and may shift with more runs. These results should not be read as a general measure of what AI can or cannot do across all work tasks.
Subjective grading. Many of our evaluation criteria involve subjective judgment, so the numerical scores should not be treated as ground truth. Specific numbers could change with different rubric item definitions or different graders. We include these numbers to give a rough sense of how models compare and how close each task is to being fully automated. We believe that the qualitative findings are more informative than these exact numbers.
Epoch-specific tasks. Our tasks are scoped to Epoch work, and thus are not representative of all work tasks. We also evaluate these model outputs by Epoch’s standards, which may not align with how other work tasks should be evaluated.
Different agent harnesses. We test each model on its own harness, such as GPT-6 Astra in Codex. Some observed differences may reflect the harness rather than the model itself. For instance, GPT-6 Astra in Codex with its built-in browser extension may help boost performance compared to Qwen 3.8 Max in Qwen Code with the Chrome DevTools MCP. We chose this setup to reflect how people actually use these models in their work, but we acknowledge it may introduce advantages for some models over others.
Conclusion
The frontier is evolving rapidly, but we still see gaps in the more open-ended parts of work. The best-performing closed-weight models are reliable on well-defined tasks, such as coding, data analysis, and computer use. However, they consistently fall short on the judgment that drives real work. This includes learning generalizable patterns from examples, designing good research experiments, and drawing meaningful conclusions from results.
In studying AI capabilities at open-ended tasks drawn from Epoch, we have two goals. First, we aim to provide a fuller view of AI capabilities than current benchmarks, which typically focus on well-defined, easy-to-grade tasks. Second, we want to shed light on how close AI is to automating work end-to-end. We find that it cannot yet replace workers, at least not at Epoch. We plan to add results as new models are released, informing questions about AI adoption and automation over time.
來源:Epoch AI:研究、數據與評測 · epoch.ai