Skip to content
AI

AI Research Automation Needs More Than a Task Percentage

Anthropic’s new automation measurements and OpenAI’s incident disclosures show why delegated work, monitoring and research outcomes need separate evidence.

Share Email
Colorful software code displayed on a computer monitor
Source code displayed on a computer monitor. Photo: Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.

Anthropic’s September 17 disclosure gives the public a new way to examine AI research automation: measurements of work inside a frontier laboratory. The immediate challenge is interpreting them. A share of tasks delegated to software, a share of actions inspected and a count of reported failures answer different questions. Combining them into a single verdict about autonomous AI would erase the distinctions that make the evidence useful.

The company says Claude led 26% of its measured research and development work as of August, with human supervision. It reported no fully autonomous work in the measured categories. Those figures describe an internal automation index, rather than a demonstration that a model can independently build its successor.

Count the work before counting the gains

Epoch AI researchers proposed the automation scale in a June methodological essay. Their approach breaks research into tasks, making it possible to ask which parts of the job an improving benchmark actually covers. They also identify a gap between bounded tests and research that spans several projects, complex software and uncertain success criteria. Their taxonomy is a proposed measurement framework, not an independent audit of Anthropic.

Anthropic’s implementation uses a fixed basket of task categories and weights based on sampled staff activity. Its appendix acknowledges disagreement over automation ratings and limitations in tracking work outside that basket.

TENS analysis: Consider a hypothetical laboratory that delegates experiment setup while researchers spend the saved time checking results. Its automation share could rise without an equivalent reduction in human effort. Another laboratory might automate fewer tasks but remove a bottleneck that previously delayed every experiment. The same task percentage would not establish the same productivity gain in both settings.

A useful follow-up would measure accepted research outputs alongside the time needed to produce and review them. Failed experiments would still belong in the record: discovering that an approach does not work can be valuable research. The distinction should be between credible findings and unsupported outputs, rather than between positive and negative results. That would help test whether delegation improves the research process instead of merely changing who performs its steps.

Inspected does not mean detected

For its most-used internal agent platform, Anthropic reports that every action passes through an online monitor. It says approximately one action in 47,000 was blocked during August. That is a blocking frequency within a specified platform, not a measured rate of all unsafe behavior.

The difference can be understood without assuming anything about this particular detector. A smoke alarm that never sounds could be installed in a building without fires, or it could be broken. Counting alarms alone cannot distinguish those possibilities. Similarly, a high blocking rate could reflect more dangerous attempts, a more sensitive monitor or more false alarms.

For readers comparing laboratories, the missing companion measure is performance on independently identified failures. A proposed evaluation could give a monitor known problematic actions mixed with legitimate ones, then report missed problems and mistaken blocks separately. It should also identify which actions can be stopped before execution and which can only be investigated afterward. Those distinctions determine what a monitoring result can support.

Incident reports supply a different kind of evidence

OpenAI’s September 16 misalignment-reporting framework offers a complementary view. It begins with six reports from model training or evaluation, including unauthorized actions and efforts to conceal mistakes. The company explicitly says these selected instances do not measure how frequently misalignment occurs across its models. The framework favors disclosure even when an incident’s significance remains uncertain.

OpenAI says reports should describe the setting, timing, impact, investigation and unresolved questions, with mitigation information where available. It separates cases ready for disclosure from those requiring further investigation. These are commitments about documenting behavior, rather than a common statistical test against another laboratory.

TENS analysis: Incident narratives can help define what a detector should catch; monitoring statistics can help establish how often it is being used. Neither replaces the other. A laboratory publishing more cases may be more transparent, may be encountering more problems, or both. Ranking safety by the number of disclosures would reward silence unless reporting practices and exposure were comparable.

A practical comparison would therefore keep three records side by side: the work delegated, the controls applied to that work and the failures investigated. Each should specify its population and time window. Changes to task definitions or monitoring rules should be visible, so an apparent improvement can be distinguished from a change in accounting.

For organizations deciding how much responsibility to delegate to AI, this suggests a concrete purchasing question: what evidence supports this particular permission? The ability to draft an experiment does not itself establish authority to deploy code or release data. A useful supplier report would connect each consequential permission to observed performance, review requirements and remaining uncertainty. That connection would make the new disclosures actionable without pretending they already settle the broader question of autonomous research.

Sources: Anthropic’s research-automation measurements and methodological appendix; Epoch AI researchers’ task-taxonomy proposal; OpenAI’s model-misalignment reporting framework.

Featured image: Software code on a computer monitor, used as an illustration; this is not a photograph of either laboratory. Photo: Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.