Anthropic uses a six‑level scale, adapted from Epoch AI's automation framework, to describe individual kinds of research work. At AL3, Claude collaborates: it handles substantial parts of a task under close human direction. At AL4, it leads: a person states the problem, then the agent carries out most of the task and brings the result back for review. AL5 would mean the agent notices, scopes, completes and deploys the work without needing a human in the loop. Anthropic reports Claude at AL4 for 26% of its measured R&D basket, at AL3 or higher for more than 90%, and at AL5 for none of the measured categories.
The paper's own example is a failed overnight data pipeline. At AL4, Claude could inspect logs, find the broken stage, write and test a fix, handle surprises and document the result. An engineer would still decide whether to release it. That is a meaningful change in who performs the investigation, but it is not a claim that Claude chooses the lab's research direction or deploys changes on its own. Anthropic's example and footnotes make the human handoff part of the definition, not an incidental caveat.
The change over time also needs careful wording. Anthropic's figure rose from less than 1% in February to 26% in August. That is a rise of more than 25 percentage points in its index. Calling it a 26‑fold increase would imply a known February baseline when the company only supplies an upper bound; a zero starting value would make a ratio meaningless. Nor does the increase alone tell us whether model R&D became 26% faster. An automation rating concerns who does a type of task, while speed depends on the time saved, review burden, failed experiments and other bottlenecks.