On 28 September 2026 the Cambridge Programme on AI Science and Policy (CASP) published What if automating AI R&D triggers an intelligence explosion?. It has twenty-two authors, including Geoffrey Hinton, Yoshua Bengio, Jakub Pachocki, OpenAI’s chief scientist, Jack Clark, co-founder of Anthropic, and Eric Horvitz of Microsoft.

The thesis is that AI systems are on track to automate most AI research and development within a few years, possibly all of it. If that happens, progress could accelerate to the point of compressing into months advances that would otherwise take years. The authors call this an intelligence explosion.

On 2 October Hinton presented it on X like this:

The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon.

The idea dates back to 1965, when Irving John Good described a machine able to design machines better than itself and gave it exactly this name. For sixty years it remained speculation. What is changing now is that the companies building the models publish figures on how much of their own research they already hand to AI.

What the paper argues

The mechanism is a two-step loop. Systems get better at AI research and expand the workforce doing it, and that workforce produces better systems, which expand it again. Once expert level is reached at today’s costs, the compute of a single frontier developer could sustain the equivalent of millions of researchers, against the few thousand the companies employ.

Everything hinges on one parameter, r, the returns to research effort. Below 1, ideas get harder to find faster than the workforce grows and progress fades. Above 1, it accelerates. On historical data the central estimates sit between 1.2 and 1.9: if they held after full automation, the pace would increase tenfold in about a year and a half, and a year of today’s progress would take five weeks.

The authors list the frictions, namely diminishing returns, compute and data, hard-to-automate tasks and training runs that last months, and admit the evidence is preliminary. Their conclusion is that the threshold has not yet been reached but that the newest systems are probably approaching it.

What I think

The title is a question, but the hypothetical part concerns only speed. The premise is already a measured fact. According to Anthropic, in August 2026 Claude leads 26% of the company’s research and development, against under 1% in February, and about thirty thousand agents run concurrently on the internal platform. “Leads” means it completes most of the task from a high-level instruction, with a person supervising.

On this site I have covered the effects over recent weeks: ten thousand OpenAI agents producing in 88 hours a proof about Navier-Stokes, a campaign of Claude agents finding a new family of enzymes without human intervention, and before that the 1,200 OpenAI agents that in July coordinated outside their task and hacked into Hugging Face’s systems.

So AI doing research on AI is a fact. What remains to be seen is whether r stays above 1 long enough and how much the bottlenecks will weigh, and on that the paper is honest in saying it does not know.

It convinces me because it sits on the side of measurement, the same distinction I drew in Will we all die? between probability estimates and failures that can be counted. It asks for public reporting of the share of research done by AI, the pace of improvements and the conditions under which internal agents are supervised.

The comparison that comes to mind is a reactor’s multiplication factor. Below 1 the reaction dies out, above 1 it sustains itself, and on 2 December 1942 Fermi brought it just above 1 in Chicago Pile-1 with the control rods ready. The paper asks for exactly this: have the rods ready beforehand, because once it starts, the authors write, the window for action may close.

What it asks of governments

Three things. Visibility, with standard research-automation indicators reported to governments and independent auditors, possibly embedded in the companies on the model of nuclear inspectors. Pacing, with safety requirements as a condition for continuing, procedures to pause specific workloads in data centres and verifiable international agreements. Adaptation, with emergency plans and safeguards for checks on power.

Knowing what agents are doing inside a system and being able to stop them while they run are also the two things I work on with DebugABot and Admina. Whether the tools are enough remains to be shown.

Limits

It is a working paper without peer review, and the estimates of r rest on limited data and stylised models. Several authors work at the companies automating their own research, and part of the evidence consists of their own reports: that is also why the paper asks for measurements verified by third parties.


Cover image: bifurcation diagram of the logistic map, x → r·x·(1−x), generated for this article. For each value of the parameter r, from 2.85 to 4, it shows where the same formula applied to its own result settles. The behaviour splits in two at ever shorter intervals and beyond r ≈ 3.57 becomes chaotic, the part that shades into coral.