Recursive, Toward What? Recursive Self-Improvement and the Question of Aim.
The debate over recursive self-improvement has fixed on whether the loop will close. The more consequential question—already being answered by funding, benchmarks, and publication norms—is what it closes around.

In 1965, the mathematician I.J. Good wrote a paper that included this sentence:
“the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.”
Six decades of commentary have worn a groove into the first half of that sentence—the promise, or threat, of a machine that redesigns itself again and again, each version smarter than the previous one, until something stands on the other side of the loop that no longer needs us to invent for it.
Good called it the last invention. He did not call it the last decision.
The second half of the sentence—the provided that—is the part the groove keeps skipping. Someone or something still has to decide what the machine is improving toward, and whether it keeps answering to anyone once the self-improvement loop starts turning on its own. That is a question about aim, not speed. We keep asking whether the self-improvement loop will close. We ask far less often what it will close around.
*****
Recursive self-improvement, stripped of its mythology, describes something fairly simple: a system that uses its own outputs to improve its capacity to produce outputs. A model that helps design its successor's training run. A research pipeline that does the work a building full of people used to do. Good gave the idea of an “intelligence explosion” its first serious statement in 1965; Vernor Vinge and later Nick Bostrom gave it its modern vernacular through phrases like “technological singularity” and “seed AI.” All of them had a sense that there was a threshold that, once crossed, cannot be uncrossed. The industry has since translated the same idea into a flatter phrase, “AI R&D automation,” which has the advantage of sounding like a banal line item and the disadvantage of hiding what is actually being claimed.
What is actually being claimed is an event. Somewhere out past the horizon, the self-improvement loop closes; before that point, humans are still load-bearing, and after it, they are not. Framed this way, the only serious questions are whether and when, and the debate splits, predictably, into camps that argue about timelines. It is a debate built for headlines. It is a generally unhelpful debate when it comes to trying to govern anything. You can only argue about a future event before it happens. You can only prepare for it in the abstract. And abstract preparations rarely hold once the incentives on the ground start pulling in a different direction.
The more useful frame, and the one this essay will use, treats recursive self-improvement not as an event but as a rate—an improvement loop with feedback, and a feedback that is already being tuned in public by choices being made this year. That reframing does not resolve the anxiety attached to the idea of recursive self-improvement, but it does relocate it from a future we can only speculate about to a present we can actually examine.
*****
Here the economics turns out to be more informative than the science fiction. A recent formal treatment of the problem, Cunningham and Althoff's “The Economics of Recursive Self-Improvement,” models the self-improvement loop as a pair of feedback effects rather than a single switch. One is a self-loop: the existing stock of algorithmic knowledge makes the next improvement to that knowledge easier, or harder, to find. The other is a core loop, routed through capability: better algorithms produce systems that are better at the research that produces better algorithms. Self-sustaining acceleration in their model requires the elasticity of research productivity with respect to capability—how much a gain in raw model capability translates into a gain in the rate of useful research—to clear a specific threshold, something above roughly 0.15. Their current empirical estimate, calibrated against the period since coding agents entered the research pipeline, is closer to 0.09. Rising, but short of the singularity.
This is how recursive self-improvement is not a threshold you cross, but rather a gain you tune. The gain is the feedback multiplier that determines how much a change at one point in the loop gets amplified (or damped) by the time it comes back around. A high-gain loop takes a small input and blows it up into a large output each pass; a low-gain loop takes that same input and lets it die out. Whether a feedback system runs away, stabilizes, or oscillates depends almost entirely on that one number.
The idea of tuning a gain is worth focusing on, because the natural next thought—tuned by what or by whom?—has an answer, and the answer is uncomfortably familiar. Elasticities like these are not laws of nature. They are shaped by what gets funded, what gets benchmarked, what gets published in the open and what gets kept behind a paywall, which experiments a lab decides are worth running and which get quietly shelved. This is directed technical change again, the same logic these Notes have examined in the context of the direction of pharmaceutical innovation—except the “invention” now under direction is the machine that does the inventing. The gain dial exists. Someone's hand is already on it.
*****
An abstraction this clean invites a test, and one arrived this August. Researchers at Princeton, led by Peter Kirgis and Sayash Kapoor, ran what they called a shadow evaluation: they gave Claude Opus 4.8 a set of unpublished research problems submitted to NeurIPS 2026, along with six days and three thousand dollars in compute, and let it work the way a research team would—reviewing literature, running experiments, drafting papers, iterating on results. Then they submitted on a blind basis the output back to the original authors of the problems it had been assigned.
Both papers were rejected. Not for lack of effort, and not for lack of competence at the mechanics. The system could execute—run the experiments, survey the field, produce prose that read like a paper. What it could not do was exercise judgment: explore genuinely different approaches before committing to one, notice when a strategy was failing and abandon it for another, absorb feedback and change course rather than polish the same idea more carefully. “The papers were nowhere close to the mark,” Kapoor said, “when it came to being at the quality of a top AI conference.”
These Notes made a version of this same distinction recently, writing about AI agents and biomedical coordination costs: agents can search faster, draft faster, monitor more continuously than any human team, but none of that touches the deeper question of whether the parties involved actually want the same thing. The lesson generalizes. Execution is not invention, any more than capability is delivery. Current systems are collapsing the execution layer of research—the searching, the drafting, the running of experiments—well before they touch the judgment layer that decides which experiments are worth running in the first place. That gap is not a permanent feature of the technology. But it is the honest state of it right now, and it is doing a great deal of the work that keeps the elasticity below 0.15.
*****
None of this has stopped the louder claim from being made anyway. In July, OpenAI's Sam Altman told an audience that the singularity was, in effect, already underway. It is worth noticing what kind of claim that is: not a measurement, but an announcement, made by someone with every incentive to make it. The Princeton study and the elasticity estimate are not counter-announcements. They are measurements, and measurements are less exciting, and also more trustworthy, precisely because they were not built to be exciting.
The more interesting response to the moment has come not from the frontier labs' marketing departments, but from their governance documents. Anthropic's Responsible Scaling Policy, now in its third version, defines an “AI R&D automation” threshold with unusual specificity: a model capable of compressing roughly two years of 2018-to-2024-pace AI progress into a single year. What is notable is not the number itself so much as the framing around it. The policy does not treat this automation threshold as the finish line, that is, the moment the last invention occurs. Rather, it treats it as a gateway—a point at which the risks compound across every other domain the technology touches, and a point that accordingly demands its own countermeasures: continuous monitoring of development activity, deeper interpretability work, published risk reports subject to outside review.
That is direction-setting, in something close to real time, and it is worth recognizing as such. No single lab owns the gain function of this loop in full. Compute allocation, publication norms, safety thresholds, benchmark design—these are the AI-era equivalent of the advance market commitments and priority review vouchers these Notes have described previously: mechanisms built not to further fuel an engine but rather to point it in a particular direction. The self-improvement loop is being governed already, if unevenly, by people who seem to mostly agree it should not be governed by accident.
*****
Here is why the distinction between event and rate matters more than it first appears to. With ordinary directed technical change—the kind that has decided for a century which diseases get treated and which get neglected—a society is capable of redirecting mid-flight. A market commitment can be created after the market has already failed to produce a vaccine. A foundation can seize the direction signal away from the price system years into a field's neglect, and the field can respond. Direction, in that world, is a standing responsibility precisely because it stays open; it can be picked back up whenever it was set down.
A self-improvement loop whose elasticity has genuinely become recursive may not extend the same offer. By the time its direction is visibly wrong, its velocity has already made correction expensive, possibly prohibitively so. This is the sense in which recursive self-improvement is not simply directed technical change wearing a new label. It is directed technical change with a closing window—and the evidence published so far this year says the window has not closed, and perhaps may not be closing as fast as the loudest voices claim. But that is no reason to stop paying attention to the narrowing gap.
*****
Good's sentence, read in full, was never really a prophecy. It was a condition. The first ultraintelligent machine would be the last invention we needed to make—provided the machine stayed answerable to us. Sixty years on, we are closer to being able to measure that provision than Good ever was: we can estimate the elasticity, watch the shadow evaluations, read the thresholds a serious lab is willing to write down and defend in public. What that measurement shows, right now, is not a machine straining against its harness. It is a machine that is very good at the parts of research that were always going to be automatable first, and not yet good at the part that decides what the research is for.
That is the cautionary note, and it should be taken as one: the elasticity is rising, the execution layer keeps closing the distance to the judgment layer, and the governance documents attempting to track that distance are themselves works in progress, written by parties with obvious interests in how the story gets told. None of that is a reason for false comfort.
But it is also, still, a reason for something better than dread. The loop has not yet closed around a direction nobody chose. The elasticities in that model are not physical constants; they are being set by current decisions about what gets funded and published and governed—decisions that remain, for now, ours to make and remake. Good called the ultraintelligent machine the last invention we would need to make. He was careful enough to add the provided that condition. We should be careful enough to keep meeting it—not once, as a threshold crossed, but continually, as a choice we keep making for as long as the making is still ours to do.♦