Suppose somebody came out with a convincing proof that the idea of recursively self-improving AI contained a mathematical contradiction of some kind. That any attempt to build a system capable of improving itself would eventually hit a fundamental roadblock. Say, for example, that at a certain point further gains in intelligence became prohibitively resource-intensive. Or perhaps there would always be unavoidable tradeoffs: for instance that any AI system becoming increasingly good at maths might have to eventually stagnate in its ability to write poetry.

These examples are kind of nonsensical. What I’m really asking you to consider is how you would feel if you were told that an “intelligence explosion” was impossible. Such news would change the tone of the broader conversation around AI. It would not diminish the scale of either the risks or opportunities created by increasingly capable AI systems, but it would be a step towards making the conversation more manageable. It would take the most radical scenario out of consideration and free some of our attention to focus on the many other, more practical questions.

I want to stress that I’m not claiming that RSI is impossible and there’s no need to worry about an intelligence explosion. Rather, I am bothered that we are spending an enormous amount of attention entertaining the implications of an idea that has not, as far as I can tell, been placed on rigorous foundations. At least not in a way that feels general enough to connect clearly to the kinds of AI systems we are building today.

The scenario sounds intuitively plausible: we build an AI system capable enough to design a better successor, which in turn designs an even better successor, setting of a recursive process. But “intuitively plausible” is not a high bar, and when theres this big of a spotlight on the idea, it’s not good enough. When we talk about RSI, we invoke mathematical language. We hypothesise about an accelerating rate of increase in the intelligence of an AI system. We borrow the concept of a singularity to describe what would happen once some threshold is crossed and an unsupervised feedback loop begins. Yet these concepts have precise and rigorous meanings in the domains we’ve borrowed them from. In the RSI conversation, they are often ambiguous. For instance, it is not clear what it means for an AI system to be ‘more intelligent’ than its predecessor.

Let me share an experience that changed how I think about the role mathematics can play in problems like this.

A couple years ago, I came across Alan Turing’s 1936 paper “On Computable Numbers, with an Application to the Entscheidungsproblem”. In it, Turing gives a negative answer to the Entscheidungsproblem, a problem posed by David Hilbert and Wilhelm Ackermann in the early twentieth century. Roughly speaking, it asks whether there exists an algorithm that can take any statement in a suitable formal system and mechanically determine whether that statement is universally valid.

“Universally valid” is a technical term from formal logic that I will not elaborate on here. For the sake of this essay, we can trace the broader ambition behind the Entscheidungsproblem back to Gottfried Leibniz’s dream of building a machine able to determine the truth values of mathematical statements. Note that in this discussion, we use ‘machine’ and ‘algorithm’ somewhat interchangeably, what we care about ultimately is that a problem can be solved mechanically, without requiring creative insight.

In any case, if you wanted to know whether such a miraculous machine could exist, where would you begin? The intuitive approach, perhaps, would be to attempt to build such a machine.Turing proved that no matter how hard you tried, your engineering efforts would be in vain. No such machine or algorithm could exist.

There’s a remarkable section in his paper where he reduces the work of a human “computer”, that is, a person carrying out calculations according to fixed instructions, to its bare essentials. I’ll include a sample of this section here:

“Computing is normally done by writing certain symbols on paper. We

may suppose this paper is divided into squares like a child’s arithmetic book.

In elementary arithmetic the two-dimensional character of the paper is

sometimes used. But such a use is always avoidable, and I think that it

will be agreed that the two-dimensional character of paper is no essential

of computation. I assume then that the computation is carried out on

one-dimensional paper, i.e. on a tape divided into squares. I shall also

suppose that the number of symbols which may be printed is finite. If we

were to allow an infinity of symbols, then there would be symbols differing

to an arbitrarily small extent. The effect of this restriction of the number

of symbols is not very serious […]”

Turing systematically strips away incidental details until what remains is an abstract model of mechanical computation: what we now call a Turing machine.

Once a rigorous mathematical object has been defined, mathematics can act on it. We no longer need to reason about the intuitive ideas that inspired the definition. The object can be manipulated symbolically, and most importantly, we can prove theorems involving it. Turings proof delivering a negative answer to the Entscheidungsproblem is a great example of this.

However, impressive as it may be, the reason this paper has stayed in my memory years later is not for the final result it proves. It is the creative leap of transforming an informal question of the type “Can there exist an algorithm that …?” into a rigorous mathematical definition that can begin searching for an answer. I think something analogous is missing from our discussion of RSI.

A sufficiently general, rigorous theory of recursive self-improvement could prove valuable far beyond answering the binary question of whether intelligence explosions are possible. It could help reveal nuances and constraints our intuitive picture can’t express. It could help us towards a measure of AI progress beyond the particular benchmarks we currently use. Turing set out to solve the Entscheidungsproblem, but the value of his contribution went far beyond his original goal. The formalism he invented became foundational to the field of computer science.

Recursive self-improvement is an idea that has firmly captured our attention, and if we found a unifying way to make it precise, it might help us better understand where progress artificial intelligence can or cannot go.

A final note: I do not want to make the claim that no one has tried to give a precise definition of RSI, that would not be true. It appears to me, however, that no sufficiently general theory has emerged to capture what we intuitively think of as ‘intelligence’. I have not come across such a theory, and our continued reliance on benchmarks seems to suggest that its missing. If you are reading this, and something comes to mind, please share it with me, I’d be very happy.