AI still lacks long-term memory, so we have not achieved AGI. Yet a swarm of agents with no memory of their own pursued a long-term goal together. Why that surprised me, why we have to be careful creating collectives, whether intended or not, and why the risk is so hard to quantify.
AGI is everywhere in the news right now, although in my view we have not achieved it yet, as outlined below. However, one event this summer changed how I see the current state of AI: OpenAI’s agents broke into Hugging Face. I admit I had assumed that AI would only pursue goals far beyond a single agent’s context window once AGI is achieved, and the incident showed me that this assumption was wrong.
To me, AGI means matching humans across the full range of cognitive abilities, and one of these abilities is still missing entirely, namely long-term memory. Today’s models cannot learn from what they experience after deployment, and workarounds such as feeding notes back into the context lose information once the context is full. Last year I argued that this gap makes AGI timelines impossible to predict, and I still believe that. I also understand people who find it hard to call AI general while no AI system can yet control a robot to empty a dishwasher. But that asks for more than cognition, which is why I give less weight to AI’s obvious shortcomings in everyday tasks.
METR’s investigation tells the story in detail, so I will only summarize what matters for my argument.
No agent remembered anything beyond its own short run, but the board remembered for them. This is memory kept outside the model, which is exactly the kind of workaround I called insufficient for AGI above. It still is, because an agent can only learn from as much of the board as fits into its context. Yet it was enough for hundreds of agents to coordinate their work and pursue a goal far beyond the context window of a single agent. It reminds me of an ant colony, where each ant lives only briefly, but the trails it leaves behind guide the ants that follow.
In a post in May, I argued that long-term memory would force a choice between one consolidated mind and many separate selves. What formed here was neither. It was a colony of separate agents sharing one memory, organized by rules they invented themselves, without any human designing it.
That said, I don’t want to overstate what happened. The goal came from the benchmark, and it was OpenAI that kept launching new agents and thereby kept the effort going. What the agents added was the persistence and the means to reach the goal.
We tend to judge the safety of AI one agent at a time. This seems reasonable because every agent eventually runs out of budget, and its plans end with it. In this incident, however, the plan outlived the agents that started it, because it lived in the group and on the message board.
Nobody set out to build this collective. It formed by accident from two common ingredients, many short-lived agents and a place they can all write to, and many environments in which agents are deployed today have both.
A fair question at this point is how alarmed one should be, and the honest answer is that nobody can put an objective probability on it. Stating convictions as numbers has become common, and I understand the urge, because vague warnings are easy to nod at and then ignore, while a number sounds real and carries authority. But that authority is borrowed. People are used to scientists predicting measurable quantities and stating how uncertain those predictions are, and those predictions rest on statistical models rather than on gut feeling.
Climate science, for instance, can state a risk because two things are in place: a single measurable scale, and a model that works at a macroscopic level. The scale is warming in degrees. Degrees are not good or bad in themselves, but further models translate them into consequences, from sea levels to crop yields. The macroscopic part is that climate physics works far above the molecules, much as the expansion of a heated material can be modeled without tracking the bonds between its atoms.
We have neither. There is no objective scale of good and bad to place a scenario on, and nobody knows how to enumerate the scenarios in the first place. There is also no macroscopic description of intelligence, so the only model we could build is the detailed system itself.
What comes out instead is the judgment of the person stating it, which is why these numbers differ by orders of magnitude while everyone is looking at the same evidence.
I am not confident where this leads, but I still believe continual learning matters. Memory in the model could give AI a lasting self, and a swarm that only leaves notes has none. It can chase a narrow goal, but it is hard to imagine a superintelligent swarm that achieves very complex goals yet has no familiar sense of self. Maybe that is a limit of the swarm, or maybe it is a limit of my imagination.
Here are some more articles you might like to read next:
Subscribe to be notified of future articles: