Diverting trains of thought, wasting precious time

I am still dreaming that our society might wake up from the present LLM mania. Now is a difficult time. To pick from today's doom-press, “Fast food for the brain” seems an appropriate metaphor. LLMs are especially corrosive to education, and therefore to our skills and abilities as a society and our survival as a species. I am pinning hopes on hypothetical future regulation (no, I don't know either) and the mother of all stock market crashes (more likely, at least) have any hope of putting LLMs back in their box (but not bottle).

Peer review processes are having an especially difficult time, owing to inundation with LLM-generated submissions. Many smart people are right now thinking up guidelines attempting to stem the tide, but these are emergency flood defences rather than sustainable land management. What might the latter involve?

One response is simply to “lean in” and let the torrents rage: use LLMs for review also! I think this is a terrible idea. Given that the value proposition of LLMs is faster but lower-quality knowledge work at a cheaper-than-human prices, as arbiters of quality they are not fit for purpose. “Quality of what?” is also worth asking: working code is one thing, thoughtful criticism quite another.

The most terrible pattern of LLMs is their escalatory nature, common to many technologies. In response to things getting unmanageably large or complex, over-powerful machines promise to tame the complexity for us, but in reality simply generate more of it while making us more beholden to the machines. This is what Illich referred to, mockingly, as “attempting to solve a crisis by escalation”.

How can we “lean out” of this dynamic? I don't know, but one tentaive proposal is that the only acceptable use of LLMs, at least in the context of peer review, is using them to combat LLMs. If an LLM thinks that a submission is generated by an LLM, pencil that submission for rejection. This feels uncomfortable not so much for its risk of false positives (humans remain in the loop) but because it is also escalatory: it creates an arms race, for undetectable LLMs and ever-more sophisticated detection abilities. However, as applied to peer review, this struggle is at least one level removed: the cat-and-mouse game is in the quality of pre-filtering on a logically shallow criterion with a knowable answer (“Was this text generated by an LLM?”), rather than on a deep one that even humans are not reliably good at answering (“Is this paper worth accepting?” or “Is this proposal worth funding?”).

It feels important to keep this “deep versus shallow” distinction in mind, vague though it is. Being syntax-extruding machines, LLMs are able to do shallow work to an acceptable quality level, apparently somewhat reliably. The question remains of whether it's wise to make them our default tool for that shallow work, and I still believe not. That's partly because humans need to learn from the shallow tasks as well as deep, and partly because removing any brake on shallow work will have unintended escalatory ill effects. But from my armchair, non-expert perspective, it seems deep work is something LLMs per se cannot reliably do. In apparent success stories, such as in mathematics, they are only useful if combined with agents that have strong external validity checks, such as Lean mechanisations, providing a feedback loop. In other words, when errors cause a hard bounce onto another path, a fast-iterating dullard can eventually stumble on a useful result. Review of papers or proposals does not come with that kind of bounce. I don't want them reviewed by a dullard, either human or machine, and nobody should.

[/research] [all entries] permalink contact