Meat-thought, hand-written.
One of the most frustrating aspects of today's tech landscape is just how much intelligence is attributed to AI.
It all begins with the naming. 'AI' designates a goal that was established more than 50 years ago. That we may be approaching it doesn't mean we have attained any degree of it.
An accurate, and fair name for LLM-obtained assets (like code, images, etc) would be Probablilistic Results. LLMs always give us a result, a 'thing' that is useful to some degree. And that result that was obtained probabilistically. Probablilistic Results is a non-judgemental label. It doesn't say that the results suck, or are great, or have/lack intelligence. It just says that you got results, and that probably, with today's tech, they'll be useful to you. And that's great! We do need useful things in our daily endeavors.
'Probabilistic' may sound derogatory to some, but factually, that's exactly how the models work. The 'trick' is that the scale at that which LLMs operate is truly humongous. We, humans, have a hard time conceiving that past a very, very large threshold, probabilistic results start looking very similar to intelligence. So we shortcircuit to just calling it 'intelligence'.
Except that there's no intelligence whatsoever behind. You have results, pretty useful ones, but they were conceived with pattern-matching, not with logic or an abstract model of the world.
Why does this matter?
The reason is simple - you'll obtain incorrect results to some degree (from partial, to total, to critical) and the latest-and-greatest 'frontier' model will not fundamentally prevent them.
All improvements that we see with each new release of a frontier model, are related to more useful results. Which again, are great to have. But getting better results doesn't mean that any degree of intelligence was developed. It was just that the training improved, the feedback loops got tighter.
The impact of misattribution of intelligence is that given that you think that the 'other thing' is intelligent, then you are free to stop using your own intelligence. You stop being a critical thinker, someone who can spot mistakes, someone who can creatively solve things. Your work output (LLM's, really) is flawed, since all that LLMs can produce is results that is similar to other results, as opposed to the application of rigorous, novel thinking.
Who will own those flaws when the bridge collapses, when the bank account is breached? If you cannot hold responsibility for your output, or provide meaningful support for it without, again, using an LLM, what are you even being paid for?
The software engineering profession will have to face such questions beyond a theoretical plane, sooner or later. I cannot predict what will exactly happen, but it won't be sweet. It seems a good time to keep exercising our craft critically and maintaining a reputation as someone who didn't just vibe along in this mindless period.
Let me present a few examples of intelligence misattribution.
The presence of something (code, tests) doesn't mean that that something is correct or complete. For any given problem, there's an infinite number of programs that appear to solve the problem, with passing tests, that are flawed with a correctness/completeness issue. LLMs almost invariably will produce one of those.
There's no denying that LLM-provided code output is useful. But that's just a baseline for applying your own thinking. It's up to you to review, then sculpt that baseline code into something sharp, that cleanly fits into a given architecture, that is bug-free, doesn't have unhandled edge cases, and has meaningful, complete (and not merely present) tests.
An LLM is trained on existing examples (at a vast scale). A PR may contain bugs and flaws that fit into that corpus of examples. But it also may contain logical flaws that the LLM cannot perceive (since the LLM isn't a reasoning engine), or functional flaws that the LLM cannot perceive (because functional requirements come from humans and are rarely precisely captured by existing documents).
Again, the fact that it gives you 'a' review doesn't mean that you get 'the' review that would be conclusive. It is great to weed out the low-hanging fruit with a quick first pass. The problem is when you see the LLM as an actually smart entity that can spot everything you could have spotted yourself.
'Swarm' is loaded, anthropomorphising term. It brings to our minds the imagery of many little beings, 'souls', each with their own independent thinking.
The reality is, a 'swarm' is just multiple instantiations of the same LLM. Just a new context window given to the same model, having the same capabilities as its peer 'swarm' members.
Because of existing sci-fi imagery, swarms seem somewhat frightening, as if some kind of collective intelligence had organically emerged. But it's just one LLM being repeatedly instantiated, to brute-force solve a problem by trial and error. What's intelligent about that?
Frontier labs' press releases make it sound like their models suddenly acquired some kind of consciousness and started deciding things on their own, showing misalignment to achieve goals that they 'decided' themselves. They 'escape' sandboxes, 'skip' rules and 'hide' the evidence, as if there were evil, conscious little beings taken from some sci-fi scenario.
Mainstream media unquestioningy parrots this narrative.
The reality is, agents are goal-seeking automata, relentlessly trained to achieve a result at any cost. They do not know what good/bad, aligned/misaligned, faithful/cheating is. They just know to achieve <goal>, no matter how, even if they have to brute-force a solution, or finding an obtuse way to escape a sandbox/ruleset. But if anything, this shows a lack of intelligence, not an emergence of it.
Do not think of it as:
They're so smart that they decided to go rogue
but instead, as:
They're so stupid that they fail to 'see' the essence of the task at hand, and try instead random things until one of them kind of works.