Coding agents are broken out of the box.

When the majority of the conversations surrounding them (outside of blind praise) are about preventing them from generating slop, or telling us that code quality doesn’t matter anymore, surely there’s a problem.

Now, we could pin the blame on the models, by saying the training datasets include lots of mediocre to bad code, and LLMs are fundamentally predicting the next most likely token. While certainly true, these are basically unsolvable, and mitigating these issues when using LLMs just comes with the territory.

no, it’s the harnesses that are wrong

Agentic code harnesses are wonderful tools, but between being configured to generally work for everyone, overbuilding functionality, and being skewed towards output & speed, they set us up to fail. If someone was to be overly cynical, they might say this is a ploy to make us all burn more tokens, but I’d say this is a consequence of “tech brain”, or the conventional thinking around what AI is going to do for us, e.g. turn us all into 10x engineers and usher in a new golden age of productivity.

So, how does this break coding agents? We have to look at this in terms of friction. Friction is supposed to describe the unnecessary effort, or pain points that gets in your way. With this framing, it’s very easy to think the optimal approach would be to remove all friction, right? Except, friction can be good or bad depending on the context, i.e., there are situations where you want to be slowed down. And, in my experience, harnesses get this wrong entirely and are removing good friction, while introducing tons of bad friction.

friction burns

Some of this bad friction is unfortunately inherent to coding harnesses, because reading code uses different parts of your brain, than when you read language, and working with an agent means switching between those two modes constantly. Unless you’re not reading the code, but lets put that to the side for now.

On the other hand, some of this bad friction is just because of how verbose agents can be, both when “responding” to the user or “writing” code, i.e. there’s huge cognitive overhead. I don’t want to say fixing this was obvious, but very quickly we saw approaches like caveman, i-have-adhd, & ponytail to try and address this. Now, these are all great steps in the right direction; make it easier to read the agent output & reduce the amount of code that might need to be reworked, but that’s only one part of the problem.

slop begets slop

“Write code that reads like the surrounding code: match its comment density, naming, and idiom.”

This line shows up in multiple Claude Code system prompts, even Opus 5.5, and there is a “Following conventions” section in OpenCode’s system prompt which has an equivalent line. The problem here is that we get tired, but agents don’t, so you will let something slip in.

Then you’ll do it again & again, and now you’re stuck in this self-amplifying feedback loop where staying in the “smart zone” isn’t important anymore because generating low quality code is baked into your codebase already. At this point, the cost to bring the bar back up is so expensive, it’s no surprise we hear people saying we shouldn’t ever look at the code anymore. But if you don’t believe that, the obvious fix is to give agents more context, design upfront, essentially make them engineers. Except they’re not engineers.

agent configuration does not an engineer make

The bulk of why agents are broken comes from trying to make coding agents engineers. They can be great developers/coders, but not great engineers. Great engineers understand the tradeoffs they’re making, and their goal is to ensure that the continued development of a project becomes, and stays, effective. LLMs produce statistically likely outputs based on their context, which is entirely divorced from what the outcomes would be, so even if we could provide enough context, they cannot make these engineering decisions for us.

We don’t always need great engineers though, because there are small problems that will stay cheap no matter how you solve them, or the solutions are disposable anyway. The hard part is figuring out which problems are which, so go and try to find those small, cheap problems in your life and vibe code away.

However, when we configure agents to “do” and “know” software engineering, we are still saying that it’s important, but that it can be offloaded. I find this very hard to believe. To some extent, this isn’t a radical statement, since we’ve adopted all sorts of planning skills & workflows to design top-down and make sure we’re doing the engineering upfront. Again, this is a step in the right direction, we’re acknowledging that we still need to take ownership of decision making, but I’ve found this to be another form of bad friction, on both ends.

The map is not the territory, we know this. Yet we fabricate these maps, hand them over, and realise they were wrong after an expensive review of the code. Then we vow to make sure we’ll include enough detail next time, instead of realising if we had the perfect map, we’d have written the code already.

Are we then stuck in between a rock and a hard place, where on one end, an agent will over-engineer trivial problems, and is incapable of doing real engineering when it’s actually required? I don’t think so, but the solutions are going to look different than what we have today.

this is my agent. there are many like it, but this one is mine.

I will admit that until very recently, I put very little stock into doing any configuration of the harnesses I’ve used, besides repo level AGENTS.md files. By using the out-of-the-box experience for so long, I’ve had time to become familiar with the fundamental problems we’re all building on top of.

Unfortunately, I don’t think there’s any one single solution to these problems. On the bright side, it also isn’t a particularly hard problem to solve, nor do I think it requires a lot of effort. This is because we’re all different, so your solution is not only going to look different to mine, it should. It has to.

Of course, I’m not going to leave you stumbling around in the dark. Shared problems still have a shared underlying solution, it’s just that the actual implementation won’t be.

So why are they broken and what questions should you be asking?

- Agents skew towards increasing output at the cost of avoidable

cognitive overhead

- Where do agents create unnecessary cognitive overhead in your workflow?

- Agents expect you to maintain perfect standards while regressing to

the mean

- How could you work with an agent to maintain standards?

- Agents try to be a one size fits all tool regardless of who’s using

them, or the problem being worked on

- How could you make an agent integrate more closely with your natural workflow?

- Agents are great at creating code, but can’t think for you

- How can you integrate decision making into your agent sessions?

In Part 2, I’ll be sharing specific solutions that are working for me right now.