I like vibe coding. I want to say that up front because this is going to read like a complaint, and it isn’t one. Describing what you want, watching an agent build it, and having something working an hour later is a real pleasure. For a solo developer or a two-person team it feels like finally having the staff you could never afford.
The problem is what that hour does to your relationship with the code.
When you write a system yourself, even badly, you come out the other side with a map. You know where the state lives, which function is doing too much, where the ugly workaround is and why it’s there. When an agent writes it for you, you come out with a product and no map. For a while that doesn’t matter. The product works. Then a bug shows up in production that the agent can’t fix on the first try, or a customer asks for something that cuts across the architecture, and you discover you aren’t the owner of this system. You’re a spectator to a black box that happens to live in your repo.
I’ve watched this happen to people who are good engineers. It isn’t a competence problem. It’s a side effect of the workflow, and it has now been measured.
Back in January, Anthropic published a randomized controlled trial on exactly this. Fifty-two developers, most of them junior, were asked to build two features with an unfamiliar async library, half with AI assistance and half without. Afterwards everyone took a quiz on the concepts they had just used. The AI group scored 17 percent lower, roughly two letter grades, and the widest gap was on debugging and tracing execution flow. They also weren’t meaningfully faster. The people who delegated the whole task to the model averaged under 40 percent.
The detail that matters most is what separated the AI users who did fine from the ones who didn’t. The high scorers used the model to generate code and then kept asking it questions: why this approach, what happens if this fails, where does this value come from. They were still prompting. They just refused to stop at the output.
That lines up with what Microsoft Research and CMU found in a survey of knowledge workers the year before: the more confidence people placed in the AI’s output, the less critical thinking they reported doing. Trust the tool, stop checking the tool. Nobody decides to do this. It’s the path of least resistance, and the tool is designed to be a very smooth path.
What’s being lost here isn’t a specific stack. It’s the portable stuff. The instinct for where a bug probably is before you open the file. The ability to read a stack trace and have a theory. Those are the skills that transfer between jobs and languages, and they’re the ones that get you from junior to senior. They are also the first thing to go when every error message gets pasted straight into a chat window.
So what do you do, short of refusing to use the tools, which is both impractical and a bad idea?
The honest answer is that you have to make understanding a deliberate step rather than something you hope happens along the way. The Anthropic data says the step doesn’t have to be expensive. Generation followed by a handful of real questions got people to 86 percent. The problem is that under deadline pressure the step is the first thing you skip, and there’s no record of whether you skipped it.
That’s the gap we built Ninchi for. Ninchi doesn’t verify AI-generated code. It verifies that a human understood the AI-generated code. On each code submission it asks the developer a few questions about what they’re shipping, the kind a good senior would ask in review, and keeps an inspectable record of the answers. If you can explain the change, the record shows it. If you can’t, you find out before production does.
For a solo developer that record is mostly for you. It’s the forcing function that turns “the agent handled it” into “I know what the agent did.” For a small team it’s the thing that lets you keep moving fast with agents without quietly becoming a group of people who can’t maintain their own system.
Vibe coding is a fine way to build. It’s a terrible way to own. The difference is whether a human ever closes the loop, and that is still a choice you get to make.