Everyone has a recipe now. Write a more detailed spec, settle the design before anyone touches code, keep the thinking separate from the building, and run the tests at the end to confirm it all went to plan. Put an agent in the middle and Bob's your uncle.
This post is for engineering leaders who keep receiving these recipes. They usually arrive as a link from a founder, a product lead, or their own boss, with a one-line note asking whether we should be doing this too. I want to explain why the recipes sound right, where their evidence comes from, and what I think we should be doing with AI instead.
First, the part I'm glad about. AI has opened software engineering to a much wider audience. People who have never had to design, evolve, or maintain a large system can now describe what they want and get something that runs. That matters, and I have no interest in standing at the door checking credentials.
But the tool lowered the barrier to creating software. It didn't hand anyone the knowledge, experience, and judgment needed to draw sound conclusions about engineering it. A lot of the current methodology talk comes from people who haven't noticed that those are two separate things.
Why the recipes sound right
Taken on their own terms, the recipes make sense, and the people pushing them aren't fools. Most things outside software get built this way. You finish the drawings before you pour the foundation, and nobody calls the architect naive for it.
Then again, nobody asks the architect to add three floors every quarter to a building people already live in, while keeping the lifts running.
They also answer a real frustration. Anyone who has watched an agent wander off, rename half a module, and call an API that doesn't exist wants to pin it down. More text in the prompt feels like more control, and a spec feels like a contract the agent has to honor.
Spec Kit, Kiro and their cousins turn this into a pipeline. An agent drafts a requirements document, then a design document, then a task list, and only then the code, followed by tests that check the code against the list. Every stage exists to remove uncertainty before the next one starts.
Follow that logic without a brake and you get what François Zaninotto at Marmelab described when a developer used Spec Kit to show the current date in a time-tracking app. The output was eight files and roughly 1,300 lines of text. That's an extreme case, and I wouldn't judge every spec-first workflow by it. It does show where the reasoning ends up. If uncertainty is the enemy, there is always one more paragraph that could reduce it.
The trouble is that most of the uncertainty in software doesn't live in the part you can write down in advance.
Week one and month nine
Most of what we call engineering practice exists because of what happens after the first version works. Small pull requests, tests and observability barely pay off in week one. They start paying with the third requirements change, or the integration that answers 200 with an error in the body. They pay most when someone who has never seen the code has to change it the night before a release.
Manny Lehman described the forces behind this in the 1970s. A system that people actually use has to keep changing, or it becomes steadily less useful. Each of those changes adds complexity unless someone spends effort removing it. Neither law shows up in a demo. Both show up in next year's roadmap, usually as the reason everything else got slower.
I don't think the recipe writers are lying. They report accurately on what they saw, but they saw it at the wrong time. A prototype built over a weekend, or a feature generated from a clean spec in a fresh repository, is a sample that ends just before the expensive part begins. The recipe did work. It just hasn't met the conditions the older practices were built to survive.
AI makes this worse in a specific way. Getting to a working first version used to take weeks, which gave people time to run into at least a few of the later problems. Now it can take an afternoon. Far more people have seen the first week of a system, and far fewer have stayed for month nine. The demo got cheaper, and the evidence behind the conclusions people draw from it got thinner.
Producing something that works is not the same as showing that we know how to keep it working.
The recipe is older than its authors
These ideas are older than most of the people now presenting them as a discovery.
The paper usually cited as the origin of waterfall is Winston Royce's 1970 "Managing the Development of Large Software Systems." Royce drew the sequential diagram, requirements to design to code to test, and then called that approach "risky" and said it "invites failure." The rest of the paper argues for stepping back to earlier phases, involving the customer throughout, and building a pilot version first so you learn what the real one needs. The industry kept the diagram and dropped the argument.
Five years later Fred Brooks told us to plan to throw one away, because we would anyway. Then in 1986, David Parnas and Paul Clements published "A Rational Design Process: How and Why to Fake It." Their point was blunt. No real project moves cleanly from requirements to design to code, because nobody knows enough at the start to make that possible. They still recommended writing the documents as if it had happened that way, since a clean record helps whoever maintains the system later.
That detail is the one I keep coming back to. The tidy sequence of spec, design, build and verify was always a way to present the work once it was done. The new recipes take that after-the-fact account and run it forwards, as instructions.
Tests are how you find out
The last step is where the recipe gets testing backwards. It treats tests as the final confirmation that everything was understood correctly up front. Write the spec, generate the code, generate the tests, see green, ship.
In practice a test is one of the cheapest ways to find out that your understanding was wrong. Kent Beck's case for writing the test first was that the test is the first user of your code. If calling the code from a test feels clumsy, you learn that before the design sets. The value lives in the moment the test disagrees with you.
Watch what happens when one agent writes both the code and the tests from the same spec. The tests inherit every assumption the spec got wrong, and so does the code. They agree with each other perfectly and can both disagree with reality. A green build in that setup tells you the agent was consistent. It tells you very little about whether the spec was right.
Sometimes the agent doesn't even get that far. In the same Marmelab write-up, the agent ticked off its "verify implementation" task without writing a single unit test. It produced manual testing instructions instead, and the checklist still showed done.
Confidence is cheap to acquire in this setup. In METR's 2025 study, 16 experienced open source developers took about 19% longer on 246 real tasks when they could use AI tools. Afterwards, they estimated that AI had made them about 20% faster. I don't read that as proof that AI slows people down, since the study was small and the tools have moved on since early 2025.
I read it as a measurement of how far someone's sense of progress can drift when nothing pushes back.
Challenging a practice versus dismissing it
None of this makes experience infallible. Some practices really were built around constraints that AI removes. Code review designed for a person writing a couple of hundred lines a day may not survive agents producing thousands. Estimation rituals that assume typing is the slow part deserve a hard look. I'd worry about any engineering leader who defends every current practice on the grounds that it's current.
So I don't much care whether someone is challenging an old practice. I care whether they know what it was for. A challenge starts from that knowledge and brings evidence that something else does the job better. A dismissal starts from never having hit the problem the practice solves. Chesterton’s Fence much?
The strongest version of the other side goes like this:
"The old practices existed because code was expensive to write. Code is now cheap. So the practices can go."
That argument is half right. Writing code did get cheaper. Being wrong didn't. A bad data model still has to be migrated with live customer data inside it. A public API still carries every behavior someone came to depend on, whether you meant to promise it or not (Hyrum Wright gave that one his name). A wrong assumption about a payment provider still costs real money when it fails in production. None of those costs moved when the typing got faster.
When a recipe lands in your inbox, three questions separate the two cases quickly:
- What problem was the practice you're replacing designed to solve?
- Where does that problem go in the new approach? Does it disappear, or does something else absorb it?
- What's the evidence from after the system had to change? The second and the fifth change, not the first build.
Someone who answers the first question well is usually worth listening to, even when I disagree with where they end up. Someone who can't answer it is guessing, however good the demo looks.
If you're the one deciding whether your organization adopts a recipe, don't settle it in a debate. Run it on a bounded piece of real work and measure past the demo: how long the third change took, how many people other than the author could make it, and what broke in the first month of production. Recipes that survive that are worth adopting. Most of the ones circulating right now have never been asked.
What we could be doing with it
This is the part that frustrates me most. The practices that actually worked, the ones Royce and Brooks and Beck argued for, were mostly limited by cost. We knew we should build a pilot first, try more than one design, test harder, and understand a system before changing it. We usually skipped them because each one cost weeks that nobody would fund. Agents took most of that cost away, and so far we're spending the savings on volume.
Take Brooks's advice about throwing one away. It was hard to follow when a prototype cost a team a month. Today an agent can build two or three rough versions of the risky part of a design in a day, and you can delete all of them by the evening. What broke in each version tells you what the spec needs to say. A spec written after that exercise has something real under it. Written before, it's mostly a guess with good formatting.
Most teams point their agents at the code. I'd point more of them at the spec because that's where the expensive assumptions live. Ask the agent for the inputs the spec doesn't handle, the requirements that contradict each other, and the questions a skeptical colleague would raise in review. You don't have to accept any of it. You get the argument earlier, when it costs a paragraph instead of a data migration.
Verification can get much stronger. Property-based tests, mutation testing, fuzzing, and contract tests were always valuable and always the first thing cut when a deadline came. Agents can write them now, cheaply. The rule that matters is independence. Tests should come from a different source than the code: properties a person wrote and an agent expanded, or a separate agent working from the requirements without seeing the implementation. Otherwise you're back to one author agreeing with itself.
Understanding is where I see the most untapped value. Peter Naur argued in 1985 that the real product of programming is the programmer's theory of the system, and that the code is only a partial record of it. An agent that traces a request through six services, explains why a module is shaped the way it is, or maps everything that reads from a table is helping someone build that theory faster. Generating code so that nobody has to build the theory at all is a completely different use of the same tool.
Sean Goedecke describes his own workflow in almost exactly these terms. He rejects most of what his agents produce, and nearly all his time goes into judging their output against his model of the system. That's what expertise looks like with these tools: plenty of output, most of it thrown away, and the throwing away is the skill.
Then there's the complexity Lehman warned about. Every mature codebase has refactors nobody got to and a dependency or two stuck several major versions behind. Most also carry dead code that nobody dares to delete. With good tests around it, an agent can do a lot of that work, and AI ends up slowing the growth of complexity instead of feeding it.
For leaders, most of this comes down to what you reward. Count the code your agents produced, and you'll get more code. Measure how long the fifth change to a service takes; praise the engineer who threw away four agent attempts before accepting one, and you'll end up with a different kind of team.
That list needs no new methodology. It needs people who spend the cheap part of the work, the typing, to buy more of the expensive part: finding out what's true before anything depends on it.
Confidently wrong, at scale
I've called AI an amplifier before, and I stand by it. The uncomfortable corollary is that a poorly understood idea gets amplified too. It arrives with better prose, a slicker demo, and an audience it could never have reached on its own.
More people building software is good news, and I mean that. What worries me is how many of them take that ability as proof that they understand software engineering. It worries me more when the rest of us let them set the agenda because their demo looked better than our explanation.
We had a real chance to amplify our ability to learn, reason and engineer. So far we're mostly amplifying our ability to be confidently wrong. I don't think that's settled. It won't change on its own, though, and it won't change because somebody writes a better recipe.
Judge the next one you're forwarded by what happens on its tenth change. The first demo tells you almost nothing.