There are people who’ll tell you that you can’t or shouldn’t estimate software work, to the point of outright refusing to do it. I fundamentally disagree with that.
In reality, businesses have to juggle tradeoffs, deadlines, and limited resources, and “it’ll get done when it gets done” doesn’t help decide whether to build feature A or feature B. Or how many people you’d need to finish the work by May.
That said, getting a decent estimate is hard, especially for a multi-quarter project. I’ve tried a number of things over my 15 years in software, with results ranging from spectacularly wrong to actually not too bad. And, based on the last few years, I think I’ve settled on something that works OK-ish.
This involves two components:
- Using the right methodology to get to the initial guess, and
- Layering in the right amount of uncertainty.
This post is meant to be a companion piece (sort of Part 2.5) for my Intro to Project Management for Developers series, though you don’t have to read the rest of the series before reading this post.
What doesn’t work
There are two approaches that I’ve seen fail repeatedly, both producing wild underestimates.
The first one is: any guess that’s based on initial gut feel, especially when not backed by experience delivering software end-to-end. I’ve often seen it from “business”-side people who don’t know what it takes to deliver software. But it’s also very common among developers who haven’t managed a project end-to-end, because they picture just coding the core functionality changes and miss the other 90% of the work needed to complete the project. This is how you get an estimate of “a couple months” on a project that ends up taking a year.
The second approach that doesn’t work, surprisingly, is breaking up the delivery into small steps, each a few hours to a few days long, estimating them individually, and then adding them up.
Well, actually, it works pretty well for projects that take a single developer a month or so to complete. But it often fails for larger projects. This happens for a couple of reasons:
- No matter how thoroughly you design and plan the project, you will always have course corrections, inter-team frictions, missed or changed requirements, and other surprises that consume a lot of energy but are difficult to size upfront.
- Sizing this way is a lot of work. Most people get tired after they size the first couple sprints, give up, and go with their gut feel for the rest of the project. Kind of like drawing this owl:
To be clear, with the second approach, you’d still likely get a better estimate than with the first. But unless you’re really thorough, don’t be surprised if your estimate is off by an order of magnitude.
So, how do you estimate a large project?
The best method I’ve found so far is to use comparable projects. This is supported by research (e.g. this and this) and has been replicated across industries—for example, by Flyvbjerg and Gardner in their book How Big Things Get Done.
The approach goes something like this: suppose you’re estimating a new feature that’s similar to one you built last year. Last year’s project took six months with four developers, or about 24 developer-months total.
Say about a quarter of that effort went into the general feature code, while the rest went into integrating with six banks, each with its own quirks and idiosyncrasies. Your new feature touches the same front-ends and back-ends but only needs four bank integrations instead of six.
A rough estimate for the new feature could then be:
- General feature work: ~6 developer-months (same as before)
- Bank-specific work: ~12 dev-months (4/6 of the previous ~18 dev-months)
- Total: ~18 dev-months
You can refine this further; for example, if you’re really confident that you can reuse some bank integration code and save one dev-month, feel free to subtract that. Though in practice, the uncertainty on this estimate will be at least ~25% (see next section), so there’s no point in stressing over any tiny adjustments.
The main benefit of this approach is that most of the time it automatically accounts for the course corrections, inter-team frictions, and other surprises that you’d have a hard time sizing when estimating bottom-up from small parts.
But what if you don’t have such a great comparable? Then you’d have to reach for less similar projects and add appropriate adjustments. Or maybe look at other teams’ comparable projects, or other projects from your experience at other companies. If all else fails, you can break up the project into coarser chunks and estimate each chunk using comparables. But be aware that once you do that, you go back to missing some of the surprises that inevitably arise when you try to fit the various parts together.
Converting the estimate to a range
So now you have a decent initial guess. But in reality, all estimates come with a range or an uncertainty level, and now we have to figure out the right uncertainty level for this project.
This is where we have to leave any pretense of science or rigorous methodology behind and go just off past experience.
And from my experience, I’d group the projects like this:
- Known territory: You have a solid comparable and you’re touching mostly the areas where your team regularly works. Your estimate is probably off by around ±25%.
- Something kinda new: You don’t have a good comparable, but you’re still mostly on familiar ground. Maybe you’re building a new kind of feature, integrating with teams or APIs you haven’t worked with, or setting up new infrastructure. In this case, the actual work will likely be up to 2x the original estimate.
- Outer wilds: You’ve never done this before and you had to piece your estimate together from POC outcomes or small individual work items. You’re likely undersizing by at least 4x. From past experience, a guess of 4-8x the original estimate is likely closer to reality.
With a more concrete example of the payroll feature above, I’d present the estimate like this in each situation:
- Known territory: “It should probably take 3-4 people a couple of quarters to do this, based on how long it took us to make {other feature}.”
- Something kinda new: “3-4 people in two or three quarters, depending on what kind of surprises we encounter and what we can reuse.”
- Outer wilds: “We’ve never done anything like this. On the surface, this doesn’t look like much, but I’ve seen projects like this really blow up. I wouldn’t be surprised the whole team ends up taking a year or two to finish this.”
…with the last estimate likely leading to a lively discussion on how to limit the scope and deliver at least some sort of useful result as soon as possible.
Alternatively, you can give estimates with an uncertainty level
Sometimes, instead of ranges, you might be answering a question like “Can we finish this by March next year?”
You can apply the same methodology and check where exactly March falls relative to the estimate’s range—though don’t forget to offset it from when you’d start the work, not from right now.
Then, your responses could be something like:
- “That’s really unlikely; other similar features have taken us longer.” Or,
- “It’s plausible, but only if everything goes perfectly and the whole team works on it. It would be safer if we reduce the scope.” Or,
- “We almost certainly can, but I’d prefer to have 4-5 people on this just to be sure we make it.” Or,
- “Yes, definitely.”
Final thoughts
So ultimately, yes, the estimation is imprecise and often wrong. But that doesn’t mean that you can’t help the decision-makers by doing your best. If you ground your estimates in reality and honestly communicate the uncertainty level, you can give people useful information without pretending to know more than you do.
And that’s really the goal: not to predict the future perfectly, but to make better decisions with the information you have.