What "software estimation techniques" actually means (and what it doesn't)
Software estimation techniques size the work: how big is this task, this sprint, this backlog. They don't set a price, and they don't count how many hours a team has free this week. Those are separate problems with their own methods, and a lot of what ranks for this term collapses all three into one guide, which is usually a sign the page is really a vendor's pricing pitch wearing an estimation-technique title.
Sizing sits next to forecasting and delivery metrics in HighCircl's Engineering management guides for CTOs and VPs of Engineering. Seven techniques cover how teams put a size or a probability on their own work: story points, planning poker, t-shirt sizing, PERT, Monte Carlo simulation on throughput, reference-class forecasting, and the #NoEstimates position, which argues some of this shouldn't be attempted with a single number at all. If you came here wanting a dollar figure for a client project instead, pricing the whole project in dollars is a different article for a different question.
Relative-sizing techniques
These three don't try to predict time directly. They rank work against other work, which is a smaller, more answerable question.
Story points
A story point sizes effort relative to other work and says nothing directly about hours. Mountain Goat Software's definition is explicit about this: "It is the ratios that matter, not the actual numbers," so a story assigned two points should be roughly twice the effort of a one-point story. The scale itself, usually a Fibonacci-like sequence, exists to force a decision between two rough sizes instead of arguing over whether something is a 6 or a 7.
Ron Jeffries, credited as one of the people who popularized the technique, has published his own regret about how teams use it. On his site, he writes: "I like to say that I may have invented story points, and if I did, I'm sorry now." His objection targets what happens after the sizing: teams convert points to velocity, velocity to a date, and then hold the team to that date as though the estimate carried that kind of precision. He calls using story points to predict when a team will be done a weak idea, and argues that comparing teams by velocity or estimate quality does active harm. Use points to size work inside a team. Stop before using them to make a promise to someone outside it.
Planning poker
Planning poker is a collaborative technique agile teams use to estimate backlog items as a group, using private, simultaneous reveals instead of a public go-round. Mountain Goat Software's explanation of the mechanic lays out why the private step matters: "Each estimator thinks independently before seeing anyone else's answer. The simultaneous reveal also gives quieter team members a stronger voice." Without it, the first number spoken anchors everyone who answers after it, and a senior engineer's guess quietly becomes the team's guess.
The meeting time is worth it when the estimate genuinely benefits from more than one perspective, a cross-cutting story that touches parts of the system different people own, or a backlog where anchoring has been a recurring problem. It's not worth convening for a backlog of small, well-understood tickets; async sizing or a quick team call covers those without a formal round.
T-shirt sizing
T-shirt sizing swaps points for XS, S, M, L, and XL, trading precision for speed. It's the right tool exactly once: triaging a large, unscoped backlog before anyone has looked closely enough at any item to size it properly. Forcing a fresh, half-understood epic into a Fibonacci point value produces false precision; calling it an L and moving on is more honest about how little is actually known.
Where it breaks is sprint commitment. A sprint needs items sized precisely enough to know whether they fit inside the sprint's capacity, and "M" doesn't carry that information. Re-size to story points once an item is about to enter a sprint, and treat the t-shirt size as a triage label that expires the moment real planning starts.
Probabilistic and statistical techniques
These three treat uncertainty as the thing being measured, rather than something a single number quietly hides.
Three-point (PERT) estimation
Three-point estimation asks for three numbers instead of one: an optimistic estimate, a pessimistic estimate, and the most likely one. Wikipedia's entry on three-point estimation gives the standard PERT formula as E = (a + 4m + b) / 6, where a is the optimistic figure, m the most likely, and b the pessimistic one. Weighting the most-likely estimate four times over either extreme keeps the result from being dragged around by a single bad-day guess in either direction.
This is the same formula the cost-estimation piece above uses, pointed at a different scope. There, it prices a whole project. Here, it sizes a single task or a sprint's worth of work, still narrow enough that a team has real grounds for an optimistic and pessimistic bound instead of padding a guess with round numbers.
Monte Carlo simulation on throughput
Monte Carlo forecasting doesn't estimate individual backlog items at all. It runs against how many items a team actually finished per week or day, historically, and simulates thousands of possible futures from that record. Industrial Logic's one-line definition puts the idea simply: "If you combine daily throughput (the actual time it takes to complete work items) with thousands of Monte Carlo simulations, you get probabilistic forecasting, a way to predict date ranges for finished work."
Here's how the simulation works, step by step, with made-up numbers as an illustration. Say a team's last ten weeks of throughput looked like 3, 5, 4, 6, 3, 5, 7, 4, 5, 6 items finished per week. A simulation samples one of those ten values at random, subtracts it from the remaining backlog, and repeats week by week until the backlog hits zero, recording how many weeks that run took. Do that ten thousand times and you get ten thousand different finish weeks. Sort them, and the week by which 85 percent of the runs finished is your 85th-percentile delivery estimate. Swap this fictional throughput history for a team's real one and the mechanic produces a real forecast: a range with a stated confidence level, instead of a date someone will hold the team to regardless of what the data actually supports.
Throughput history is a flow metric, and a team that doesn't have it yet needs to start tracking before Monte Carlo has anything to sample from. Measuring delivery performance instead of point counts is a reasonable place to build that habit.
Reference-class forecasting
Reference-class forecasting starts from outside the item you're trying to size. Wikipedia's entry on the technique describes it as "a method of predicting the future by looking at similar past situations and their outcomes," with the underlying theory developed by Daniel Kahneman and Amos Tversky. Instead of estimating a new work item from first principles, a team finds a set of past items that resemble it closely enough, checks what those actually took, and uses that outcome as the estimate.
The point is countering optimism bias. An estimator looking only at the item in front of them tends to imagine the smooth path: no blocked dependency, no unexpected edge case, no code review round-trip. A reference class of five similar past stories, complete with however long they actually took, doesn't let that optimism through, because the data already includes the blocked dependency and the edge case from last time. Applying it takes some discipline built up front: tagging past work well enough that a comparable reference class is actually findable six months later, not just a backlog of closed tickets nobody labeled.
The #NoEstimates position
#NoEstimates isn't a rejection of planning. Woody Zuill and Vasco Duarte started making the case in blog posts and conference talks around 2011-2012, and the argument is narrower than the hashtag makes it sound. T2informatik's explainer on the idea's origin puts it precisely: "despite its name, the movement does not explicitly oppose all estimates of effort. It addresses projects that develop so unpredictably that useful estimates are substituted by random guesswork."
Reserve #NoEstimates for exactly that situation: a backlog too volatile, a codebase too unfamiliar, a domain too new for a point estimate to mean anything more than confidence theater. It's a poor fit for a stable team shipping predictable, well-understood work; there, an estimate earns its keep and a stakeholder asking for one is asking something reasonable. The case against dropping estimates entirely is mostly this: forecasting still needs some input, throughput history at minimum, and a team that drops estimates without picking up flow metrics to replace them has traded one guess for a different kind of silence.
Which technique fits which team
Team maturity decides more of this than any framework admits. A team with six sprints of real history has throughput data to run a Monte Carlo simulation against; a team on its first sprint together doesn't, and relative sizing or three-point estimates are the honest options until that history exists.
Sprint-based teams tend toward story points and planning poker because the ceremony fits the cadence: a backlog refinement session already exists to hold the poker round. Kanban and other flow-based teams, without a sprint boundary to size against, get more out of Monte Carlo on throughput, because it doesn't need a sized backlog item to start, only a count of what finished each week.
Who's asking for the number matters too. An internal planning conversation between engineers can run on relative sizes and shared context. A stakeholder who wants a date is really asking for a probability, whether they'd word it that way or not, and a percentile range from a Monte Carlo run answers that question more honestly than a story-point total converted to a date with a made-up velocity.
| Technique | Best for | Needs |
|---|---|---|
| Story points | Sprint-based teams comparing relative size | Shared team context, no history required |
| Planning poker | Reducing anchoring bias in a group estimate | A refinement meeting, private simultaneous reveal |
| T-shirt sizing | Coarse backlog triage before anything is scoped | Almost nothing; least precise on purpose |
| Three-point (PERT) | A single task needing a range instead of a guess | Three inputs: optimistic, likely, pessimistic |
| Monte Carlo on throughput | A delivery date range with a stated confidence | Historical throughput or cycle-time data |
| Reference-class forecasting | Countering optimism bias on a novel-looking item | A set of comparable past items and their outcomes |
| #NoEstimates | Work too volatile for a point estimate to mean anything | Flow metrics to replace the estimate you dropped |
None of this replaces knowing how many hours your team actually has before committing to a sprint's worth of sized work, and none of it means much if the team hasn't agreed on what counts as finished before you size it; a story sized as a 3 under one person's definition of done and a 5 under another's isn't the same unit at all.
FAQ
Are story points better than estimating in hours?
Better for what. Story points are relative by design, so a 2-point story should be twice the effort of a 1-point story regardless of how many hours either one takes, which makes points more stable across a team with mixed skill levels than hour estimates, which vary by who's doing the work. Hours earn their place when the audience genuinely needs a date and the team has enough history to convert points to time reliably. Pick whichever one the team actually has the supporting data for.
What is #NoEstimates and is it actually "no planning"?
It's often misread as a call to stop planning entirely. It targets one specific failure mode: a point estimate substituting for a real answer on a project too unpredictable for that estimate to mean anything, the situation Woody Zuill and Vasco Duarte were describing when they started making the case around 2011. A stable, well-understood backlog is still a fine candidate for story points or PERT; #NoEstimates is for the backlog too volatile for either to mean anything.
How does Monte Carlo forecasting work without a full backlog estimate?
It doesn't estimate the backlog items at all, and that's the point. The simulation runs against how many items the team actually finished per week or day historically, samples repeatedly from that distribution, and produces a spread of possible finish dates for however many items remain. No item needs a size on it anywhere in that process, only a count of items finished.
Do agile teams still need PERT or three-point estimation?
Sometimes, and less often than they used to. PERT earns its place when a team has genuine grounds for an optimistic and a pessimistic bound on a single task, a specific technical unknown, a piece of work novel enough that a flat guess would be dishonest. That's a narrower job than sprint-level story pointing, and a different scope than the whole-project pricing PERT gets used for in cost estimation: here it sizes one task, there it prices one project.
