
AI made your team faster, then velocity went flat again. Story points re-anchor to the team's own speed — here's why the gain vanishes and how to fix it.
Two numbers that shouldn't be able to coexist.
An engineering leader I've been trading notes with reports 3.5x capacity growth on his teams since they moved to AI-assisted development. Our own field study, run across a 250-engineer organisation, measured a 14% gain in implementation time - and only about 6% of it surviving to delivery once review and rework had taken their cut.
3.5x and 6%. Both from people looking hard at their own teams. Both, as far as I can tell, honest.
I spent a while assuming one of us had to be wrong. The answer turned out to be more interesting than that. We were counting in different currencies, and one of them inflates.
A story point isn't a unit of work. It's a unit of effort as this team currently experiences it, calibrated against what this team currently gets done. That calibration is the whole point - it's why points aren't comparable across teams, and everyone knows that part.
What gets less attention is the other consequence: if the unit is calibrated against the team's own throughput, then when the team gets faster, the unit quietly shrinks with them.
Here's the shape of it.
A team runs at 40 points a sprint. That's their baseline, built up over a year of estimating against each other's memory of how long things take.
AI lands. For a sprint or two they hit 46 against the old estimates, because the estimates were formed before the tooling changed. The gain shows up on the chart. Someone screenshots it for the leadership deck.
Then comes refinement. Someone says the obvious, sensible thing: "we've been crushing these - that's not really a 5, it's a 3." Nobody is being dishonest. They're doing exactly what estimation is supposed to do, which is stay calibrated to reality.
Next sprint, the same body of work is estimated at 40 and delivered at 40.
The chart reads 40 → 46 → 40. It says the gain came and went. Meanwhile the team is permanently faster and shipping more features than it was a quarter ago.
It's a scale that recalibrates to call your new weight normal.
When I put this to the leader reporting 3.5x, his answer was better than my objection.
He coaches his teams not to re-frame points because AI is doing the work. The complexity is still there. The effort is still there. The only thing that changed is who - or what - performs it. So refinement is still done on the basis of a developer doing the work, even if it gets handed to an agent afterwards.
That's not a workaround. It's a deliberate choice about what the unit measures, and it changes the meaning of the whole chart. By holding the human reference fixed, points stop tracking elapsed time and start tracking human-equivalent effort delivered. The unit stops moving, so the gain has somewhere to show up.
Economists have a name for this. Nominal figures are measured in the currency of the day. Real figures are measured in constant dollars, deflated to a fixed base year, so you can tell whether things actually got bigger or the money just got smaller.
Almost everyone reporting velocity gains from AI is reporting nominal velocity in a currency that inflates. He kept the deflator constant. That's why his number can be 3.5x while an org-level cycle-time measurement lands at 14% gross and 6% net - the two aren't in competition, they're denominated differently.
Holding the unit fixed depends on a reference class that erodes.
"How long would a developer take to build this" works as a shared anchor for exactly as long as the people in the room remember building that kind of thing by hand. Two years from now, a meaningful share of any refinement session won't have hand-written the thing they're estimating. The reference stops being memory and becomes folklore - and folklore drifts, usually in whatever direction makes the current sprint look reasonable.
I don't know how long the anchor holds. Neither, when I asked, did he. Strong refinement standards clearly slow the drift. Whether they stop it is an open question, and it's the one I'd most like to see someone measure.
The reason our own study reports in time rather than points isn't methodological purity. It's that we needed a unit that couldn't do this.
We decompose the delivery timeline - task started, branch created, implementation, merge request opened, review and rework, merge - and measure the duration of each segment. An hour is an hour whether or not the team is fast. The base year never changes, because there isn't one.
That's a narrower instrument in most respects. It tells you where time goes, not what the work was worth. But it does have the property that the number can't quietly re-baseline against itself while you're watching it.
If someone shows you a chart demonstrating that AI gains arrived and then faded, the first question isn't whether the tooling underdelivered. It's what the denominator was doing while the numerator moved.
Three ways to check:
None of this makes points a bad tool. It makes them a tool with a known failure mode that gets much worse in a period of rapid capability change - which is precisely the period everyone is currently trying to measure with them.
The teams still getting clean signal from points are, almost without exception, the ones who decided in advance what the unit meant and refused to let it move.