Against Measuring Developer Productivity: The Calorie-Counting Problem
McKinsey published a framework for measuring individual developer productivity in August 2023, and the software engineering community delivered the same rebuttal it has delivered every decade since 1968 — most memorably from Tom DeMarco, who spent 1982 telling the industry you can't control what you can't measure, then spent 2009 recanting it in print. Counting a developer's output is exactly like counting calories: it treats all activity as equal and can't tell nutrient-dense work from junk. The newest wave of Goodhart's-Law casualties runs the identical experiment on GenAI tooling — measured in tokens instead of lines, with the same ending.
None of this means productivity is unmeasurable, or that accountability stops mattering — the ambition behind every one of these frameworks is real, and so is the underlying question. What leaders need isn't a sharper individual metric, though: it's to stop hunting for one and start measuring what fifty-eight years of evidence already show predicts whether an organization delivers — system-level delivery health, developer experience, and outcomes customers can actually feel. Protecting the motivation that makes any of that possible belongs on the same list, because a demoralized team optimizes whatever number is put in front of it, while a trusted one is the kind that ships something with no prior version to compare against.
In August 2023, McKinsey & Company published "Yes, you can measure software developer productivity" — a framework promising to quantify individual engineers through a blend of inner/outer-loop activity metrics and a capability rubric. CFOs and boards forwarded it around corporate America within days, because it seemed to finally answer the question every non-technical executive has been asking since the discipline existed: why can't we just measure developers the way we measure everyone else?
The backlash was immediate, specific about names, and — unusually for a critique this pointed — co-authored. Kent Beck and Gergely Orosz published a joint, 12,000-word rebuttal across Part 1 and Part 2, opening with as sharp a line as either has published in their own name: "the thinking that's evident in the article is absurdly naïve and ignores the dynamics of high-performing software engineering teams." Their complaint about the metrics themselves was just as specific — "nearly every single one of its custom metrics which differ from DORA and SPACE metrics, measure effort or output," not anything a customer, executive, or investor would recognize as value: "the only folks who care about these metrics are the people collecting them."
They didn't stop at diagnosis. The joint piece warns that "collecting & evaluating these metrics interferes with the team delivering on the measures downstream folks actually do care about, like profitability" — the framework doesn't just fail to help, it actively competes with the work that does. Dave Farley titled his video response, without hedging, "the Nonsense McKinsey Article." CodePulse, an engineering-analytics vendor, published its own guide to answering leadership pressure for exactly these metrics. Dan North reached for the analogy that outlasted the news cycle:
"Attempting to measure the individual contribution of a person is like trying to measure the individual contribution of a piston in an engine. The question itself makes no sense."
None of this was new, and that's the actual story. This is roughly the fifth time in fifty-eight years someone has tried to reduce software engineering to a single quantitative ruler, and the fifth time the field has produced an almost identical rebuttal — Fowler said as much in 2003, the metrics book's own author recanted it in 2009, and the pattern traces back to a 1968 study that found individual output varying by more than 20 to 1 before anyone even picked a metric. The section below titled "A History of Failed Productivity Metrics" has the full rap sheet.
What's actually new in 2026 is the unit under measurement, not the mistake.
GenAI vendors and the companies that bought their tools have spent two years measuring adoption by the token instead of the line, and a survey of 415 developers just documented the identical mechanism running one layer up: activity goes up, verification burden goes up more, and net productivity does not move.
This post covers why measuring individual developer output fails structurally, starting with the same reason a calorie count can't tell you whether you're healthy; what fifty-eight years of prior attempts already proved about it; and what to measure instead.
01 The Seduction and Fallacy of the Single Metric
The desire to measure developer productivity comes from a real managerial problem, not an invented one.
Software engineering is one of the largest line items in a modern enterprise's budget, and unlike a sales division (measured in revenue) or a recruiting team (measured in candidates placed), its output is opaque to the people funding it. A C-suite that already has a ruler for every other function naturally wants an equivalent one for engineering.
That ruler doesn't exist for a reason older than McKinsey's report. Martin Fowler made the underlying point in 2003, on what he calls his bliki — his term for a blog entry that gets revised in place like a wiki page instead of staying frozen at its publish date.
Nobody can measure developer productivity, he argued, because nobody can measure software output in the first place: manufacturing output is a countable physical unit, while software output is non-linear, context-dependent, and routinely negative in the metric meant to track it.
Three separate reasons software output specifically resists a formula:
- Output is proxy-driven. Lines of code, feature velocity, and commit counts are proxies for effort, not indicators of value.
- Less is often more. The best fix for a hard problem is often deleting thousands of lines of legacy code, simplifying an architectural boundary, or holding a twenty-minute whiteboard conversation that prevents six months of unnecessary work.
- Business impact lags. The value of a release may not show up for months or years, once the business processes around it adapt.
When an organization forces a simple quantitative formula onto that reality, it does not gain clarity. It gains distortion.
02 Calorie Counting and Software Engineering
In many ways, counting calories is exactly like measuring raw developer output — lines of code, story points, or tokens. Both approaches reduce a complex, qualitative system to a single trackable number, and both fail to capture health or value.
Same Calorie Count, Not the Same Food
When you count calories, you treat all inputs as equal, whether the food is nutrient-dense or junk. Measuring lines of code, commit counts, or completed story points does the same thing: it treats all activity as equal, confusing volume with value.
The comparison holds up with real numbers, not just an analogy. TGI Friday's published nutrition data puts its Million Dollar Cobb Salad with grilled chicken, ranch, and honey mustard at 1,220 calories — within 60 calories of a bacon cheeseburger and seasoned fries at 1,160. Nutrition coach Graeme Tomlinson, who posts as The Fitness Chef, built a whole genre of content around exactly this gap: the calorie count says these two orders are nearly interchangeable, and nothing else about them is.
That's not a contrarian take anymore — it's the federal government's own position.
The 2025–2030 Dietary Guidelines for Americans, released jointly by HHS and USDA, reoriented national nutrition advice around "real food" — whole, nutrient-dense, minimally processed — instead of a calorie target alone. Measuring an engineering org the way the old guidelines measured a diet has the same blind spot.
Brazil got there first, and went further. The Ministry of Health's 2014 Dietary Guidelines for the Brazilian Population didn't just de-emphasize calories — they dropped numeric nutrient targets from the recommendations entirely, replacing them with the NOVA classification, which sorts food by how processed it is rather than by what's in it. The golden rule: prefer natural or minimally processed foods over ultra-processed ones. No calorie icon, no portion count — a decade before the U.S. reached a version of the same conclusion.
The Metric That Registered Its Best Work as Negative
As Andy Hertzfeld recorded on Folklore.org, Apple ran into this exact failure mode decades before "spurious productivity" needed a name.
When Apple engineer Bill Atkinson spent a week rewriting QuickDraw's region-calculation routines to run six times faster, he also removed 2,000 lines of code.
The Lisa team's weekly form asked every engineer to report lines written that week. Atkinson entered -2000.
Management, confronted with a metric registering its best engineering work of the year as negative productivity, quietly stopped asking him to fill out the form.
The same reversal shows up wherever the metric rewards volume over judgment: an engineer who refactors 10,000 lines of bug-prone legacy code down to 400 clean lines has done valuable work that a naïve output tracker records as a collapse. The GenAI era runs the identical junk-food diet on a newer input — eating empty calories (tokens, in this case) to hit a number, bloating code and inflating review burden along the way, producing what researchers now call spurious productivity (the next section covers the token-specific version; the GenAI section below covers it in full).
Economics has a name for why this keeps happening no matter which proxy gets picked. Charles Goodhart described it in 1975 in the context of UK monetary policy; anthropologist Marilyn Strathern gave it the phrasing that stuck, in a 1997 paper about university audits:
"When a measure becomes a target, it ceases to be a good measure."
When management demands higher velocity or more story points, developers inflate their own estimates, split tasks into trivially small pull requests, or merge code without proper review just to hit the number. The dashboard looks fantastic. System health degrades underneath it.
Measuring Health, Not Just Weight
The alternative isn't a better formula — it's a different kind of looking. Checking your health means looking in the mirror, not stepping on a single scale. The engineering equivalent is measuring delivered business value, customer satisfaction, and team flow, together:
- Outcomes over outputs. Track whether shipped features increased sales, reduced support load, or made customers happier — not how fast they shipped.
- Holistic system health. Frameworks like SPACE and DX Core 4 look at the whole organization, balancing speed against quality, business impact, and developer well-being. Shipping fast while your change failure rate climbs isn't health.
- Team morale. If the metrics look terrible but the team is happy, you're probably measuring the wrong thing. If the metrics look great and the team is miserable, you have a real problem the dashboard can't see.
True measurement in software development starts with accepting that individual productivity cannot be captured by a simple formula — no more than fitness can be captured by a single number instead of diet, exercise, sleep, and genetics working together.
03 The Flaws of Output-Based Metrics: From Lines of Code to Tokenmaxxing
Management has cycled through proxy metrics for decades trying to quantify engineering activity. Every output-focused metric so far has failed to correlate with business outcomes while introducing its own way to be gamed.
| Metric | Intended Signal | Gaming Strategy | Consequences |
|---|---|---|---|
| Lines of Code | Volume of work done | Copy-paste, avoid refactors, expand formatting | Code bloat, higher maintenance cost, clean design penalized |
| Pull Request Count | Engineering velocity | Split simple tasks into trivially small PRs | Reviewer burden, elevated context switching |
| Cycle Time | Speed of execution | Prototype locally before opening the ticket | Hidden duration, bypassed review quality |
| AI Token Consumption | Tool adoption & leverage | Run automated scripts that burn tokens | Financial waste, code bloat, spurious productivity |
Function Points arrived to fix Lines of Code's obvious flaw by counting delivered functionality instead of raw text. They still cannot tell valuable functionality from bloat: a developer who spends a year shipping 100 function points nobody uses is far less productive than one who ships 20 that generate real revenue. The full history of that back-and-forth — and how each fix became the next decade's failure — gets its own section below.
The GenAI era has produced a direct descendant of the Lines-of-Code trap: tokenmaxxing. Organizations tracking the volume of AI tokens a developer consumes assume higher usage means more leverage. In practice, it rewards feeding unvetted text and prompt iterations into an LLM at volume — flooding codebases with unverified boilerplate, raising reviewer fatigue, and introducing subtle architectural bugs nobody asked for.
The term traveled fast once engineers noticed the pattern repeating. Dev Interrupted, the engineering newsletter run by LinearB, put it as bluntly as anyone: tokenmaxxing and its opposite, capping token spend, are both "optimizing the wrong half of the equation" — neither one asks whether the work shipped, stayed stable in production, or created value.
Even the executive running the highest-profile leaderboard agreed once the data came in. Meta's own CTO, Andrew Bosworth, summed it up in one line: "All motion is not progress and token usage alone is not a measure of impact of any kind."
Not everyone treated it as a permanent shift, either. Carmen Li, founder of the AI-infrastructure startup Silicon Data, called the whole leaderboard era "a phase thing" — a marketing tool dressed up as management, not something built to last. By July 2026, Forbes was reporting that token spend, not token volume, had become the number boards wanted — cost efficiency in, vanity leaderboard out, the same rotation the history section below tracks across the previous fifty-eight years.
04 Goodhart's Law and Perverse Incentives
Whenever an organization ties individual output metrics to compensation, promotion, or performance review, it switches Goodhart's Law on. Smart, rational engineers adapt fast — optimizing for the personal metric instead of the organizational outcome it was meant to stand in for.
In 2026, Meta ran an internal dashboard nicknamed "Claudeonomics" that ranked more than 85,000 employees by AI token consumption. It tracked 60 trillion tokens burned in thirty days; one engineer alone accounted for roughly $500,000 a month in usage. Employees who ranked low reported feeling exposed to layoff risk, so some set up idle agents to run all day purely to climb the board. The dashboard produced no measurable productivity gain, leaked externally, and was shut down within two days of the leak becoming public.
Meta wasn't alone.
Amazon ran a similar leaderboard pushed by an internal tool, producing "performative usage" — workflows run to signal adoption rather than to accomplish anything. It's a corporate-scale cargo cult: mimicking the form of AI-driven productivity without any of the substance that made the original worth imitating.
Unlike Meta, Amazon left the system running long enough for the behavior to calcify into a team norm. Duolingo's CEO went the other direction, publicly retracting his own viral "AI-first" memo once the same pattern showed up internally. Three companies, one failure mode, discovered independently within the same news cycle.
Every company that stood up a leaderboard ran the identical experiment and, with one exception, got the identical result — down to the same handful of names showing up in trade coverage within weeks of each other:
| Company | What Happened |
|---|---|
| Meta | Ran an internal dashboard nicknamed "Claudeonomics" that ranked more than 85,000 employees by AI token consumption, tracking 60 trillion tokens burned in thirty days. Produced no measurable productivity gain and was shut down within two days of leaking externally. |
| Amazon | Ran a similar leaderboard pushed by an internal tool, producing "performative usage" — workflows run to signal adoption rather than to accomplish anything. Left running long enough for the behavior to calcify into a team norm. |
| Microsoft | Ran its own token leaderboard from January 2026 with senior engineers and VPs at the top; one participant admitted inflating usage with needless lookups and prototypes nobody meant to ship. Canceled most of its Claude Code licenses within months. |
| Salesforce | Set minimum and maximum per-tool spend targets and a widget showing colleagues' usage updated every fifteen minutes; developers burned tokens on unrelated work just to clear the minimum. |
| Uber | Burned through its entire 2026 token budget in four months. COO Andrew Macdonald acknowledged the company couldn't connect the spend to any company-wide productivity gain. |
| Shopify | The exception: renamed its own 2025 leaderboard to a plain "usage dashboard," added circuit breakers to catch runaway agents, and had engineering leadership check top spenders' work directly instead of trusting the ranking. |
Shopify is the tell. The same tool, pointed at verification instead of ranking, produced none of the casualties above — the leaderboard wasn't the problem, treating it as a target instead of a dashboard was.
A parallel structural failure runs through promotion-driven cultures, sometimes called promomaxxing. At companies where promotion criteria have historically rewarded engineers who work on large, technically complex systems, clean design creates a direct conflict: build something simple and maintainable that doesn't read as promo-worthy, or manufacture complexity to justify the next level.
Engineers routinely resolve that conflict the rational-for-the-individual, value-destructive-for-the-business way: excessive documentation, unneeded microservices, bespoke frameworks where an off-the-shelf tool would suffice. Sales teams show the identical dynamic when measured on a single number — a consultancy case study found a top-performing salesman who won every internal award by booking large deals at unrealistic timelines and below-margin discounts. Every project he sold landed late, ran at a loss, and damaged the client relationship. Measuring developers on commits, PR velocity, or story points rewards the identical shortcut.
05 A History of Failed Productivity Metrics
None of the last four sections describes a new failure. Software engineering has run this exact experiment, with this exact outcome, roughly once a decade since 1968: pick a proxy → watch it get gamed or debunked → watch someone senior recant it in print → repeat with a new proxy.
The first study is also the one that should have ended the conversation. In 1968, Sackman et al. ran the first controlled comparisons of programmer performance and found individual variance so large — roughly 20 to 1 in coding time, over 25 to 1 in debugging time, across programmers doing the identical task — that a formula applied at the individual level was statistically hopeless before anyone had even chosen one. They also found no relationship between years of experience and code quality, which should have been the second reason nobody trusted a simple ruler.
The proxies kept coming anyway. Thomas McCabe's 1976 cyclomatic complexity measure was a genuinely useful diagnostic for testability — until organizations used it as a talent or output signal, at which point engineers flattened it artificially to hit a target. Allen Albrecht's Function Points, developed inside IBM in 1979 explicitly to fix Lines of Code's flaws, ran into the exact same wall it was supposed to solve: counting delivered functionality still can't distinguish functionality anyone wants from functionality nobody asked for.
Tom DeMarco gave the era its best-known slogan in his 1982 book Controlling Software Projects: "you can't control what you can't measure." That line licensed a generation of engineering metrics programs.
It took twelve years for the first serious empirical rebuttal: Perry, Staudenmayer, and Votta tracked engineers through their actual workdays in 1994 and traced the bulk of productivity variance to organizational and social factors — meetings, cross-team dependencies, interruptions — invisible to any metric attached to an individual's commits.
The paper's own calibration exhibit makes the point better than its statistics do: one engineer, one day, two logs.
Same developer, same day, two irreconcilable accounts of the work — and the mismatch is the observer's notes catching what the self-report couldn't. Any metric built on what engineers report doing inherits this gap by construction.
1976 also produced Campbell's Law, Donald Campbell's more general version of the same idea Goodhart articulated the year before: the more a quantitative social indicator gets used for decisions, the more it corrupts the process it was measuring.
Goodhart's Law explains why software metrics get gamed. Campbell's Law is the reminder that this was never a software-specific problem — it's what happens to any simple number attached to any human incentive.
Martin Fowler's 2003 bliki entry said the quiet part out loud: the field should stop trying, because a false measure is worse than no measure. Eleven years after that, Meyer et al. surveyed 379 developers and observed 11 more directly, and found that even developers' own felt sense of a productive day didn't line up with any output trace of that day — self-perception resists a formula as thoroughly as the underlying work does.
The most striking reversal belongs to DeMarco himself. On the fortieth anniversary of the NATO conference that named the field, he published "Software Engineering: An Idea Whose Time Has Come and Gone?" in IEEE Software — recanting his own 1982 book and writing that software projects are fundamentally experimental, not controllable:
"The more important goal is transformation: creating software that changes the world, or that transforms a company or how it does business."
By 2018, Nicole Forsgren, Jez Humble, and Gene Kim's Accelerate and the DORA research program behind it tried something the previous fifty years hadn't: a framework designed from the start to resist collapsing into one gameable number, by measuring the delivery system rather than the individual inside it. The SPACE framework extended the same idea to developer experience in 2021. McKinsey's 2023 report ignored that entire lineage and tried the individual ruler anyway — which is why the rebuttal it got back was, almost word for word, the rebuttal the field had already delivered in 2003, 2009, and 2014.
| Year | Milestone | What Happened |
|---|---|---|
| 1968 | Sackman, Erikson & Grant | 20:1 individual variance, no link to experience — a formula was already hopeless |
| 1976 | McCabe's cyclomatic complexity | Useful diagnostic, gamed the moment it became a target |
| 1979 | Albrecht's Function Points | Fixed LOC's flaw, inherited its inability to value functionality |
| 1982 | DeMarco, Controlling Software Projects | Licensed a generation of metrics programs |
| 1994 | Perry, Staudenmayer & Votta | Traced variance to organizational factors, invisible to output metrics |
| 2003 | Fowler, "Cannot Measure Productivity" | Declares the whole quest a dead end |
| 2009 | DeMarco recants, IEEE Software | The 1982 author renounces his own book |
| 2023 | McKinsey tries again | Same rebuttal, delivered by name |
The pattern in that table isn't a coincidence of bad luck striking the same idea eight times. It's what happens whenever an industry needs a proxy for something it can't directly see: the proxy gets adopted, gets gamed, gets debunked by the people closest to the work, and the debunking gets ignored by whoever wasn't in the room the first time.
McKinsey wasn't naïve so much as unlucky in its timing — arriving in year fifty-five of a fifty-eight-year cycle, with the rebuttal already written five separate times by five different people who had no reason to coordinate. The next attempt won't be the last, either, unless the measure changes from the person to the system around them.
06 GenAI and "Spurious Productivity"
GenAI coding tools escalated the push to measure productivity rather than settling it. Vendors cite headline numbers like Peng et al.'s GitHub Copilot randomized trial, where developers given the tool finished a scoped task 55.8% faster than a control group. That number is real, and it is also the initial-speed half of a two-part story the industry mostly skipped.
The second half arrived in 2026. Afroz et al. surveyed 415 professional developers using the SPACE framework and documented what they call spurious productivity: surface acceleration that obscures effort simply moving downstream rather than disappearing.
The study's numbers are specific enough to be uncomfortable. Frequent GenAI users reported far higher commit volume (48.3% vs. 7.9% for infrequent users), more test cases generated (56.5% vs. 24.8%), and more completed work items (55.9% vs. 9.1%). None of that showed up downstream: 67.4% of frequent users reported no improvement, or a decline, in test-pass rates, and 58.6% reported the same for how fast they learned new APIs. Most damning of all, 84.3% of frequent users said GenAI did not reduce the time they spent on code review — the flood of AI-generated boilerplate simply moved the cost onto whoever reviewed it next.
DORA's 2025 State of AI-Assisted Software Development report reaches the same conclusion from a different survey: AI is an amplifier, not a fix. Teams with strong existing delivery practices get faster. Teams with technical debt and process chaos get the same problems, just generated faster.
The report's own numbers make the case better than a paraphrase could. Drawing on more than 100 hours of qualitative research and nearly 5,000 survey responses, DORA found adoption is now nearly universal — 90% of respondents use AI, and more than 80% believe it has increased their productivity — but 30% report little to no trust in the code it generates.
For the first time in the 2025 report, the throughput number moved: AI adoption now measurably improves software delivery throughput. Delivery instability rose right along with it, which means teams are shipping faster on infrastructure that hasn't caught up to the new pace yet. DORA's own conclusion could double as a chapter title for this post: "Simple metrics are not enough." Its answer was to stop looking for one number and sort teams into seven profiles instead — from "harmonious high-achievers" to a "legacy bottleneck" cluster still servicing old technical debt — because the identical throughput number means something completely different depending on which profile produced it.
| Guardrail | What It Does |
|---|---|
| Surface uncertainty | Require developers to flag AI-assisted sections of a PR, with the verification steps taken |
| First-pass reviewing | Position GenAI as a preliminary checker, not a review-ready generator |
| Strict quality gates | Automated structural checks before a human reviews a large AI-heavy pull request |
| Rich architectural context | Feed the model project-specific constraints instead of generic boilerplate defaults |
07 The Galácticos Problem
Software engineering is an interdependent, team-based endeavor. Isolating individual productivity from team context — the piston-in-an-engine problem from the top of this post — is not just uncharitable, it's scientifically unsound.
The Galácticos Precedent
Sports proved this on a pitch, at nine figures a season, well before software had the vocabulary for it.
Real Madrid spent the early 2000s assembling the most expensive collection of individual talent soccer had ever seen — Figo, Zidane, Ronaldo, Beckham: one marquee superstar signed every single summer under president Florentino Pérez's Galácticos policy. The opening years delivered two league titles and a Champions League.
Then, for the back half of that same presidency, 2003 to 2006, the club won nothing at all — the longest trophy drought in its history, fielding a roster that individually outclassed almost anyone else on the planet.
Individual skill never automatically synthesized into team performance. Software engineering teams show the identical dynamic: a roster of brilliant individual "rockstars" routinely underperforms a cohesive team of mid-level developers with high trust and clear communication norms.
What the Research Says About Teams
Ice hockey's plus-minus statistic measures goal differential while a specific player is on the ice — a genuinely useful number, in a sport with discrete time limits, fixed 5-on-5 rules, and one scoring mechanism.
Software development spans months or years, involves non-linear tasks, and has no equivalent single scoring event to attach a plus-minus to.
Google's own multi-year Project Aristotle study set out to find what distinguished its highest-performing engineering teams, across 180 teams and roughly 250 measured attributes. Individual skill, seniority, and background all came out secondary. The single strongest predictor was psychological safety — the shared belief that a team is safe for interpersonal risk-taking.
Project Aristotle wasn't the last word on this, either — Google had already run one comparable study, and Microsoft has since run two of its own. Project Oxygen, Google's earlier investigation into whether managers mattered at all, found eight (later ten) specific manager behaviors that predicted team outcomes; Aristotle was its team-level sequel.
Microsoft's 2018 version — 37 interviews plus a 3,646-person survey — reached the identical conclusion from a different angle: technical skill, the study found, "is not the sign of greatness for an engineering manager." One engineer described exactly what happens when a manager substitutes a fixed definition for that judgment instead:
"I had a manager try to mold me to their definition of what a good engineer does, and I was probably working the hardest and yet my output was probably the least."
Microsoft asked the same question one level down, too. A separate study of 59 interviews across 13 divisions, by Li, Ko, and Zhu, asked engineers what actually makes a colleague great, and the 53 attributes that came back skew personal and relational, not quantitative — passionate, curious, self-aware, aligned with the team's actual goals. One manager summed up what happens when that alignment slips: "A mismatch of value... their number one goal is really to learn... you are paid because we are a business." None of the 53 attributes was a commit count.
Surveilling developers through individual commit counts, cycle times, or PR scores runs directly against that finding.
An ExpressVPN survey of 1,500 U.S. employers and 1,500 employees found half of workers already suspect they're being monitored without being told, and roughly half would consider quitting outright if that monitoring increased.
An organization cannot demand creativity, architectural risk-taking, and mentorship from engineers while tracking their every commit — those two asks cancel each other out.
08 What to Do Instead
Abandoning individual metrics does not mean abandoning accountability. It means shifting measurement from individual surveillance to systemic capability and business outcomes — five concrete moves, not one dashboard.
Peter Drucker named knowledge-worker productivity the central management challenge of the century, and his prescription runs through everything below: manage it "by objective and trust," not by direct observation of activity, because the worker usually knows more about the task than anyone watching.
Track System-Level Delivery Health
Rather than assessing individual developers, evaluate the health of the entire delivery pipeline with the evidence-backed DORA metrics:
| Signal | Metric | What It Measures |
|---|---|---|
| Throughput | Deployment Frequency | How often code reaches production |
| Throughput | Lead Time for Changes | Time from commit to deployment |
| Stability | Change Failure Rate | Share of deployments causing outages |
| Stability | Time to Restore Service | How fast an outage gets remediated |
DORA metrics work because they assess the capability of the system as a whole. If deployment lead time stalls, that points at pipeline bottlenecks, environments, or review protocols — not one developer's slackness.
Evaluate Holistic Developer Experience
Combine those pipeline metrics with multidimensional survey frameworks like SPACE and DevEx: satisfaction and well-being, performance, activity, communication and collaboration, efficiency and flow. None of these five resolve to one number by design — that's the point.
One of that framework's own co-authors watched the same mistake resurface in real time. In a 2025 fireside chat, SPACE co-author Brian Houck and Meta's Nachiappan Nagappan described executives, faced with a new AI rollout, reaching straight back for raw lines of code to measure its impact — the identical metric this post opened with. Their fix: evaluate AI as an interactive collaborator across velocity, volume, and quality together, and measure the system it operates in rather than its output faucet.
They pointed to Microsoft's Engineering Thrive program: it tracks total cycle time from an idea's conception through customer deployment, not commits logged along the way.
The other lever they kept returning to was toil: administrative overhead, bureaucratic ticketing, and meeting load tax innovation directly and drive burnout, and cutting that toil is a clearer, more durable win for AI tooling than shaving milliseconds off code generation.
A separate 2026 survey of nearly 3,000 developers reached a compatible conclusion from the ground up. Chen et al., working with data from BNY Mellon, combined 2,989 survey responses with 11 in-depth interviews and found six distinct productivity factors that split cleanly into short-term and long-term buckets — not one number, six.
Survey respondents disagreed sharply over whether their AI tools even helped, but the interviews converged on the same long-term factors regardless of that disagreement: technical expertise and ownership of the work predicted developer-perceived productivity better than any short-term output measure did. The paper's own framing lands almost exactly where this post does — a multifaceted approach is needed, because reducing AI's productivity impact to a single score throws away the factors that predict it.
Protect Developer Motivation
All else equal, a motivated developer outperforms an unmotivated one, and the mechanism connecting the two isn't folklore — it's fifty years of motivation research converging on the same answer. Self-determination theory identifies autonomy, competence, and relatedness as the three needs any intrinsic motivation depends on, and its founders were explicit about what threatens autonomy specifically: external monitoring and controlling incentives.
Motivation crowding theory gives the same finding an economist's name — the hidden costs of control — and documents it across more than twenty independent studies: raising external control over a task can reduce, not increase, the effort people put into it.
Software engineering has its own, more specific version of this literature. Graziotin et al. found that happier developers solve analytical problems better and ship more; a 2025 review of 44 studies and 16,000 engineers reached the same conclusion from the outcome side, tracing well-being — autonomy, competence, relationships, a sense of meaning — directly to engineering performance. A 2025 CHI study went further, interviewing 31 developers directly about the tools they touch every day, and found the same three needs at stake in something as mundane as which IDE or ticketing system they were allowed to use.
One participant, required onto tools their leadership team had mandated instead of the ones they already knew, described it plainly: it "felt like I was being told 'how' to do my job instead of being told to do my job."
Another put the direction of the effect even more bluntly: "the more restrictions I have on tools that I'm allowed to access, the more annoyed I feel, or the less likely I feel I'm able to do my work."
The paper names the failure mode outright — companies that nominally encourage autonomy while managers micromanage underneath it produce what the authors call fake autonomy.
That's the same gap this post keeps finding everywhere else: the dashboard says one thing, the incentive structure underneath it says another, and developers respond to the incentive, not the slogan.
Aggressive bean-counting attacks exactly the need this literature identifies as central. Individual commit tracking, PR-count dashboards, and token leaderboards all convert a developer's daily judgment calls into monitored, externally-scored events — the precise condition self-determination theory and motivation crowding theory both flag as corrosive. The workplace-surveillance data earlier in this post — developers who'd consider quitting over monitoring — isn't a separate phenomenon from the productivity story.
It's the same mechanism, measured on the way out the door instead of on the way to the ticket queue.
Treat Engineering Investment Like Capital Allocation
The honest answer to "how much should we invest in engineering" looks more like R&D exploration than cost-center optimization — and this isn't just an analogy, it's a formal discipline in both venture finance and internal innovation management. Real options reasoning, the finance theory Rita McGrath applied to entrepreneurial investment, treats each engineering bet as an option to expand rather than a fixed commitment: pay a small premium up front for the right, not the obligation, to invest further once the uncertainty resolves.
Eric Ries's innovation accounting gives that theory its operational form for a single team — track a small portfolio of leading indicators tied to validated learning instead of one vanity number, and let the next round of investment follow the evidence.
Google's own 70/20/10 resource-allocation model runs the same discipline at company scale: roughly seventy percent of engineering investment to the core business, twenty percent to adjacent bets, and ten percent to genuinely speculative ones. The split isn't symmetric with the returns — by most accounts, the bulk of long-term value comes back from that smallest, riskiest tier — which is exactly what a portfolio is supposed to show and a single quarterly output number never could.
Measure the Previously Impossible
Every metric this section has recommended so far — DORA's deployment frequency, SPACE's efficiency and flow, the R&D portfolio's early-traction signal — is still, underneath the sophistication, a faster horse, so to speak.
All of them answer some version of "are we doing the same work faster?" That question is comfortable, board-slide-friendly, and almost beside the point.
- An earlier post on this blog about software maintenance ran the numbers on exactly that comfortable question: issue-closure velocity up 8×, PR-merge velocity up 10×, after agentic adoption. Impressive, fully audited, and still the wrong headline — the same trap this entire post has spent seven sections dismantling, just wearing a nicer chart.
- Yet another earlier post drew the distinction that matters: the job isn't to rebuild the familiar slightly faster, it's to build the architectures everyone previously wrote off as science fiction. Riot Games running 4 million simulated matches a week for regression testing isn't a faster version of a QA process a human used to run. No human process like it ever existed to be faster than.
Apply that test to any productivity claim, including the ones in this post, before you believe it. A team shipping the same features ten times faster is still playing last decade's game, just faster.
The spicy take is this: A team shipping something with no prior version — a feature, an architecture, a scale of verification nobody could run by hand — is the only one that actually changed what's possible. Everything else is a nicer dashboard on the same ceiling.
09 Conclusions
McKinsey's claim that individual developer productivity reduces neatly to a proxy metric was not merely inaccurate. It offered non-technical leadership a false sense of control while licensing practices that actively damage the organizations that adopt them — and it is, at minimum, the fifth time in fifty-eight years that exact claim has been made and rebutted.
None of this is really a software story. Donald Campbell said the same thing about standardized tests a year after Goodhart said it about interest rates, and Tom DeMarco proved it about his own career: the man who told the industry in 1982 that you can't control what you can't measure spent 2009 recanting the book that made him famous, under his own name, in a peer-reviewed journal. A field doesn't get a cleaner demonstration than its own most-quoted advocate reversing himself in print.
Physics has a name for this same failure mode: the observer effect — the general principle that the act of measuring a system can change the state of the system being measured, whether that's a thermometer drawing heat out of the fluid it's reading or an apparatus disturbing whatever quantum state it's trying to pin down.
It's easy to conflate with Heisenberg's uncertainty principle, but the two aren't the same claim: uncertainty is a hard limit on what can be jointly known even under an ideal, non-disturbing measurement, while the observer effect is specifically about the disturbance the measurement itself introduces. Software doesn't need quantum mechanics to reproduce the pattern. Watch a developer's commits closely enough and commits stop being a byproduct of the work and become the work — the instrument doesn't just fail to capture the system, it becomes part of it, and rarely for the better.
Every proxy in this post fails the same way a calorie count fails: it treats all activity as equal and calls that equality objectivity. A calorie count can't tell nutrient-dense food from junk, and a commit count can't tell a real fix from padding. True software productivity is an emergent property of a healthy, well-aligned socio-technical system. It flourishes where trust is high, architectural boundaries are clear, cognitive load is low, and psychological safety is real — not where a dashboard says a number went up. Bean-counting doesn't just fail to capture that system; it actively degrades it, and the mechanism isn't abstract — it's the developer who feels surveilled instead of trusted, does worse work because of it, and starts looking for the exit.
Measure the system, not the piston.
Engineering leaders don't need a better individual metric. They need to stop looking for one, and start measuring what the last fifty-eight years already proved predicts whether an engineering organization delivers: system-level delivery health, developer experience, and outcomes that customers can feel.
References
- Gnanasambandam, C., Harrysson, M., Hussin, A., Keovichit, J., & Srivastava, S. (2023). Yes, You Can Measure Software Developer Productivity. McKinsey & Company. link
- Beck, K., & Orosz, G. (2023). Measuring Developer Productivity? A Response to McKinsey (Parts 1 & 2). The Pragmatic Engineer. part 1 · part 2
- Farley, D. (2023). My Response To The Nonsense McKinsey Article On Developer Productivity. Continuous Delivery. video
- North, D. (2023). McKinsey Developer Productivity Review. dannorth.net. link
- CodePulse. (2026). McKinsey Developer Productivity Metrics: How to Respond. link
- TGI Friday's. Nutritional Information. link
- Tomlinson, G. The Fitness Chef. link
- U.S. Department of Health and Human Services & U.S. Department of Agriculture. (2026). Dietary Guidelines for Americans, 2025–2030. fact sheet
- Hertzfeld, A. Negative 2000 Lines of Code. Folklore.org. link
- Engineer's Codex. (2026). Tokenmaxxing, Promomaxxing, and Misaligned Incentives in Tech. link
- Victorino Group. (2026). Three Companies, One Failure Mode: Goodhart's Law Comes for AI Adoption. link
- Fortune. (2026). A Meta Employee Created a Dashboard So Coworkers Can Compete to Be the Company's No. 1 AI Token User. link
- The Pragmatic Engineer. (2026). The Pulse: "Tokenmaxxing" as a Weird New Trend. link
- Fortune. (2026). Tokenmaxxing Is Over. It Was a Flawed Way to Measure a Company's ROI From AI. link
- Zigler, A., & Lloyd Pearson, B. (2026). Whether Tokenmaxxing or Tokenminimizing, You're Measuring the Wrong Thing. Dev Interrupted (LinearB). link
- Keary, T. (2026). After "Tokenmaxxing," Token Spend Has Become the New Metric to Watch. Forbes. link
- Sackman, H., Erikson, W. J., & Grant, E. E. (1968). Exploratory Experimental Studies Comparing Online and Offline Programming Performance. Communications of the ACM, 11(1), 3–11. link
- McCabe, T. J. (1976). A Complexity Measure. IEEE Transactions on Software Engineering, SE-2(4), 308–320. paper
- Albrecht, A. J. (1979). Measuring Application Development Productivity. Proceedings of the IBM Applications Development Symposium. link
- Goodhart, C. (1975). Problems of Monetary Management: The U.K. Experience; popularized by Strathern, M. (1997). "Improving Ratings": Audit in the British University System. European Review, 5(3), 305–321. paper
- Campbell, D. T. (1976). Assessing the Impact of Planned Social Change. Occasional Paper Series, Public Affairs Center, Dartmouth College. link
- DeMarco, T. (2009). Software Engineering: An Idea Whose Time Has Come and Gone? IEEE Software, 26(4). paper
- Perry, D. E., Staudenmayer, N. A., & Votta, L. G. (1994). People, Organizations, and Process Improvement. IEEE Software, 11(4), 36–45. paper
- Fowler, M. (2003). Cannot Measure Productivity. martinfowler.com. link
- Meyer, A. N., Fritz, T., Murphy, G. C., & Zimmermann, T. (2014). Software Developers' Perceptions of Productivity. Proceedings of FSE 2014. link
- Forsgren, N., Humble, J., & Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps. IT Revolution Press. link
- Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590. paper
- Afroz, S., Feng, Z., Menezes, T., Kimura, K., Trinkenreich, B., Steinmacher, I., & Sarma, A. (2026). The Fast and Spurious: Developer Productivity with GenAI. Proceedings of FSE 2026. arXiv:2510.24265. paper
- Chen, V., He, J., Williams, B., Valentino, J., & Talwalkar, A. (2026). Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants. arXiv preprint arXiv:2602.03593. paper
- Google re:Work. Understand Team Effectiveness (Project Aristotle). link
- Los Galácticos: Real Madrid's Greatest Team, 2000–2007. video
- Kalliamvakou, E., Bird, C., Zimmermann, T., Begel, A., DeLine, R., & German, D. M. (2018). What Makes a Great Manager of Software Engineers? IEEE Transactions on Software Engineering, 45(1), 87–106. paper
- Li, P. L., Ko, A. J., & Zhu, J. (2015). What Makes a Great Software Engineer? Proceedings of ICSE 2015, 700–710. paper
- ExpressVPN. (2026). Workplace Surveillance Trends in the U.S. link
- Forsgren, N., Storey, M. A., Maddila, C., Zimmermann, T., Houck, B., & Butler, J. (2021). The SPACE of Developer Productivity. Communications of the ACM, 64(6), 46–53. link
- Noda, A., Storey, M. A., Forsgren, N., & Greiler, M. (2023). DevEx: What Actually Drives Productivity. Communications of the ACM, 66(11), 44–49. link
- DORA. (2025). State of AI-Assisted Software Development Report. link
- Houck, B., & Nagappan, N. (2025). Beyond the Commit: A Fireside Chat on AI and Developer Productivity. DPE Summit. video
- Ryan, R. M., & Deci, E. L. (2000). Self-Determination Theory and the Facilitation of Intrinsic Motivation, Social Development, and Well-Being. American Psychologist, 55(1), 68–78. paper
- Frey, B. S., & Jegen, R. (2001). Motivation Crowding Theory: A Survey of Empirical Evidence. Journal of Economic Surveys, 15(5), 589–611. link
- Graziotin, D., Wang, X., & Abrahamsson, P. (2014). Happy Software Developers Solve Problems Better: Psychological Measurements in Empirical Software Engineering. PeerJ, 2, e289. paper
- Godliauskas, P., & Šmite, D. (2025). The Well-Being of Software Engineers: A Systematic Literature Review and a Theory. Empirical Software Engineering, 30(1). link
- Wong, N., Cheng, N., Oewel, B., Genuario, K. E., Stoeckl, S. E., Schueller, S. M., Ahmed, I., van der Hoek, A., & Reddy, M. (2025). "It's a Spectrum": Exploring Autonomy, Competence, and Relatedness in Software Development Processes and Tools. Proceedings of CHI 2025. link
- McGrath, R. G. (1999). Falling Forward: Real Options Reasoning and Entrepreneurial Failure. Academy of Management Review, 24(1), 13–30. link
- Ries, E. (2011). The Lean Startup. Crown Business. innovation accounting
- ITONICS. The 70-20-10 Rule of Innovation. link
- Drucker, P. F. (1999). Knowledge-Worker Productivity: The Biggest Challenge. California Management Review, 41(2), 79–94. link