Analyzed document · Engineering

Measuring Developer Productivity in the Generative AI Era: What's Still Unsolved

In 2023, McKinsey published an article that’s still cited in every debate about “can software developer productivity be measured at all”. It emphasis a multi-level hierarchy of system, engineering team and individual metrics like DORA, backed by the promise that the old “software is a black box” excuse no longer holds.

Two years in, and with AI agentic engineering impacting on developers ways of working faster than any measurment approach can track and adapt to it, that promise is worth checking against what has actually shipped.

We ran the full article through https://app.assay.it. Eleven hypotheses and four problems extracted, each verified independently against public sources. Seven hypotheses came back supported, four contested — and all four problems the run surfaced remain open.

The metric hierarchy and methodology itself holds true. It is established practice, not a new insight. None of the article claims came back contradicted. But the harder question of the article, the one the article’s title asks, cannot be resolved. Measuring whether remote work and AI tooling investment are actually paying off is still, per every source the run could find, an unsolved problem if we look on this holistically from the organization perspective.

Causal releation between day-to-day work signals and delivery outcomes is still “a vision and roadmap,” not a shipping tool. Existing frameworks such as DORA and SPACE remain partial. And the fastest-moving part of the picture is generative AI’s and its actual effect on the businesses. Existing litirature is least equipped with an answer how to measure the return on investements.

Two years since publication, the field has better vocabulary for the problem than it had in 2023. It still does not have a common answer.

Your first analysis is on us — try the app on your own document.

The analysis

Yes, you can measure software developer productivity

Industry report · McKinsey & Company · analyzed

Hypotheses · 11
none novel
Problems · 4
4 open gaps
Holds — supported / solved Contested Fails — contradicted / open gap Unverifiable Novel — no prior art

The question the paper's title asks, still open two years later

Critical gap — unresolved two years after the original 2023 piece
“Absence of reliable productivity metrics prevents confident validation of remote work policies and AI tooling investments.”
Evidence

GetInt.io lists DORA, SPACE, and flow frameworks but "does not claim one is definitive." Jellyfish "acknowledges multiple frameworks coexist, implying none is universally sufficient." Wikipedia - Software Metric: "attempts at complete measurement create unintended negative side effects."

Analysis

The Brief confirms this directly. DORA, SPACE, and other frameworks are partial tools. The Brief's own summary states: "Multiple competing frameworks exist, but none is described as eliminating the core opacity problem."

All findings

Unstable Measurement Model in Changing Software Development

A critical gap — no tool in the brief bridges individual, team, and organizational measurement into one stable model that also accounts for unmeasured GenAI impact.

Critical gap

Hidden Drivers of Deployment Frequency

A critical gap — no mature tool automates the causal link between work metrics and deployment frequency changes, and causal software engineering tooling remains an unfinished roadmap, not a shipping product.

Critical gap

Context-Dependent Productivity Defeats Uniform Measurement

A critical gap — no off-the-shelf tool resolves cross-context productivity measurement, and the closest workarounds (Emerald hybrid principles, OECD research projects) remain fragmented and research-phase.

Critical gap

Simple Metrics Create Unintended Consequences

Metric gaming when a measure becomes a target is a fully established mechanism under Goodhart's Law since the 1970s, with no contradicting evidence found, making the software-context framing the only element not already named in prior literature.

Supported

Inner-Outer Loop Time Allocation

A fully established mechanism for developer productivity and satisfaction, but the claim's unqualified form is contested because narrowly maximizing inner-loop time can harm system-level delivery outcomes, a condition the claim does not state.

Contested

Capability Mapping Targets Talent Development

Assessing skills against benchmarks and using personalized learning to advance developer proficiency is fully established, but the specific claim that 30% of developers will advance one level in six months has no direct empirical support in the evidence.

Contested

…plus 8 more hypotheses, all 30 recorded entities, and every citation, in the full dossier

Including the assumption lists behind each hypothesis and the fourth open problem not covered above.

Continue to the complete report

How this analysis was produced. Assay extracted every substantive claim from the source document, then ran an independent research pass per claim against public sources — vendor docs, independent benchmarks, papers, postmortems. 15 claims · 74 sources · analyzed 22 Aug 2026. All four problems the run extracted resolved to a critical gap, not a solved one. The opacity gap tied to remote work and AI tooling investment is the one the article's own title promises to close, and it is still open: two years on, the run found no source — including McKinsey's own later coverage — that closes it. Verdicts are evidence-backed judgments, not oracles; every citation is linked so you can check the checker. How Assay works →

Your first analysis is on us — try the app on your own document.

Paste the link, tell us who you are, and we'll run it through the same pipeline. No card, no subscription.

assay.it logo assay.it

Paste the document your decision depends on. Assay checks every claim it makes against public sources — verdicts and citations in minutes.

© 2026 assay.it. All rights reserved.