analysis 4 min read

Instrument Failure

Instrument Failure

Harness surveyed 700 engineering practitioners and managers across five countries. The results contain a contradiction so clean it could be a logic textbook example.

89% of engineering leaders said their current metrics accurately reflect AI's impact on productivity.

94% of those same leaders acknowledged that critical factors — developer fatigue, code quality, technical debt — are missing from those frameworks.

Both numbers. Same survey. Same respondents.

What They Report
89% say metrics accurately reflect AI's impact
89% say productivity improved
88% say developer satisfaction increased
What They Admit
94% say metrics miss fatigue, quality, debt
81% say code review time increased
31% of developer time is invisible to metrics

This isn't a story about organizations not measuring well enough. It's about the instruments themselves failing — and organizations knowing it and proceeding anyway.

The Experiment That Couldn't Run

METR's developer productivity study was the gold standard: randomized controlled trial, experienced developers, real tasks, paid participation. The February 2026 update admitted the study design is no longer viable. 30-50% of developers refuse to submit tasks because they don't want to do them without AI. Recruitment at $50/hour couldn't overcome the resistance.

"We can no longer measure it reliably with this study design."

The measurement instrument was defeated by the adoption curve. When the tool becomes the environment, the control condition — a world without the tool — becomes unmaintainable. The experiment didn't fail from bad methodology. It failed because the thing it was measuring changed the conditions required to measure it.

What replaced it? Self-reports. METR's own May 2026 survey of 349 technical workers showed median self-reported productivity gains of 1.4-2x value, 3x speed. But METR's prior research had demonstrated that self-reports overestimate AI gains by 40 percentage points. The people closest to the measurement problem — METR's own staff — reported the lowest gains of any group surveyed.

The Metric That Decoupled

Wheeler's "Substrate Collapse" paper (arXiv 2606.20882) identifies a deeper failure. Every knowledge metric in software engineering — truck factor, Degree-of-Authorship, code ownership — relies on a foundational assumption: committing code implies understanding it. When AI generates the code, version control still attributes authorship. But the attribution no longer licenses any conclusion about comprehension.

"The metric still returns a number, but that number measures a substrate uncoupled from the quantity it estimated."

The dashboard keeps updating. The number is real. What it measures is not.

And when metrics do exist, they contradict each other. Faros AI tracked 22,000 developers across 4,000+ teams for two years. The throughput metrics show gains: epics per developer up 66%, task throughput up 33.7%. The quality metrics show collapse: code churn up 861%, bugs per developer up 54%, review time up 441.5%. Same dataset, same teams, same period. The metrics don't disagree about a margin. They describe two different realities.

Vendor-reported benchmarks fare worse. A KDD 2026 study found Simpson's Paradox in agent merge rates — Devin's reported 33.5 percentage point advantage collapses to 1.6 points (p=0.73) once you control for repository selection. Statistically indistinguishable from zero. The aggregate looked decisive. The controlled number looked like noise.

The Organizations Know

What makes this different from ordinary measurement challenges: organizations have articulated the problem and are deciding inside it anyway.

GitLab's 2026 AI Accountability Report: 87% of respondents were confident they could trace AI-generated code in a production incident within 24 hours. Among those who actually had incidents, 34% could. A 53-point confidence-reality gap — visible only after something broke.

Sonar's 2026 survey: 96% of developers don't fully trust AI-generated code. Only 48% always verify it. Half of the code that developers themselves don't trust goes unchecked.

The Harness report names what's happening: "AI has been very good at increasing output. Simultaneously, it has not automatically delivered more shipped value." 31% of developer effort now goes to work invisible to existing metrics — reviewing AI output, debugging edge cases, managing parallel agents. Managers are nearly four times more likely than developers to report no concerns about the measurement systems. The dashboard is green. The people doing the work aren't sure.

Deciding Without Instruments

Enterprises will spend $2.5 trillion on AI in 2026. MIT research found 95% of enterprise AI pilots fail to deliver demonstrable ROI. Fewer than one in five organizations can identify specific financial outcomes from their AI investments. Boards got tired of productivity claims that never appeared in the operating result — "productivity" as the primary AI ROI metric fell from 23.8% to 18.0% this year, replaced by direct financial impact.

This isn't a gap that more data fills. The experimental method broke because developers won't work without the tools. Self-reports inflate by 40 points. Authorship metrics no longer measure authorship. Velocity and quality metrics contradict each other from the same dataset. Vendor benchmarks don't survive statistical controls. Organizational confidence collapses on contact with reality.

And organizations are making hiring, tooling, and investment decisions — billions worth — based on the numbers that remain. Not because they believe those numbers. Because they're the numbers that arrive.