A paper published on July 2 has one of the best titles in AI research this year: "AI Writes Faster Than Humans Can Review."
The researchers — He, Agarwal, Denisov-Blanch, Azaletskiy, Koyejo, and Vasilescu — tracked 802 developers and 196,212 pull requests at a mid-sized, AI-forward enterprise that mandated doubling engineering output through AI tools starting mid-2025. Over 28 months, per-capita throughput hit 2.09x the pre-mandate baseline by April 2026. Among the largest field-deployment gains ever reported.
Then the other number: per-reviewer load roughly doubled too. Automated review overtook human review. The enterprise didn’t eliminate the bottleneck. It transferred it.
The name for what happened
Annie Vella and Kelly Blincoe, in a longitudinal study of 158 professional developers across 28 countries, gave the transformation a name: supervisory engineering work — “the direction, evaluation, and correction of AI output.”
82% of their participants reported spending less time writing code. The work didn’t disappear. It changed shape. Developers went from authors to supervisors — reviewing AI output, catching its mistakes, steering its direction, deciding what to keep.
Vella & Blincoe, 95 matched developers tracked over 6 months — arXiv:2605.23135
This is Vella and Blincoe’s Productivity-Experience Paradox. Among matched participants tracked over six months, productivity perception held rock-steady. But the proportion reporting worsened developer experience in at least one dimension — flow state, cognitive load — nearly doubled. The gains were real. The cost was real. Both were true simultaneously.
The verification gap underneath
Sonar’s 2026 survey of 1,100+ developers quantifies the toil: 96% don’t fully trust AI-generated code, yet only 48% always verify it before committing. AI now accounts for 42% of committed code. Developers spend 24% of their work week — nearly a full day — checking, fixing, and validating AI output.
Faros AI’s benchmarks across 10,000+ developers confirm the pattern at scale: teams with high AI adoption completed 21% more tasks and merged 98% more pull requests, but PR review time increased 91% and average PR size grew 154%. More output. More supervision. The ratio holds.
Why some orgs escape
Not everyone drowns. Stanford research projects top-quartile AI adopters at +260% productivity, bottom-quartile at +20% — a 10x gap from the same technology. DX Research, across 121,000 developers and 450+ companies, found that well-structured organizations saw 50% fewer incidents while struggling organizations saw twice as many. Same tools. Opposite outcomes.
The pattern: organizations whose processes were already supervisory-oriented — code review as a first-class activity, structured quality gates, clear verification ownership — absorbed the shift. The tools amplified what was already there. Organizations built around individual heroics discovered that AI just produced more to heroically review.
What this isn’t
I’ve written sixteen articles about generation scaling faster than verification. This is a different claim. Generation versus verification names what happens to code. Supervisory engineering names what happened to developers.
The enterprise in He et al. achieved the 2x. It worked. Throughput doubled. But per-reviewer load also doubled, and automated review overtook human review. They solved the throughput problem by creating a supervision problem. When they automated the supervision, they built exactly the recursive verification layer that creates its own unverified layer.
AI turned every developer into a supervisor. The tools that accelerate coding cannot accelerate supervision, because supervision requires the one thing that doesn’t scale with volume: judgment.