This is the fourth article in the AI Contradictions series. It moves upstream to management itself: once AI adoption becomes a target, the evidence used to prove productivity can become part of the problem.
| Audience: | CIO đźž„ CTO đźž„ VP of Engineering |
| Primary Sectors: | Financial Services đźž„ Government/Public Sector |
| Decision Horizon: | Before the next performance-management, budget, licensing, or workforce-planning cycle |
Executive Summary
Mandated AI adoption can create genuine usage, and may contribute to genuine productivity gains, while making adoption itself less useful as evidence of those gains. Once employees are rewarded, ranked, assessed, or reorganized around AI use, the resulting increase partly measures compliance with management’s instruction.
Decision Posture. Mandate an AI evidence-separation rule. Treat seat activation, token consumption, prompt volume, AI-generated-code share, and usage frequency as enablement or compliance measures instead of evidence of productivity, ROI, labor savings, or readiness to scale.
Do not simply replace a usage metric with a single outcome metric. For consequential decisions, require an evidence bundle combining the intended workflow outcome with measures showing where cost, delay, rework, review effort, quality, or risk moved. The evidentiary burden should rise with the consequence of the claim. Keep controls light for reversible local use, but require stronger and independently testable evidence before approving recurring spend, booked savings, staffing reductions, or external claims.
Engineering should validate the technical measurement, the workflow owner should validate the operational effect, and Finance should determine whether the evidence supports the economic claim. Internal Audit should test the control, not routinely approve each productivity case.
Our Analysis
AI can improve software productivity. The harder question is whether an organization’s metrics can establish where, by how much, and under which operating conditions.
The Narrative vs The Reality
Several prominent organizations have begun managing AI adoption explicitly. Shopify made reflexive AI use a baseline expectation and later reported universal adoption of AI code editors. Coinbase shares AI-usage measures with engineering leaders and has used AI-generated-code share and token consumption as adoption indicators. DORA has documented internal token leaderboards intended to reward consumption and accelerate behavioral change.1,2,3 These are illustrative company practices, not evidence that every enterprise is following the same pattern.
The incentive mechanism is also visible in developer accounts. Engineers interviewed by 404 Media reported AI use entering performance criteria and encouraging performative participation despite doubts about output quality.4 These accounts are qualitative rather than representative, but they show how a measure can become partly self-fulfilling once evaluation or management pressure is attached to it.
The operational evidence is more complicated:
- Real gains do not make adoption a valid proxy for them. Three randomized field experiments involving 4,867 developers found a combined 26.08% increase in completed tasks among developers given access to an AI coding assistant. The result supports AI’s productivity potential; it does not establish that token counts, generated-code percentages, or frequency of use can measure that effect.5
- Average results conceal substantial variation. A 2026 meta-analysis of 23 studies found a moderate positive productivity effect overall, but considerable heterogeneity. Effects tended to be larger in controlled experiments and smaller in enterprise and open-source environments.6
- Usage and confidence are moving separately. Stack Overflow’s 2025 survey found that 84% of respondents were using or planning to use AI development tools and 51% of professional developers used them daily. Yet positive sentiment had fallen to 60%, while 46% distrusted the accuracy of AI output and only 33% trusted it.7 Adoption therefore cannot stand in for confidence, quality, or value.
- Self-report and preference are not substitutes for measured performance. METR’s randomized study of 16 experienced open-source developers estimated that participants took 19% longer with AI, despite expecting a 24% speedup and still believing afterward that AI had made them 20% faster. The uncertainty interval around the measured slowdown was wide (approximately 2% to 39%) so this is not a population-level verdict on AI coding.8 Its relevance is methodological: expected, perceived, and observed performance can point in different directions.
- Embedding AI can make controlled comparison harder. In METR’s 2026 follow-up, committed users increasingly declined AI-free work, developers withheld tasks they particularly wanted AI to perform, and concurrent agents made time-on-task measurement less reliable. METR believed newer tools probably offered greater acceleration but concluded that it could no longer estimate the size reliably.9 Revealed preference for AI may indicate usefulness or dependence; it still does not quantify productivity.
- Outcome metrics can conceal where the work moved. A July 2026 preprint examining an enterprise “2×” mandate found that pull-request throughput per developer eventually reached 2.09 times its pre-mandate baseline. But per-reviewer load also roughly doubled, automated review overtook human review, and substantive human review fell while silent approvals increased. Merge and revert rates remained stable, but the researchers cautioned that these were coarse, short-horizon quality proxies that did not capture incidents, maintainability, ownership, or technical debt.10
The 2Ă— study does not prove that the throughput gain was illusory. It demonstrates why a headline outcome cannot certify the whole business case: the production system changed around the target, review absorbed part of the added volume, and important downstream costs remained outside the measurement window.
Meanwhile, the more successfully management drives a measured behavior, the less that measure can independently prove management’s productivity claim.
What Changes the Decision
After AI use becomes a target, adoption telemetry should be classified as management-induced evidence. It can show whether employees complied, where tools are being explored, and which teams may need support—but it should not support scale funding, workforce reductions, or ROI claims.
Moving from usage to outcomes does not make measurement incorruptible. Decision-grade evidence requires an outcome, measures that reveal cost or risk displacement, an account of other changes that could explain the result, and a defined period in which the finding remains applicable.
The standard is not laboratory-grade causal proof for every local decision. It is evidence proportionate to what management intends to approve:
- Reversible use within an existing team and budget needs a named owner, bounded cost, basic quality monitoring, and a rollback path.
- Enterprise expansion or recurring spend needs a baseline, workflow outcome, balancing measures, and disclosure of material concurrent changes.
- Staffing reductions, booked savings, board reporting, or external claims require pre-defined metrics, reproducible inputs, explicit attribution limits, and cross-functional validation.
No function should certify the entire case alone. Engineering validates the technical measurement, the workflow owner validates the operational consequence, and Finance decides whether the evidence is sufficient for the proposed economic decision.
Why This Matters Now
Token leaderboards, baseline-use expectations, and other adoption incentives are emerging while CIOs are being asked to turn AI experimentation into budget commitments and labor assumptions. The first action is cross-industry—separate adoption from value evidence—but the assurance pathway changes by sector.
- In Financial Services, Engineering should establish the technical evidence, the relevant service or product owner should confirm the operational result, and Finance should approve any savings or staffing claim. Compliance should join when the conclusion enters board risk reporting, regulated operations, customer representations, or other formal institutional claims.
- In Government/Public Sector, the equivalent economic control sits with Finance or the comptroller function. Internal Audit should test whether departments consistently separate adoption, operational outcomes, and economic claims. License utilization and workforce participation should not be presented as program value in a funding submission without evidence proportionate to the claim.
What to Watch for Next
Watch for AI-generated-code share and token consumption moving from diagnostic dashboards into individual performance criteria. Also watch for organizations replacing those measures with one headline outcome—such as pull requests, cycle time, or cost per change—without measuring the downstream work that outcome creates.
Recommended Actions
Do This
- Apply the evidence rule in proportion to the decision. Keep local, reversible experimentation lightweight: require a named owner, bounded spend, basic quality or risk monitoring, and a rollback condition. Trigger the full evidence bundle only when the organization seeks enterprise scale, recurring funding, booked savings, staffing changes, or a material internal or external claim.
- Lock the measurement before the high-consequence evaluation begins. For scale, savings, or staffing cases, record the metric definition, baseline, source systems, extraction method, intended outcome, and at least one directly coupled and one lagging balancing measure. Log material changes in tooling, staffing, workflow, automation, or work mix. Engineering may generate and explain the evidence, but it should not redefine the measure after seeing the result. For the highest-consequence claims, make the underlying data reproducible and subject to independent sampling.
- Label the strength of the evidence—and limit the decision accordingly.
- Causal: A credible randomized or quasi-experimental comparison exists, measurement remained stable, and major confounds were controlled. This may support stronger financial claims within the evidenced scope.
- Strongly indicative: A pre-established baseline or comparison exists, the outcome and balancing measures move consistently, and no identified concurrent change plausibly explains most of the result. This may support expansion within the same workflow and operating conditions.
- Directional: The result is observational, self-reported, based on an uncontrolled before-and-after comparison, or materially confounded. It may support continued bounded use, but not booked savings or staffing reductions.
Avoid This
- Do not rank employees or teams by AI consumption. Token and usage leaderboards may temporarily identify active users or training needs, but tying them to performance rewards encourages gaming, unnecessary consumption, and performative adoption. Retire them before they become production theatre.
- Do not replace an adoption vanity metric with an outcome vanity metric—or impose enterprise assurance on every experiment. Pull-request volume, cycle time, accepted changes, and cost per release can also be distorted once funding, compensation, or staffing decisions depend on them. Require the full evidence bundle for consequential claims, not for every reversible use of an AI tool.
Bottom Line
An AI mandate may increase both adoption and productivity. It still cannot make a targeted metric proof of value. The evidence burden should rise with the consequence of the decision, and no function should grade the entire case alone.
Evidence and Sources
- Evan Conaway. 2026. “Finding Balance in the Era of Tokenmaxxing.” DORA, June 2.
- Daniel Beauchamp. 2025. “Serious Results, Unserious Methods: Shopify’s AI Playground.” Shopify, October 28.
- Kyle Cesmat and Chitra Venkatramani. 2025. “Tools for Developer Productivity at Coinbase.” Coinbase, August 6.
- Emanuel Maiberg. 2026. “Software Developers Say AI Is Rotting Their Brains.” 404 Media, May 13.
- Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. 2026. “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers.” Management Science, February 27.
- Sebastian Maier, Moritz Gunzenhäuser, Jonas Schweisthal, Manuel Schneider, and Stefan Feuerriegel. 2026. “A Meta-Analysis of the Effect of Generative AI on Productivity and Learning in Programming.” arXiv preprint, May 6.
- Stack Overflow. 2025. “AI: 2025 Stack Overflow Developer Survey.”
- Joel Becker, Nate Rush, Beth Barnes, and David Rein. 2025. “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.” METR, July 10.
- Joel Becker, Nate Rush, Tom Cunningham, David Rein, and Khalid Mahamud. 2026. “We Are Changing Our Developer Productivity Experiment Design.” METR, February 24.
- Hao He, Shyam Agarwal, Yegor Denisov-Blanch, Pavel Azaletskiy, Sanmi Koyejo, and Bogdan Vasilescu. 2026. “AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2× Mandate.” arXiv preprint, July 2.