| Audience: | CTO đ CIO đ VP of Engineering/Director of Application Development |
| Sector Applicability: | Cross-industry |
| Decision Horizon: | Before using successful releases to justify broader agent autonomy, recurring savings, lower maintenance funding, or engineering-capacity reductions |
Executive Summary
AI-assisted code may be faster to produce, and current academic evidence does not establish that it is inherently less maintainable than human-authored code. What the evidence does show is that functional success, static quality, and downstream maintainability are different questions answered at different points in the software lifecycle.
Decision Posture: Make release success inadmissible as evidence of durable maintainability. Deployment may proceed under existing controls, but proposals to expand agent autonomy, recognise recurring savings, or reduce maintenance capacity should receive conditional approval only when they contain relevant post-release evidence from the affected software environment.
Our Analysis
Every AI-assisted change does not quietly damage the codebase. But that does not free management to award AI a quality gain before the organization has observed the part of the lifecycle where maintainability is actually tested.
The Narrative vs. the Reality
The market narrative treats working code as increasingly measurable: the task is completed, tests pass, static checks are clean, and the change clears review. Those signals matter, but they establish immediate correctness against known conditions, not how readily another developer can understand, modify, extend, or recover the code later.
Direct comparative evidence on downstream maintainability remains narrow. The strongest cited study tests one subsequent evolution task; the remaining research primarily identifies observational patterns, measurement limits, structural differences, and possible debt pathways. The studies also examine different development modes and use different methods to identify AI involvement. Human-led assistance, agent-led generation, and commits detected through AI attribution should not be assumed to carry the same oversight or maintenance profile.
Direct downstream evidence
- AI-assisted development can accelerate the initial change without improving or degrading the next one. A preregistered, peer-reviewed experiment involving 151 participants (95% of them professional developers) found a 30.7% median reduction in initial completion time among developers using AI assistance. When different developers later evolved that code without AI, however, the researchers found no significant difference in completion time or code quality and no systematic maintainability advantage or disadvantage.1
Observational maintenance evidence
- Agent-generated files may genuinely require less maintenance. An accepted EASE 2026 study compared 508 agent-generated files with 508 human-generated files across 100 repositories. The agent-generated files received fewer and smaller subsequent changes during the observation period. That may indicate lower maintenance requirements, although the six-month observational design could not fully control for differences in file importance, use, complexity, age, or developer behaviour.2
- Some issues introduced by commits identified through explicit AI attribution persist after the original change. A large open-access preprint tracked more than 302,000 such commits and reported that some detected issues remained in later repository versions. Its broader results were mixed: the commits could also remove existing maintainability issues. The study has no matched human-authored baseline and should be read as evidence that AI-associated issues can persist, not that AI-produced code inevitably ages worse.3
Scope and mechanism evidence
- Accuracy benchmarks do not cover the entire quality question. An ICSE-NIER study introduced a benchmark specifically because conventional LLM coding evaluations concentrate on functional accuracy while overlooking code smells. Its case study found that the evaluated models could produce maintainability-related smells even when the generation task itself was completed.4 This establishes a missing evaluation dimension, not a verdict that AI-produced code is generally worse.
- Agent-led changes have distinct structural characteristics, but their durability remains unresolved. A 2026 study of more than 24,000 merged agentic pull requests found differences from human-generated changes, including generally smaller and more localized modifications. Results varied by agent, and the study did not establish whether those differences produced better or worse downstream maintainability.5
- Some unfinished assurance is knowingly deferred. A TechDebt 2026 paper found 81 self-admitted technical-debt comments among 6,540 comments referring to LLM use. The cases commonly involved postponed testing, incomplete adaptation, or limited understanding of generated code. The low count cannot establish prevalence, but it illustrates a pathway by which apparently completed work can transfer unfinished assurance into later stages.6
Meanwhile, the release dashboard can be correct and still be too early to answer the question management is asking.
What Changes the Decision
Treat release admission and durability recognition as separate management decisions. A change may be safe enough to deploy without yet providing enough evidence to claim that AI improved long-term quality or reduced the engineering capacity required to support it.
The rule should follow the softwareâs expected lifespan, frequency of change, coupling, and recovery cost, not its industry, its declared authorship, or the method used to identify AI involvement.
Ordinary testing and qualified review remain necessary release controls. They cannot establish how the software will respond to requirements, incidents, dependency changes, or architectural pressure that have not yet occurred.
Why This Matters Now
The strongest controlled evidence shows the timing asymmetry: initial implementation can accelerate while downstream maintainability remains neutral or uncertain. That study collected its data in late 2024, before more autonomous coding agents became the operating model now under consideration. Newer agentic studies describe how these tools modify code, but longitudinal evidence about how those changes age remains limited.1,5
That creates room for management conclusionsâtool expansion, greater agent autonomy, delivery commitments, maintenance cuts, or workforce assumptions, to run ahead of the evidence. Those decisions become difficult to reverse after support capacity has been removed or agent-generated change has spread through software expected to survive repeated modification.
The issue becomes material wherever software is long-lived, highly coupled, frequently changed, difficult to recover, or operationally consequential. Disposable prototypes, temporary analysis code, and low-consequence utilities do not justify the same evidence burden unless they are promoted into supported production use.
What to Watch for Next
Look for longitudinal studies that compare matched human-assisted and agent-led changes through multiple modđifications, incidents, dependency updates, and architectural transitions. Until those results mature, local service evidence matters more than broad claims that AI-produced code is inherently easier or harder to maintain.
Recommended Actions
Do This
- Separate release approval from management credit. When a business case uses AI delivery results to request wider agent permissions, recurring savings, reduced maintenance funding, or engineering-capacity cuts, the CTO or VP of Engineering must classify the supporting evidence as either release evidence or durability evidence in the existing decision paper. If only release evidence exists, the CIO should remove the durability benefit from the business case and withhold approval for the associated capacity reduction or autonomy expansion.
- Require independent approval of the evidence horizon. The application or service owner should propose the normal lifecyclđe event that will make maintainability observable: a material modification, incident-driven repair, platform change, dependency upgrade, or recurring change pattern. Engineering Excellence or Enterprise Architecture must independently approve that event and the comparison baseline before the evidence can support a savings, capacity, or autonomy decision. The function requesting the benefit may not approve the test it must satisfy; unresolved disagreements go to the CTO, who either strengthens the evidence requirement or removes the durability benefit.
- Hold agent autonomy at its evidenced scope while durability remains unresolved. Continue ordinary AI assistance and deployment under existing controls, but do not expand the class, size, coupling, or operational consequence of agent-led changes in long-lived software based only on successful releases. Expansion can proceed when the approved lifecycle event produces evidence consistent with the service baseline.
- Withhold the claim when evidence would cost more than the decision is worth. Use existing repository, CI/CD, incident, rollback, and engineering-flow telemetry wherever possible. If proving durability would require a new enterprise measurement programme for a low-value change, allow the change under normal controls but do not credit it as evidence for wider AI scale or reduced support capacity. The kill condition is not âstop using AIâ; it is âstop making claims the evidence cannot support.â
Avoid This
- Replacing a defect-count dashboard with a code-smell dashboard and calling the problem solved. Static indicators can identify structural concerns, but without a service baseline and later change evidence, they do not reveal whether the code is comparatively harder to maintain.
- Inferring quality from a quiet repository. Low churn may indicate stable code, limited use, avoidance, or a component that has not yet faced meaningful change.
- Turning uncertain durability evidence into either a blanket AI restriction or another release review. Review establishes whether the current change is acceptable; it cannot establish how the software will behave under future modification.
Bottom Line
AI-assisted code may be better, worse, or no different to maintain. A successful release cannot tell you which. Approve deployment on release evidence. Approve scale, savings, autonomy, and capacity reductions only on durability evidence.
Evidence and Sources
- Borg, Markus, Dave Hewett, Nadim Hagatulah, Noric Couderc, Emma Söderberg, Donald Graham, Uttam Kini, and Dave Farley. 2026. âEchoes of AI: Investigating the Downstream Effects of AI Assistants on Software Maintainability.âEmpirical Software Engineering 31:161. The study was preregistered and peer-reviewed. Its strongest contribution here is the controlled downstream-maintenance comparison, not a universal enterprise benchmark.
- Sawada, Shota, Tatsuya Shirai, Yutaro Kashiwa, Kenâichi Yamaguchi, Hiroshi Iwata, and Hajimu Iida. 2026. âTo What Extent Does Agent-Generated Code Require Maintenance? An Empirical Study.â Accepted short paper, 30th International Conference on Evaluation and Assessment in Software Engineering. The study covers six months, 1,016 files, and 100 open-source repositories; the authors identify uncontrolled file complexity, importance, usage, and incomplete AI attribution as limitations.
- Liu, Yue, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, and David Lo. 2026. âDebt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild.â arXiv preprint, version 2, April 26, 2026. The study uses explicit Git metadata and static-analysis attribution across open-source repositories. It is large and reproducible but remains a preprint and lacks a matched human-authored control group.
- Velasco, Alejandro, Daniel RodrĂguez-CĂĄrdenas, Luftar Rahman Alif, David N. Palacio, and Denys Poshyvanyk. 2025. âHow Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study.â In 2025 IEEE/ACM 47th International Conference on Software Engineering: New Ideas and Emerging Results, 96â100. The study demonstrates a missing evaluation dimension using two models; it does not establish comparative enterprise maintainability.
- Ogenrwot, Daniel, and John Businge. 2026. âHow AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests.â Accepted at the 23rd IEEE/ACM International Conference on Mining Software Repositories, Mining Challenge Track. The study characterizes structural differences between agentic and human pull requests but does not evaluate their downstream maintainability.
- Al Mujahid, Abdullah, and Mia Mohammad Imran. 2026. ââTODO: Fix the Mess Gemini Createdâ: Towards Understanding GenAI-Induced Self-Admitted Technical Debt.â Technical paper accepted at the 9th International Conference on Technical Debt. The study identifies a small, explicit subset of self-admitted debt and should be used to illustrate debt pathways, not estimate their prevalence.