AI Academy · Book
Executives & Directors · Module 05 · Chapter 008

AI in Software Engineering

AI has made writing code dramatically faster, but shipping working software has sped up far less. The gains reach customers only when verification, review and release keep pace with generation. Leaders decide what gets measured, where the extra checking capacity comes from, and which work an agent may finish on its own.

≈ 15 min read

After this chapter you can

  • Explain why large gains in writing code shrink before they reach shipped software, using the weak-link view of the delivery chain.
  • Describe how confidence in generated code is built in layers, and why compiling and generated tests are not enough.
  • Use DORA's delivery measures instead of code volume to judge an engineering AI program.
  • Grant coding agents autonomy by how well results can be verified automatically and what an escaped mistake costs.
  • Decide deliberately where recovered engineering capacity goes, starting with the checks themselves.

In May 2026, three economists published a large study of AI in software work. Mert Demirer, Leon Musolff and Liyuan Yang followed more than 500,000 developers on GitHub, combined with records of how much each one used AI coding tools. As the tools moved from autocomplete to interactive agents to agents that work on their own, the developers’ coding activity soared: commits rose by a cumulative 240 percent, well over three times the starting level. Then the gains thinned out. The number of projects those developers worked on rose by 80 percent. The number of releases, software actually shipped, rose by 30 percent. Across four large app marketplaces, the number of new apps jumped, but total use of them did not rise at all1.

With autonomous coding agents, commits rose 240 percent, projects 80 percent and releases only 30 percent; most of the speed-up in writing code did not reach shipped software.+240%CommitsCode written and saved+80%ProjectsPieces of work in progress+30%ReleasesSoftware actually shipped
Figure 5.8.1 The gain shrinks at every step toward the customer (Demirer, Musolff and Yang, NBER, 2026).

A second source points the same way. DORA, Google’s long-running DevOps Research and Assessment program (not to be confused with the EU financial regulation of the same name), surveys thousands of software professionals each year. Its 2024 report found that a 25 percent increase in AI adoption went with better documentation, better code quality and faster code review. It also went with an estimated 1.5 percent fall in delivery throughput and a 7.2 percent fall in delivery stability2. Teams felt faster and wrote better code, and their delivery got slower and less reliable.

The authors of the GitHub study call this the weak-link hypothesis: a production chain moves at the speed of its slowest step. AI has made one step much faster. The rest of the chain is still run by people, and that is where the executive work lies.

The goal is working software in use

Software engineering was never mainly typing. A change starts as an intent, a business need someone can state. It is designed, written, checked, reviewed, released, run in production and learned from when it fails. Writing the code is one link in that chain, and often not the longest.

A chain from intent through writing, verification and release to operation; speeding up writing alone moves the queue to verification and release.IntentA needsomeonecan stateWriteCode andtestsdraftedVerifyTests,scans,reviewReleaseIntoproductionOperateRun itand learnSpeed up one link and the queue moves to the next
Figure 5.8.2 Software delivery is a chain. AI has sped up the writing link far more than the others.

The idea of this chapter follows from the chain. AI in software engineering succeeds when business intent becomes reliable software in use faster. It does not succeed merely because more code gets written. Faster writing helps only if the links after it can absorb the extra volume, and when they cannot, it adds to the queue.

There is good news in the long-run research. The work behind Accelerate, by Nicole Forsgren, Jez Humble and Gene Kim, found that the best software organizations do not trade speed against stability. They release more often and fail less often, because they work in small changes, test automatically and get fast feedback3. Those are exactly the practices that decide whether AI-generated code turns into delivered value.

Fast at the task, slower at the system

The evidence that AI speeds up individual coding tasks is solid. In a 2023 controlled experiment, developers asked to build a small web server in JavaScript with an AI pair programmer took 55.8 percent less time than those without one: about 71 minutes on average, against about 1614. That figure measures time on one task. Field experiments measure something different, completed work, and are more modest and more telling. Three randomized trials at Microsoft, Accenture and an anonymous Fortune 100 company, covering 4,867 developers, found that access to a coding assistant raised completed tasks, measured as pull requests, by about 26 percent. Less experienced developers adopted the tool more and gained more5.

A table of three studies showing large gains on a single lab task, smaller gains in completed pull requests and modest gains in releases.StudySettingWhat was measuredResultPeng et al., 2023Lab, one new taskTime to finish55.8% less time (71 vs 161 min)Cui et al., 2026Three firms, 4,867 developersCompleted pull requestsAbout 26% moreDemirer et al., 2026500,000+ developersReleases shippedAbout 30% more
Figure 5.8.3 The closer a measure sits to the customer, the smaller the measured gain.

Read the three rows from top to bottom and the pattern is clear: the further the measure is from the keyboard and the closer it is to the customer, the smaller the gain. That is not a contradiction. A lab task has no reviewer, no integration with other teams and no release process. Real delivery has all three. And the gain is not universal: as What Does AI Value Actually Mean? showed, experienced developers working in code they knew well were slower with early-2025 tools6.

For a leader, the point is not to decide whether AI “works” for engineers. It plainly helps with many tasks. The point is to find where the time saved at the keyboard is being lost later in the chain.

Generate, then verify

The first place it is lost is verification. AI-generated code tends to look tidy and conventional, which makes its mistakes harder to see. Two studies show why that matters.

In a user study presented at a leading security conference in 2023, researchers at Stanford gave some participants an AI assistant for security-related coding tasks. Those with the assistant wrote significantly less secure code. They were also more likely to believe their code was secure. The participants who trusted the AI less produced fewer vulnerabilities7. The tool made the work faster and the people more confident, and it made the code worse.

The second study looked at a newer kind of error. Code relies on packages, ready-made libraries pulled in from public registries. When researchers generated 576,000 code samples with 16 popular models, a meaningful share of the packages the models recommended did not exist8.

Code models recommended non-existent packages at average rates of at least 5.2 percent for commercial models and 21.7 percent for open-source models, producing 205,474 unique invented names.5.2%Commercial modelsAverage share ofinvented packages21.7%Open-source modelsAverage share ofinvented packages205,474Unique invented namesEach one a name an attackercould registerSource: Spracklen et al., USENIX Security · 2025
Figure 5.8.4 A package that does not exist today can be registered by an attacker tomorrow. Checking dependencies is part of verification.

An invented package name is a gift to an attacker, who can publish malicious code under exactly that name and wait for someone to install it. This is a security problem, and AI in Cybersecurity, later in this module, takes it further. For engineering leaders the lesson is simpler: confidence in generated code has to be built in layers, and each layer catches something the one below it cannot.

Confidence in generated code is built in layers, from compiling at the bottom through tests, scans and human review to production monitoring.ProductionmonitoringWatch it runHuman reviewA named owner approvesSecurity anddependency scansKnown flaws, invented packagesAutomated testsDoes it behave as intendedIt compilesWell formed onlyTRUSTRISES
Figure 5.8.5 Compiling is the first check, not the last. Each layer catches what the one below misses.

Two cautions belong with the stack. Generated tests are not automatically good tests: a model can write a test that faithfully confirms its own mistake. And AI review may help human reviewers find problems faster, a capability with little published evidence of results so far; on systems that matter a named person still owns the decision to merge.

The bottleneck moves to review and release

The second place the time is lost is the queue after writing. When code becomes cheap, changes tend to get bigger and more numerous. DORA’s 2024 analysis traced the fall in stability partly to exactly that: AI was associated with larger batches of change, and larger batches have long been riskier to release2. Reviewers face more code, the same hours in the day, and a draft that looks finished.

A year later the picture had moved. In its 2025 report, based on nearly 5,000 technology professionals, DORA found that 90 percent used AI at work and more than 80 percent believed it made them more productive. That second figure is a self-report of perceived productivity, not a measurement of delivery. AI adoption was now associated with higher delivery throughput, but still with lower stability: more failed changes and more rework10. The report’s summary was blunt. AI is an amplifier. It magnifies the strengths of teams with good engineering practice and the weaknesses of teams without it.

DORA then asked which capabilities make the difference, and named seven11.

Seven capabilities amplify AI's benefits in software delivery - a clear AI stance, healthy data, AI-accessible internal data, version control, small batches, user focus and quality internal platforms.Clear AI stancePeople know whatis allowedHealthy dataecosystemsGood, unified internal dataAI-accessibleinternal dataTools see yourown contextStrong versioncontrolEvery change trackedand reversibleSmall batchesChanges small enoughto reviewUser-centric focusClear about whatusers needQuality internalplatformsShared pathsto production
Figure 5.8.6 DORA’s seven capabilities that turn AI adoption into better delivery. Most are engineering practice, not AI.

Notice how little of the list is about the AI tool. Version control, small batches and good internal platforms were the marks of strong engineering teams long before coding assistants existed. That is the practical meaning of the amplifier: the money that makes AI pay off in engineering is often spent on the links around the tool, not on the tool.

Measure delivery, not volume

If the chain decides the value, the measures have to describe the chain. Lines of code, commits, suggestions accepted and tokens generated describe activity at the writing link. They are useful as usage data and misleading as measures of success, a general trap that Measuring AI Business Value covers in depth. In software the alternative is unusually well established.

Five delivery measures - change lead time, deployment frequency, recovery time, change fail rate and rework rate - describe whether software reaches users quickly and safely.MeasureThe question it answersGroupChange lead timeHow long from committed changeto production?ThroughputDeployment frequencyHow often do we release?ThroughputFailed deploymentrecovery timeHow fast do we recover when a release fails?ThroughputChange fail rateHow often does a release cause a failure?InstabilityDeployment rework rateHow many releases are unplanned fixes?Instability
Figure 5.8.7 DORA’s five delivery measures. Report them before and after any engineering AI rollout.

The first four measures come from the research behind Accelerate3. DORA added the fifth, rework, in 2024, because failed changes alone understated how much unplanned repair work teams were doing12. Together they answer the two questions an executive actually cares about: are we getting change to users faster, and is it holding up when it gets there?

Two additions make them more useful for AI programs. First, add the time changes wait for review, because that is where the new queue usually forms. Second, report the measures before and after a rollout, team by team, rather than as a single average that hides the teams AI is making worse. The one-line test for any engineering AI report is whether it measures generation or delivery.

Give agents the work you can check

Coding tools now come in two broad forms. An assistant responds while the engineer drives. An agent takes a task, plans, edits several files, runs the tests and proposes a finished change. From Copilots to AI Agents explains that distinction in general. What is specific to software is that much of the checking can be done by machines: code can be compiled, tested, type-checked and scanned automatically, at scale, before any person looks at it.

That makes one question decisive when deciding how much an agent may do: how well can we verify the result automatically, and how costly is a mistake that slips through? AI Agents and Intelligent Workflows, later in this module, develops the general rule of matching autonomy to consequence. In engineering, verifiability is the half that is easiest to overlook.

Agent autonomy in software should follow how well results can be verified automatically and how costly an escaped mistake would be; well-tested, low-cost work comes first.HighLowCost of anescapedmistakeWeakAutomatic verification · StrongHumans decideSecurity, data deletion, architectureAgent draftsNamed human approvesAssist onlyWith spot checksAgents firstMigrations, upgrades, test updates
Figure 5.8.8 Start agents where machines can check the result and a missed error is cheap to undo.

The bottom-right corner is where the documented successes in this chapter sit. Large, repetitive changes across a codebase, such as upgrading a library, moving to a new framework or updating tests, are tedious for people, well defined, and checkable by the existing test suite. Google reported that 80 percent of the code modifications in its landed migration changes were fully AI-authored13. The top-left corner, by contrast, holds the changes no test suite can fully judge: authorization logic, secrets, operations that delete data, production deployment and core architecture. There, an agent can draft, but people decide. In every corner the default fence is the same: no production write access, no access to secrets and no merge without review.

Story: six weeks for a year and a half of work

In March 2025, Airbnb’s engineers described how they had handled a chore every large software company knows. Nearly 3,500 test files for the React components of its web product were written for an old testing library, Enzyme, that no longer fitted modern React testing practice. Moving them to a newer one, React Testing Library, had been estimated at a year and a half of engineering time14.

Airbnb migrated about 3,500 test files in six weeks against an estimate of a year and a half; 75 percent passed all checks within four hours, 97 percent after four days of tuning, and engineers finished the rest.Estimate1.5 yearsOf engineering timeby handFirst 4 hours75% of filesMigrated andpassing every checkNext 4 days97% of filesSample, tune, sweepAbout a weekThe last 100or soFinished byengineersSix weeksMigrationcompleteTest intent andcoverage kept
Figure 5.8.9 Airbnb’s test migration: AI did the volume, automated checks decided what counted, people finished the hard cases.

The team did not ask a model to rewrite the files and hope. It built a pipeline in which every file moved through a series of steps: refactor the test, make it pass, fix lint errors, pass the type checker. A file counted as done only when it had passed every step. When a step failed, the pipeline tried again with the error message and more context, sometimes many times, sometimes with dozens of related files included so the model could see how the code fitted together. In the first four hours, 75 percent of the files came through. Then the engineers studied the failures in samples, adjusted the prompts and context, and swept the remaining files again. Four days of that raised the figure to 97 percent. The last hundred or so files, the genuinely awkward ones, were finished by hand, and the whole migration was complete in six weeks with the original intent of the tests and their code coverage preserved14.

The figures are the company’s own account, and a test migration is close to the ideal case: a well-defined change, a large existing body of checks and a low cost if something slips. That is exactly why it teaches well. The engineers measured validated files, not lines generated. They spent their effort designing the checks and studying failures, not typing. And they reserved human time for the cases the machine could not settle. The value came from the verification design as much as from the model.

What this means for leaders

Four lessons follow, and each is a decision only leadership can make.

Ask for delivery evidence. Treat code volume and tool usage as background. Ask for the five delivery measures and review wait, before and after, by team. If delivery has not moved, the bottleneck has shifted, and you know where to look.

Fund the links around the tool. Automated tests, small changes, version control discipline, internal platforms and reviewer time are where AI’s gains are captured or lost. A budget that buys licenses but not verification buys a longer queue.

Grant autonomy by verifiability. Let agents start on migrations, upgrades and test maintenance where checks are strong and mistakes are cheap. Widen their scope as evidence accumulates, and keep high-cost, hard-to-check changes with named people.

Decide where the recovered capacity goes. Time saved at the keyboard does not become value by default. It can go to new features, paying down technical debt, security, reliability or modernizing old systems, or to cost. Each is legitimate. If nobody chooses, it disappears into longer review queues. A sensible first destination is the checks themselves, because generated code creates new demand for them. AI does not make senior engineers less important either: their leverage moves toward design, review and setting the standards the checks enforce.

Check yourself

  1. If AI-generated code compiles and passes its tests, it is correct.
  2. In large studies, the gain from AI coding tools shrinks between code written and software released.
  3. Developers using an AI assistant in a security study wrote less secure code but felt more confident about it.
  4. The best measure of an engineering AI program is the share of code the AI wrote.
  5. DORA’s research suggests AI fixes weak engineering teams.
  6. Large code migrations with strong automated tests are a good first job for coding agents.

Reflection: follow one change

What comes next

Faster, safer building raises a harder question: if software can be built more cheaply, what should we build, and can the product itself become intelligent? The next chapter, AI in Product Development, moves from how software gets built to how AI changes product discovery, design and the products themselves.

Laws referenced

EU Product Liability Directive (revised) · EU

Directive (EU) 2024/2853

No-fault liability now explicitly covers software, including AI systems and SaaS, and updates or the lack of security updates. Easier proof for claimants with complex products.

  • 2026-12-09 — Applies to products placed on the market from this date

Last verified 2026-10-06 · official text

EU Cyber Resilience Act · EU

Regulation (EU) 2024/2847

Products with digital elements sold in the EU, including software and AI-enabled products, must meet essential cybersecurity requirements across their life: secure by design, vulnerability handling and security updates, a software bill of materials, and reporting of exploited vulnerabilities. Code written with AI assistance is covered like any other code in the product.

  • 2024-12-10 — Entered into force
  • 2026-09-11 — Manufacturers must report actively exploited vulnerabilities and severe incidents
  • 2027-12-11 — Main obligations apply

Last verified 2026-10-08 · official text

References

  1. Mert Demirer, Leon Musolff and Liyuan Yang. Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools. National Bureau of Economic Research (Working Paper 35275). 2026.
  2. DORA (Google Cloud). Accelerate State of DevOps Report 2024. Google Cloud. 2024.
  3. Nicole Forsgren, Jez Humble and Gene Kim. Accelerate: The Science of Lean Software and DevOps: Building and Scaling High Performing Technology Organizations. IT Revolution Press. 2018.
  4. Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv (2302.06590). 2023.
  5. Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. 2026.
  6. Joel Becker, Nate Rush, Elizabeth Barnes and David Rein. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR (arXiv 2507.09089). 2025.
  7. Neil Perry, Megha Srivastava, Deepak Kumar and Dan Boneh. Do Users Write More Insecure Code with AI Assistants?. ACM CCS 2023, pages 2785-2799. 2023.
  8. Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath and Murtuza Jadliwala. We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs. USENIX Security Symposium 2025. 2025.
  9. European Commission. New EU product liability rules will apply to online platforms and software from December 2026. European Commission (Transition Pathways). 2025.
  10. DORA (Google Cloud). State of AI-assisted Software Development 2025. Google Cloud. 2025.
  11. DORA (Google Cloud). 2025 DORA AI Capabilities Model. Google Cloud. 2025.
  12. DORA. A history of DORA's software delivery metrics. dora.dev. 2026.
  13. Stoyan Nikolov, Daniele Codecasa, Anna Sjovall, Maxim Tabachnyk, Satish Chandra, Siddharth Taneja and Celal Ziftci. How is Google using AI for internal code migrations?. ICSE 2025, Software Engineering in Practice (arXiv 2501.06972). 2025.
  14. Charles Covey-Brandt. Accelerating Large-Scale Test Migration with LLMs. The Airbnb Tech Blog. 2025.

Further reading

Sources last verified 2026-10-08.