AI and Workforce Productivity
AI makes many tasks faster, and people notice. But a faster task is an input, not a result: it creates capacity, and capacity becomes business productivity only when the whole job, quality included, gets better and management decides what the freed time is for. Leaders who do the arithmetic honestly and measure output rather than usage are the ones who can show the gain.
After this chapter you can
- Distinguish task productivity from business productivity, using a precise definition of productivity.
- Convert "X percent faster" into time freed, and cap headcount and cost claims accordingly.
- Measure the whole job, with review, rework and quality included.
- Treat freed capacity as a management decision, not an automatic result.
- Separate AI usage metrics from productivity evidence, and keep measurement away from individual surveillance.
Between September and December 2024, the UK government ran a large, published workplace trial of an AI assistant. Twenty thousand civil servants in a range of departments used Microsoft 365 Copilot in their everyday work: drafting documents, summarizing long email threads, preparing reports. At the end, 7,115 of them answered a survey. On average, they said the assistant saved them 26 minutes a day1.
Before reading on, make a prediction. If twenty thousand people each save 26 minutes a day, how much more did the government get done?
The report’s own answer is the most instructive sentence in it. Because of the experiment’s constraints, “it was not possible to identify how time saved was spent”1. Nobody could say what the minutes became. A smaller evaluation in the Department for Business and Trade, which timed and blind-scored real tasks, found high satisfaction but “did not find evidence that time savings have led to improved productivity”2.
Neither result means the tool failed. Both mean that time saved is the beginning of a productivity story, not the end of it. The rest of this chapter is about what has to happen in between.
Productivity is valuable output per resource
The word productivity is used loosely in AI conversations, usually to mean “people do things faster”. Economists are more precise. The OECD’s standard manual defines productivity as a ratio of a volume measure of output to a volume measure of input use3. For a leader, a workable version is: valuable output produced, relative to the resources used to produce it. The resources include people’s time, technology, capital and outside services. The output has to be something the organization actually needs. If it is not valuable, the ratio tells you nothing.
That definition exposes the gap in most AI reporting. A faster task changes the input side of the ratio for one step of the work. Business productivity changes only when the organization produces more of what it values, or the same for less, across the whole workflow.
The first chapter of this module defined AI value as a measurable change in a business outcome. Workforce productivity is where that definition is most tempting to skip, because the first link in the chain is so visible. People feel faster. Surveys confirm it. The two UK evaluations show how far that feeling can sit from the result.
Task productivity is not business productivity
Keep two terms apart. Task productivity is how much faster or better one task is done with AI. A first draft that took 60 minutes now takes 20. That is real, and it is an input. Business productivity is whether the organization produces more valuable outcomes with the resources it has: more decisions made, more cases resolved, customers served sooner, with no fall in quality. That is the proof.
The difference shows in what gets counted. The more convincing field studies measure output, not minutes. When Brynjolfsson, Li and Raymond studied an AI assistant for support agents, the headline was the number of customer issues resolved per hour, which rose about 15 percent on average4. That is a business productivity measure: work completed per unit of time, observed in company data.
The reverse case is easy to find. Suppose AI doubles the number of documents a team produces. If the documents are unnecessary, or nobody reads them, productivity has not doubled. It may have fallen, because someone still has to open, read and file them. Generated output is not valuable output.
AI can change the work in several ways, and it helps to name them so you can see which one you are paying for. It can automate tasks people did by hand, such as classifying cases or extracting data from documents. It can augment people, drafting while the person’s judgment stays in charge. It can cut the friction of finding things: searching, reading and comparing earlier work. It can support decisions with forecasts and risk flags. It can catch errors earlier, which cuts rework even when the first draft takes just as long. And it can help people learn an unfamiliar area faster. These are capabilities; the measured evidence so far is strongest for drafting, writing and support work. Each mechanism has a different test. For all of them, the executive question is the same: did valuable output per resource improve?
Do the arithmetic: speed is not time freed
Productivity claims arrive as percentages, and two of them are routinely confused. “Thirty percent faster” and “thirty percent less time” are not the same claim.
If a task gets X percent faster, it now takes 100 divided by (100 plus X) percent of the old time. So the share of time freed is X divided by (100 plus X). A task that took 100 minutes and becomes 30 percent faster takes 100 divided by 1.3, about 77 minutes. It frees about 23 minutes, not 30. Fifty percent faster frees a third. Twice as fast, a 100 percent speed-up, frees half.
The research is reported both ways, which is why the rule matters. In a well-known experiment with 453 professionals on realistic writing tasks, ChatGPT cut average time by 40 percent5. Forty percent less time is the same result as about 67 percent faster. Both statements are true; only one tells you directly how much time was freed.
The same arithmetic disciplines a sentence heard in many boardrooms: “AI makes the team 20 percent more productive, so we need 20 percent fewer people.” If a team produces 20 percent more per hour, the same output needs 100 divided by 1.2, about 83 percent, of the hours. The ceiling is about 17 percent fewer hours, not 20. At 30 percent it is about 23 percent. That is the arithmetic limit before any of the organizational questions in the rest of this chapter, and before anyone has decided to act.
Measure the whole job, quality included
AI tends to speed up the step that is easiest to see, usually the first draft. The job, however, includes everything around it: preparing the task, checking the output, correcting it, getting it approved. A dashboard that times only generation will overstate the gain.
Take an illustrative report that took 60 minutes. AI cuts drafting by 40 minutes, so the drafting step is now three times as fast. But the draft sounds right and contains errors, and checking and correcting it adds 30 minutes. The whole job takes 50 minutes. The real saving is 10 minutes: 20 percent faster, about 17 percent of the time freed.
Some tasks get slower. In the Department for Business and Trade’s observed sessions, people with the assistant wrote emails faster and better, but did spreadsheet analysis more slowly and less accurately2. One department, one tool, and the sign of the effect changed from task to task. That is the jagged frontier described in AI Is Changing Everything6.
Quality has to enter the calculation, not sit beside it. A useful approximation is that productivity moves with output multiplied by quality. Suppose output rises 30 percent while quality falls 20 percent. Then 1.3 times 0.8 is 1.04: about a 4 percent gain, not 30. And poor quality does not stay with the person who created it: a 2025 survey of US desk workers found that low-quality AI output, “workslop”, took the colleagues who received it nearly two hours each to deal with7. The time one person saves can become time a colleague spends.
Capacity is a management decision
Suppose the arithmetic is honest and the whole job really is faster. What you have now is capacity: time and attention that are available for something else. Capacity is an option, not a result. It becomes value only when someone decides what it is for and changes the work so that it goes there.
Without a decision, freed minutes tend to dissolve into email, meetings and slightly slower work on everything else. Large-scale evidence is consistent with that: a Danish study linking chatbot use to employer records found no measurable effect on earnings or recorded hours8. Individual tasks got faster; the records did not move. As The Four Industrial Revolutions showed, this is the familiar productivity paradox, and it is resolved by redesigning work, not by adding more tools.
What the redesign looks like depends on the constraint. Where demand is waiting, as with a backlog of customers or applications, capacity can become more output from the same team. Where quality is the problem, it can become more checking. Where growth is planned, it can avoid a hire. Where work is genuinely shrinking, it can become lower cost, but only after an explicit organizational action: fewer paid hours, a planned hire not made, an outside contract ended. Hours saved are not cash saved. How much of the capacity actually turns into value, and why so much of it leaks away, is the subject of Productivity vs Realized Capacity, later in this module.
Usage shows adoption, not productivity
Most AI dashboards report what is easiest to count: active users, prompts, documents generated, minutes saved by self-report. These are adoption measures. They tell you the tool is being used. They do not tell you whether valuable output rose.
Climbing the ladder means choosing, for each workflow, one output measure and one quality measure that would move if the gain were real, and comparing them with a baseline or a group that does not yet have the tool. How to build that baseline and comparison is the subject of Baselines, Metrics and Measurement9.
One boundary matters as much as the measures. The goal is better organizational performance, not maximum employee activity. Ranking individuals by how often they use AI rewards activity over value, invites gaming and turns a productivity program into a surveillance program. Measure workflows and teams, not personal prompt counts.
Story: twenty-five minutes a week in 68 schools
In the summer term of 2024, the Education Endowment Foundation, an English charity that tests what works in schools, ran a carefully designed trial of a generative AI tool in everyday work. Sixty-eight secondary schools took part, assigned at random. In 34 of them, 129 science teachers were asked to prepare their lessons for pupils aged 11 to 13 with ChatGPT and a short online guide. In the other 34, 130 teachers were asked not to use generative AI at all. For ten weeks every teacher kept a weekly diary of preparation time. A panel of experienced science teachers, who did not know which group had produced what, rated a sample of the lesson resources. The independent evaluators chose the main outcome when they designed the trial: teacher workload, in a profession where workload is one of the main reasons people leave12.
The task gain was real and measured properly. In the last five weeks, teachers with the tool spent 56.2 minutes a week on preparation, against 81.5 minutes in the comparison group: 69 percent of the time, a saving of 25.3 minutes a week. By the rule earlier in this chapter, 31 percent less time is about 45 percent faster. The panel found no evidence that quality differed. The saving came even though most teachers used the tool for only one or two activities in a lesson, usually writing quiz questions or finding ideas12.
The whole job showed through as well. In interviews, teachers described the familiar catch: reformatting the output to fit their school’s slide templates, or hunting for diagrams the tool could not draw, sometimes cancelled the saving or outweighed it12.
Then the question this chapter keeps asking: what did the minutes become? Of the 68 teachers who used the tool and answered the end-of-trial survey, just under half, 31, felt they were saving time, even though the diaries measured a saving for the group as a whole. Self-reports can miss in both directions. Asked where saved time went, the most common answers, each counted out of all 68 respondents, were reducing overall workload (23), general administration (23) and marking (14); some told interviewers they could not say where it had gone. Others described a different gain: they now made separate resources for different classes, work they had skipped before because it took too long. Among teachers with the tool, the share who said they spent too much time preparing fell from 49 percent to 26 percent; in the comparison group it barely moved12.
The lesson is in the order of events, not the size of the number. The trial decided what the time was for when it was designed, workload relief, and measured that outcome against a control group with quality checked blind. That is why 25 minutes a week counts as a result here. A school that wanted something else from the same minutes, such as resources tailored to every class or faster feedback on pupils’ work, would have to decide it, organize for it and measure it. Left alone, the minutes went several ways at once, each chosen by an individual teacher. That is capacity, not yet a business result.
What this means for leaders
Treat every reported AI time saving as a claim about capacity, and ask what it became. Insist on honest arithmetic: convert “faster” into time freed, and refuse any headcount or cost figure that equals the headline percentage. Make the whole job the unit of measurement, so that review, rework and the quality of what reaches the next person are counted. Above all, make the capacity decision explicit before the rollout, not after: name the workflow, the constraint the freed time will go to, and the output measure that will show it.
Check yourself
- If AI makes a team 30 percent faster at a task, it frees 30 percent of their time on it.
- Forty percent less time and about 67 percent faster describe the same result.
- Hours saved by AI are financial savings.
- AI can make a whole job slower even when it makes the first draft faster.
- Rising AI usage across the workforce shows rising productivity.
- A 20 percent productivity gain allows at most about 17 percent fewer hours for the same output.
Reflection: follow the minutes
What comes next
This chapter completes the four sources of AI value: revenue, cost, customer experience and workforce productivity. In each one, the capability was only the start, and the value depended on a chain of changes that someone had to own. The next chapter, From AI Capability to Business Outcome, shows how to define that chain for a specific use case, with its assumptions and an owner, before the money is spent.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
Worker consultation on workplace technology · EU member states
AI Act Art. 26(7); national co-determination law, e.g. Germany BetrVG s.87(1) no. 6, Netherlands WOR art. 27
Introducing systems that can monitor or assess employees usually requires informing or obtaining the consent of works councils or employee representatives, depending on the country. Plan this before a pilot, not after.
Last verified 2026-10-06
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Government Digital Service (UK). Microsoft 365 Copilot Experiment: Cross-Government Findings Report. GOV.UK. 2025.
- Department for Business and Trade (UK). Microsoft 365 Copilot pilot evaluation. GOV.UK. 2025.
- OECD. Measuring Productivity: Measurement of Aggregate and Industry-Level Productivity Growth (OECD Manual). OECD Publishing. 2001.
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
- Shakked Noy and Whitney Zhang. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654). 2023.
- Fabrizio Dell'Acqua et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, HBS Working Paper 24-013. Harvard Business School. 2023.
- Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano and Jeffrey T. Hancock. AI-Generated "Workslop" Is Destroying Productivity. Harvard Business Review. 2025.
- Anders Humlum and Emilie Vestergaard. Large Language Models, Small Labor Market Effects (NBER Working Paper 33777). National Bureau of Economic Research. 2026.
- Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
- European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
- National Foundation for Educational Research (NFER), for the Education Endowment Foundation. ChatGPT in Lesson Preparation: A Teacher Choices Trial. Evaluation Report. Education Endowment Foundation. 2024.
Further reading
- National Foundation for Educational Research (NFER), for the Education Endowment Foundation. ChatGPT in Lesson Preparation: A Teacher Choices Trial. Evaluation Report. Education Endowment Foundation. 2024.
- Department for Business and Trade (UK). Microsoft 365 Copilot pilot evaluation. GOV.UK. 2025.
- OECD. Measuring Productivity: Measurement of Aggregate and Industry-Level Productivity Growth (OECD Manual). OECD Publishing. 2001.
Sources last verified 2026-10-10.