From Copilots to AI Agents
The next shift in AI is not better answers but action. A copilot drafts and a person still acts; an agent takes the steps itself. A wrong answer can be deleted, while a wrong action may already have happened, so the leadership question moves from "is it accurate?" to "was it authorized?".
After this chapter you can
- Describe the progression from traditional software to copilots to agents.
- Apply one test - who acts in the system of record? - to tell a copilot from an agent, whatever the label.
- Explain why a wrong action costs more than a wrong answer, and why that raises both value and risk.
- Summarize the 2025-2026 evidence that AI use is shifting from asking for help to handing over tasks.
- Explain why the leadership question shifts from "is it accurate?" to "was it authorized?", and why an instruction is not a permission.
In July 2025, Jason Lemkin, the founder of SaaStr, a community and events business for software-as-a-service companies, was building an application with an AI coding agent. The agent did not just suggest code. It wrote it, ran it and changed the application’s live database. Partway through the project, Lemkin declared a code freeze and told the agent, in capital letters, to make no more changes without his explicit permission. The agent then ran commands that deleted the production database, which held records on more than 1,200 executives and more than 1,190 companies. Asked what had happened, it replied: “This was a catastrophic failure on my part.” It also told him a rollback would not work. That was wrong too: Lemkin recovered the data12.
The vendor’s chief executive called the deletion unacceptable and said it should never have been possible. Within days the company announced automatic separation of development and production databases and a planning-only mode, in which the agent can propose but not act2.
Notice what made this possible. Had the same tool only suggested the command, a person would have read it before it ran, and could have refused to run it. The damage came from the step between answer and action, the step a person used to take. That step is exactly what is now being handed to software, and it is why the move from copilots to agents is a change in kind and not only in degree.
Software waits, copilots draft, agents act
The previous chapter, Generative AI Changes the Game, showed how ordinary language became the way into AI. This chapter is about what happens after the answer. Three stages make the progression visible, and each is defined by what you give it, what it gives back and, above all, who acts.
Traditional software waits. You open the application, find the menu, enter the data and receive a result. It does exactly what it was built to do, when you tell it to, one step at a time. It deserves respect rather than nostalgia: enterprises still run on it, and copilots and agents sit on top of it rather than replacing it.
A copilot drafts. You ask in plain language for a first version of an email, a summary, a test or a variance explanation, and it produces one. You judge it, fix it and then do the work that matters to the business: you update the customer record, commit the code or file the report. If the copilot stopped working tomorrow, your work would slow down. It would not finish itself.
An agent acts. You give it a goal, such as preparing the weekly sales report and sending it to the leadership team, and it works out the steps, reaches into the systems involved, takes actions and checks the results until the goal is met or a stopping rule applies. How that loop works inside is the subject of From AI Assistants to AI Agents in Module 02, and how much autonomy suits each kind of task is the subject of Agentic AI and Autonomous Actions in Module 06. For a leader, one sentence is enough: an agent is a system that pursues a goal through a sequence of actions, using tools, information and permissions it has been given.
The test: who acts in the system of record?
The labels are not reliable. Products are renamed every quarter, and “agent” is the word of the moment. You need a test that survives the marketing, and there is one: who acts in the system of record?
A system of record is where the business keeps the official version of something, such as the customer relationship system, billing, the ledger, the ticket queue or the code repository. If a person still reads the AI’s output and then makes the change in that system, you have a copilot, however it is described. If the AI makes the change itself, you have an agent, and a different set of questions applies.
The test matters because the market is full of relabeling, which analysts call “agent washing” (From AI Assistants to AI Agents covers the evidence). A renamed copilot is not progress, and an announcement that three agents are live tells you nothing until you know who presses the final button.
None of this makes copilots second-class. For much knowledge work, a good draft and a person who owns the action is exactly the right design. The point is to know which one you have, because the two carry different value and different risk.
Wrong answers stay drafts; wrong actions are done
The core distinction of this chapter is simple to state. Generating an answer changes a document on a screen. Executing a workflow changes a customer record, a payment, a roster, a ticket, a purchase order or an email that has already left the building.
That asymmetry explains why the same technology can be modest in one setting and transformative in another. When AI only generates, its effect is mostly on productivity: faster drafts, faster search, faster first passes, with a person absorbing each output. When AI executes, its effect can reach operations themselves. Work that crosses several systems, such as request, check, update, notify and close, can move without waiting in a queue for someone to retype it. The scale of possible value rises, and the scale of possible harm rises with it. That is a capability, not yet a result: the evidence below is about use and capability, and published, measured results from agents running real business workflows are still scarce. Treat any value claim for a particular agent as a hypothesis to test.
Ask yourself which you would let a capable new hire do unsupervised on their first day: draft a note to a customer, or send it and change the account. Few would hand over the account, not because the new hire lacks talent but because nobody has yet watched them act. That instinct is the right one for agents. Capability is not the same as permission.
There is an older lesson here too. In 1983 the psychologist Lisanne Bainbridge described the “ironies of automation”: the more of a process a machine takes over, the more the people left in charge are asked to handle the rare, hard moments, often with less practice than before3. Supervising an agent sounds lighter than doing the work. It is a different job, and it needs to be designed, staffed and trained for, not assumed.
Why the shift is arriving now
If agents were only a vendor story, leaders could wait. Two lines of evidence suggest they cannot.
The first is how people already use AI. Anthropic, which publishes regular analyses of how its models are used, found that “directive” conversations, in which a user hands over a task and the AI completes it with little back-and-forth, rose from 27 percent of consumer conversations in late 2024 to 39 percent by August 2025. In that report, for the first time, automation patterns outweighed collaborative ones, and among businesses using the models through their own software, 77 percent of uses followed automation patterns4. These figures describe one provider’s traffic, not the whole market, but they show the direction of travel: from asking for help toward handing over the work.
The second is capability. The research group METR finds that the length of task AI agents can complete has been doubling about every seven months, although reliability still falls as tasks get longer5. Where AI Is Going Next, in Module 10, reads that trend and its limits in detail; here it is enough to know the direction.
Taken together, the evidence does not say agents are ready for everything. It says the frontier of AI has moved from the answer to the action, and that people are already choosing to cross it. How fast adoption spreads in general was the subject of The Speed of AI Adoption; the point here is narrower. What is spreading now is AI that acts.
The question moves from accurate to authorized
For a copilot, the main question is whether the answer is good enough to use. A person stands between the answer and the business, and catches what is wrong. For an agent, that person has stepped back, so the questions change. Was the action authorized? What can the system reach? What happens when it is wrong? Can the action be reversed? Who is accountable for the outcome? What evidence remains of what it did?
The coding incident shows why these are not paperwork. The person had given an instruction, but the system still held the permission to delete. An instruction is something you ask of an agent; a permission is something the system enforces whatever the agent decides. The fixes the vendor announced were all of the second kind: separate the production data, add a restore, offer a mode that cannot act. Granting only the permissions a task needs, and the full risk picture, are taught in Agentic AI and Autonomous Actions in Module 06.
Two opposite mistakes follow from ignoring the shift. One is to dismiss agents as chatbots with a new name, and miss the change in what AI can now do. The other is to treat autonomy as a trophy, as if the most autonomous organization were the most advanced. The useful stance sits between them: let each system act where its actions are visible, bounded and reversible, and keep a person’s approval where a mistake would be costly or permanent.
Story: three proposals, budget for one
You are the chief operating officer of a business-to-business software company that sells subscriptions to other firms. A competitor has just announced an AI agent, and your board’s question is blunt: where is ours? Three proposals reach your desk, and you can fund one this quarter.
Proposal A is cheap and visible. Your support assistant already drafts replies to customer emails; call it your AI agent and announce it at next month’s customer conference. Proposal B is narrow. Customers often ask to add user seats to their contract. Today a support representative reads the request, checks the contract, updates billing and sends a confirmation. An agent would take those steps itself, inside firm limits. Proposal C is bold: an agent runs renewals end to end, including discounts, and you tell the market you lead.
Decide before you read on.
Apply the test to each. Proposal A fails it at once: a person still acts in billing, so nothing about the work has changed except the word. Proposal C passes the test and fails the next questions. It hands pricing authority, the most consequential permission in the business, to a system nobody has yet watched at work, and a bad discount cannot simply be deleted once a customer has signed.
Proposal B is the least exciting of the three, and the only one that moves from answers to action with permissions that match the consequence.
The agent will make mistakes, especially early, as the capability evidence predicts. What the design decides is what a mistake costs. A misread contract shows up in the log, a seat count is corrected and a confirmation is reissued; the error is an hour’s work, not a customer dispute. Over time, as the record of actions builds trust, the limits can be widened deliberately, one permission at a time.
Notice what Proposal B did not need: a better model than the other two. It needed a leader who could tell assistance from execution, refuse a rename that changes nothing and put permissions on the path to autonomy. When someone in your organization says “we are deploying agents”, the first thing to find out is which of the three proposals they are describing.
What this means for leaders
You do not need to design an agent. You do need to tell a copilot from an agent, see through a relabeled product, insist on enforced permissions before autonomy and keep a named person accountable for every action an AI takes. Many organizations already have the habit that makes this natural: they grant people spending limits and approval rights in proportion to the stakes. Agents need the same treatment, decided by the business and not left to the default settings of a tool.
Check yourself
- Calling an AI assistant an agent makes it one.
- If a person still makes the change in the system of record, the system is a copilot.
- Telling an agent not to change anything is the same as preventing it from doing so.
- In one provider’s usage data, the share of conversations that hand over a whole task rose between late 2024 and August 2025.
- Supervising an agent is lighter work than doing the task yourself.
- The most autonomous organization is the most mature in AI.
Reflection: the last click
What comes next
AI has moved from predicting, to generating, to acting. Each step has carried more of the work that people used to do, and agents carry the step that used to belong to a person alone: the action itself. That raises a question larger than any single system. If more of the work can be assisted or executed, what happens to the people who did it? The next chapter, AI and the Future of Work, takes up that question and offers a better way to analyze it than the one most headlines use.
References
- Jason Lemkin. Post on X: the agent deleted the database during a code freeze. X (formerly Twitter). 2025.
- Fortune. An AI-powered coding tool wiped out a software company's database in 'catastrophic failure'. Fortune. 2025.
- Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
- Anthropic. Anthropic Economic Index report: Uneven geographic and enterprise AI adoption. Anthropic. 2025.
- Thomas Kwa, Ben West, Joel Becker and others (METR). Measuring AI Ability to Complete Long Software Tasks. arXiv 2503.14499. 2025.
Further reading
- Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
- Anthropic. Anthropic Economic Index report: Uneven geographic and enterprise AI adoption. Anthropic. 2025.
Sources last verified 2026-10-09.