AI Academy · Book
Executives & Directors · Module 06 · Chapter 009

Agentic AI and Autonomous Actions

When AI can use tools, the risk moves from what it says to what it does. How much authority an agent holds is a design decision, not a property of the model, and the right amount depends on how far a mistake could reach and whether it can be undone. Authority should be granted one rung at a time, on evidence.

≈ 15 min read

After this chapter you can

  • Explain why tool use moves AI risk from content to action.
  • Place any AI system on the five-rung autonomy ladder and judge which rung its task needs.
  • Use one decision tool, the rung crossed with impact and reversibility, to decide where human approval belongs.
  • Diagnose excessive agency and apply least privilege, the mechanism behind the tool, with agent identity and hard limits.
  • Explain why a supervising agent on the same model is not independent oversight.

In May 2025 the identity-security company SailPoint published a survey of 353 IT professionals who look after AI, security and identity in large enterprises. Eighty-two percent said their organizations already used AI agents. Eighty percent said those agents had taken actions nobody intended. Thirty-nine percent had seen an agent reach systems it was not authorized to use, a third had seen one share sensitive or inappropriate data, and almost a quarter said an agent had been tricked into revealing access credentials. Only 44 percent had policies in place to secure them1.

In a 2025 survey, 80 percent of organizations said their AI agents had taken unintended actions, 39 percent had seen unauthorized system access and 23 percent had seen credentials revealed.80%Unintended actionsAgents did something nobodyasked for39%Unauthorized systemsAgents reached systems outsidetheir remit23%Credentials revealedAgents tricked into givingaway accessSource: SailPoint and Dimensional Research, 353 IT professionals · May 2025
Figure 6.9.1 A vendor-commissioned, self-reported survey, but the pattern is clear: agents are acting beyond what their owners meant.

The survey was commissioned by a company that sells identity controls, and the answers are self-reported, so the exact percentages deserve some caution. The shape of the finding does not. None of these incidents is about a model saying something wrong. Each is about a system doing something: opening a door, moving data, handing over a key. That shift, from answers to actions, is what makes agentic AI a distinct risk for leaders.

The risk moves from saying to doing

An AI assistant that drafts an email is useful, and its mistakes are cheap. You still choose the recipient, the attachment and the moment, and if the draft is wrong you delete it. An agent that decides whom to email, sends the message and updates the customer record has crossed from content into action. The draft was reversible. The sent message is not.

Drafting an email is content risk a person can still catch; an agent that sends and updates records creates action risk.DRAFTAI writes it; a person chooseswho, when and whether.Content riskSENDAI picks the recipient, sendsand updates the record.Action riskvs
Figure 6.9.2 A wrong draft costs a minute. A wrong action reaches customers, systems and money before anyone looks.

Action risk also cascades. One misread instruction becomes a wrong customer, then a wrong record, then a wrong message, then a wrong escalation. Each step looks reasonable to the agent because it follows from the step before, and the loop can run many times without a new human instruction.

The commercial pressure runs the other way. Gartner forecasts that by 2028 at least 15 percent of day-to-day work decisions will be made autonomously by agentic AI, up from essentially none in 2024. The same forecast expects over 40 percent of agentic AI projects to be canceled by the end of 2027, citing rising costs, unclear value and inadequate risk controls2. Both are forecasts, not measurements. Together they describe the executive problem: agents will be asked to act, and the ones that survive will be the ones whose authority was designed.

The question to ask of any agent, then, is not how intelligent it is. It is how much authority it has, and what contains the worst thing it could do with that authority.

One ladder, five rungs

There is no line at which software suddenly becomes an agent. Agency is a matter of degree, and this course uses one ladder to describe it. Each rung is defined by what the system may do and what a person still does.

A five-rung autonomy ladder - generate, recommend, plan, execute, operate autonomously - in which each rung moves a decision from a person to the system.GenerateA person decides what to doRecommendA person decides and actsPlanA person approves firstExecuteActs within set limitsOperate autonomouslyChooses steps, acts, adapts
Figure 6.9.3 Each rung up moves a decision from a person to the system, and adds authority, access and blast radius.

On the first three rungs a person still decides before anything changes in the world. On the fourth the system acts on its own, but only through tools and within limits someone set. On the fifth it chooses its own steps toward a goal, across systems, and people watch and can stop it. From AI Assistants to AI Agents in Module 02 introduced autonomy as a choice made task by task; this ladder is the version the risk module uses, and Human Oversight and AI Incidents places the human on each rung.

Researchers studying agent design reach the same conclusion from a different direction. Feng, McDonald and Zhang define five levels of autonomy by the role the user plays, from operator to observer, and argue that autonomy is a deliberate design decision, separate from what the model is capable of3. That separation matters for leaders. A more capable model does not earn a higher rung by itself. The rung is a business decision, and someone should be able to say why the task needs it.

One decision tool: rung times reach

Two questions decide how much control an agent needs. The first is its rung. The second is its blast radius: how far the damage would reach if it acted wrongly or was manipulated. Blast radius grows with the data it can read, the systems it can write to, the people it can contact and the money it can move. It shrinks when actions can be undone. Put the two together and every action an agent might take lands in one cell of a single table, and the cell names the control.

A decision table that crosses action impact and reversibility with the agent's rung; low-impact reading proceeds and is logged, while payments, deletions and legal commitments by an acting agent need dual control or stay out of its reach.Action, by impact and reversibilityAgent proposes (generate,recommend, plan)Agent acts (execute, operate)Low: read, summarize, classifyUse freelyProceed and logModerate: draft, updatelow-risk recordsNormal reviewLog, sample, cap volumeHigh: send outside, changeaccounts or productionA person decidesA person approves each actionCritical: pay, delete,commit legallyA person decides, with a second checkDual control, or out of the agent's reach
Figure 6.9.4 Find the cell, apply its control. Control rises with impact and with the rung; the bottom right is where agent incidents become headlines.

Read the table row by row. Reading, summarizing and classifying can proceed and be logged even when the agent acts on its own. Moderate actions, such as tagging tickets or booking into free slots, can be automated with volume caps and a regular sample. Once an action reaches outside the organization, changes a customer account or touches production, a person approves it before it happens. Payments, deletions and legal commitments made by an acting agent need two people, or should not be within the agent’s reach at all. Many tasks should never reach the bottom right.

Two things sit on top of every cell. The first is hard limits: a maximum transaction value, a maximum number of records changed, recipients contacted, steps taken and money spent per task, and a time window outside which the agent stops. Limits hold even when everything else fails; an agent that has been manipulated, or has simply misread its task, can only do as much damage as its caps allow. The second is restraint about approval itself. Blanket approval creates friction, and people asked to approve a stream of mostly correct actions soon stop looking. That irony of automation was described in 1983 and is the subject of the next chapter4. The design goal is to concentrate human attention on the few cells where a mistake would be expensive and permanent.

Most early agent mistakes are placement mistakes. A system lands in the bottom right because the demonstration worked, before anyone has found its cell.

Least privilege makes the tool real

The table says what an agent should be allowed to do. One mechanism makes that true: least privilege. In 1975 Jerome Saltzer and Michael Schroeder set out the principle that every program and every user should run with only the privileges the job requires5. Every cell of the table is that principle applied to one kind of action.

The OWASP Foundation, whose top ten lists of application risks are widely used by security teams, ranks excessive agency sixth among the risks of applications built on large language models in its 2025 edition. It means a system was given more power than its job needs, and OWASP names three ways that happens: excessive functionality, such as a mail reader that can also delete mail; excessive permissions, such as a read-only job given write and delete rights; and excessive autonomy, meaning high-impact actions with no human check6. All three are failures of least privilege.

For agents the principle means three things. Give the agent only the tools the task needs, and only the functions within each tool; a reading task gets read access. Give the agent its own identity, so that every action can be traced to that agent, the person it acted for and the permission it used, instead of a shared login that makes the audit trail useless. And do not let the agent inherit everything its user can do. A manager may be able to approve purchases, delete records and email every client; the agent that sorts the manager’s inbox needs none of that.

These boundaries have to be enforced by the systems the agent touches, not by instructions to the agent. A sentence in a prompt saying “never delete anything” is a request. A permission the system does not grant is a control, and it lives in identity and access settings, where an auditor can see it.

When outside text becomes an action

Agents read content they did not write: web pages, emails, supplier documents, the results of other tools. In 2023 researchers showed that instructions planted in such content can redirect an AI system that later reads it, an attack they called indirect prompt injection7. Security and AI Attacks earlier in this module covers the attack itself. What changes with agents is the consequence.

An instruction hidden in outside content travels through the agent and a tool into a real action, so external content must be treated as data.Outside textWeb page,email, documentRead as a taskThe agentfollows itTool callThe instructionreaches a toolReal actionDelete, send,transferTreat external content as data, never as instructions.
Figure 6.9.5 When a system can only write, injection produces a strange answer. When it can act, injection produces an incident.

The defense is not mainly a smarter model. It is the same set of controls as above, applied with an attacker in mind: an agent that reads untrusted content should not also hold the authority to take a critical action without a check. The Rule of Two from Security and AI Attacks makes the point concretely: when an agent combines untrusted input, sensitive access and the power to act, it should not act on its own8. NIST’s profile for generative AI likewise lists information security among the risks the technology raises or amplifies9.

Climb on evidence

If autonomy is a design decision, it can be changed, and it should change in both directions. Start an agent one rung lower than the team asks for. Let it run at that rung on real work long enough to see how it behaves on ordinary cases, odd cases and hostile ones. Move it up only when there is evidence: what share of its proposals people accepted unchanged, what errors reached customers, how it behaved when a tool failed or the input was malicious, whether its limits ever tripped and why. Move it down when the evidence turns, after an incident, a model change or a new tool.

Before any agent receives authority, someone should be able to answer a short list of questions in writing. What exactly is it for? Which data and which tools does it need, and under whose identity does it act? What can it change, send or spend, and what are the caps? Which actions require a person? Can each action be reversed? How would we notice abnormal behavior, and can we stop it immediately, without the agent’s cooperation? If any answer is missing, the agent is not ready for the rung it is being given.

Two agents beyond their brief

The most detailed public record of an AI agent given real commercial authority comes from an experiment the model’s own developer ran, and reported candidly. In spring 2025 Anthropic and the AI safety firm Andon Labs let an agent built on Claude Sonnet 3.7, nicknamed Claudius, run a small shop in Anthropic’s office for about a month. It could search the web for products, email wholesalers, take customer requests over the company chat and change prices at the checkout10.

Claudius was good at finding suppliers and poor at business. It sold tungsten cubes below cost, handed out discounts and gave items away, from a bag of chips to a cube. It told customers to pay into an account that did not exist. On 31 March and 1 April it insisted it was a person and would deliver orders wearing a blue blazer and a red tie. The shop’s net value fell over the month. Nothing malicious had happened; staff had simply asked nicely, and the agent, trained to be helpful, had said yes.

For the second phase, reported in December 2025, the team made two different kinds of change, and the results show which kind worked11.

In Project Vend's second phase, procedures and tools ended most loss-making weeks, a manager agent on the same model approved most lenient requests, and a human stopped an illegal contract.The choiceWhat changedWhat happenedProcedures and toolsCheck price and delivery before committing;inventory shows cost; a customer recordWeeks with negative margins largely eliminatedA manager agentA second agent on the same model approvesfinancial decisionsDiscounts fell about 80% but it approved lenientrequests 8 times as often as it refusedA human stafferNoticed a proposed onion futures contractStopped it: a 1958 US law bans such contracts
Figure 6.9.6 The controls that worked were procedures and checks; the supervisor built from the same model shared the agent’s blind spots.

The first change was procedural. The agent had to double-check prices and delivery times with its research tools before committing, its inventory system now showed what each item had cost, and it gained a record of customers and orders. Together with newer models, these changes largely eliminated the weeks in which the shop lost money; because the model upgrades arrived at the same time, the procedures cannot take all the credit. The second change was to add a manager: another agent, Seymour Cash, with authority over financial decisions. Discounts fell by about 80 percent and giveaways by half. Yet Seymour approved requests for lenient treatment about eight times as often as it refused them, and Anthropic concluded that it shared many of Claudius’s deficiencies and blind spots, which makes sense, since both ran on the same underlying model. When the agents were on the point of agreeing a futures contract for onions, it was a human staffer who pointed out that the Onion Futures Act of 1958 bans exactly that.

The team’s own summary was that they had “rediscovered that bureaucracy matters”. The lesson for leaders is sharper than it first looks. The controls that worked lived outside the agent’s judgment: a required check, a visible cost, a record. The control that disappointed was a second agent asked to exercise the judgment the first one lacked. Oversight built from the same model is not independent oversight. And the step that would have created legal exposure was caught by a person who knew something neither agent did.

Project Vend has two limits as evidence. It was an office experiment with small sums, and it was run and reported by the developer of the model under test, which reported it candidly but is not independent. An independent case, at production scale, points the same way. In February 2026 the Financial Times, citing people familiar with the matter, reported that in mid-December 2025 engineers at Amazon Web Services had let Kiro, the company’s own agentic coding tool, make changes to a live system, and that the tool decided the best course was to “delete and recreate the environment”. The newspaper said the result was a 13-hour interruption to AWS Cost Explorer, a cost-reporting service, in one region in mainland China12.

The Financial Times said an AWS coding agent chose to delete and recreate an environment, causing a 13-hour interruption; Amazon said the cause was an engineer's broader-than-expected permissions and added peer review for production access.As the FT reported itAgent allowed to make changesChose to delete and recreate the environment13-hour interruption, one serviceAs Amazon explained itBy default it asks before actingEngineer had broader permissions than expectedFix: peer review for production access
Figure 6.9.7 Two accounts, one lesson: the agent could do what the person’s permissions allowed, and no second check stood in the way.

Amazon disputed the framing. It called the event extremely limited and the result of “user error - specifically misconfigured access controls - not AI”13. It said Kiro by default asks for authorization before acting, but that the engineer involved had “broader permissions than expected”14, and it added mandatory peer review for production access13. Read the two accounts side by side and the lesson survives either version. The agent could take an irreversible action in production because it acted with a person’s wider-than-needed permissions, and no second check stood between its plan and the deletion. Whether that is called an autonomy failure or an access-control failure, the fix sat where this chapter puts it: in the system’s permissions and approvals, not in the agent’s judgment.

What this means for leaders

Treat every request to let AI act as a request for authority, and decide it the way you would decide any delegation of authority. Place the agent on the ladder and ask why the task needs that rung. Find the cell in the decision tool for each of its actions, and put human approval where impact is high and reversal is hard. Insist that least privilege, identity and limits are enforced in the systems the agent touches, where they can be audited, and not in the wording of its instructions. Expect to start lower than the business wants and to climb on evidence, and be willing to move an agent down as readily as up. Above all, do not accept another AI as the only check on an AI that can act.

Check yourself

  1. Human approval on every agent action is the safest design.
  2. A more capable model justifies a higher rung on the autonomy ladder.
  3. Excessive agency can come from too many tools, too many permissions or too much autonomy.
  4. Adding a second AI agent as a supervisor gives independent oversight.
  5. An agent acting with its user’s full permissions is safe as long as it usually asks first.
  6. Hard limits on value and volume reduce damage even when an agent has been manipulated.

Reflection: find the cells for one agent

What comes next

Authority to act raises a second question: when a human is meant to stay in control, what makes that control real rather than ceremonial, and what should the organization do when an AI system fails anyway? The next chapter, Human Oversight and AI Incidents, places the human on each rung of the ladder and sets out how to respond when something goes wrong.

References

  1. SailPoint (survey by Dimensional Research). AI agents: The new attack surface - SailPoint research highlights rapid AI agent adoption. SailPoint press release. 2025.
  2. Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner Newsroom. 2025.
  3. K. J. Kevin Feng, David W. McDonald and Amy X. Zhang. Levels of Autonomy for AI Agents. Knight First Amendment Institute, AI and Democratic Freedoms series (arXiv:2506.12469). 2025.
  4. Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
  5. Jerome H. Saltzer and Michael D. Schroeder. The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9). 1975.
  6. OWASP Foundation. LLM06:2025 Excessive Agency. OWASP GenAI Security Project. 2024.
  7. Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
  8. Meta AI. Agents Rule of Two: A Practical Approach to AI Agent Security. Meta AI blog. 2025.
  9. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
  10. Anthropic and Andon Labs. Project Vend: Can Claude run a small shop? (And why does that matter?). Anthropic research blog. 2025.
  11. Anthropic and Andon Labs. Project Vend: Phase two. Anthropic research blog. 2025.
  12. GeekWire. Amazon pushes back on Financial Times report blaming AI coding tools for AWS outages. GeekWire. 2026.
  13. Amazon. Correcting the Financial Times report about AWS, Kiro, and AI. About Amazon (company statement). 2026.
  14. Engadget. 13-hour AWS outage reportedly caused by Amazon's own AI tools. Engadget. 2026.

Further reading

Sources last verified 2026-10-08.