AI Academy · Book
Executives & Directors · Module 06 · Chapter 001

The AI Risk Landscape

AI risk is not the chance that a model is wrong. It is what happens to the business when an AI-enabled system is wrong, and that depends on the decision it feeds, the action it takes and how far and how fast an error can travel. Accuracy is a property of the model; risk is a property of the use.

≈ 15 min read

After this chapter you can

  • Define AI risk as the business impact of an AI-enabled system's errors, not the model's error rate.
  • Explain why risk can enter at every layer of an AI system, including people and the organization.
  • Name the six families on the AI risk map and recognize that real failures span several.
  • Show how the same capability carries different risk in different uses, as the EU AI Act also assumes.
  • Size a use case with four multipliers - reversibility, reach, detectability and speed.

Pages like the illustrative one below are a familiar sight in steering-committee packs. A team has built an AI system. It agrees with experienced staff 98 percent of the time on a test set of 5,000 past cases. Strictly, that is an agreement rate with past human decisions, not proven accuracy, but the page calls it accuracy. The status light is green, and the team would like a signature before Monday’s go-live.

An illustrative go-live dashboard showing 98 percent accuracy, 5,000 test cases and a ready status, with nothing about consequences.98%AccuracyAgreement withexperienced staff5,000Test casesPast cases reviewedReadyStatusGo-live MondayILLUSTRATIVE NUMBERS
Figure 6.1.1 An illustrative go-live page. Every number on it describes the model. None describes what happens to the business when it is wrong.

Look at what the page does not say. It does not say what the system does with its answers, who acts on them, or what happens in the two percent of cases where it is wrong. It does not say how quickly anyone would notice an error, how many customers one error could reach, or whether the result can be undone. Every number describes the model. None describes the business.

That gap is not hypothetical. In 2018 Zillow, the US property website, began buying homes directly, using its own pricing models to decide what to offer. In the third quarter of 2021 the company bought 9,680 homes, sold 3,032, and wrote down the value of its inventory by 304 million dollars because it had paid more than it now expected to sell for. In November it announced it would close the business and cut about a quarter of its workforce. The shareholder letter was candid: the company had been unable to forecast home prices “by much more than we modeled as possible”1.

In Q3 2021 Zillow bought 9,680 homes, wrote down inventory by 304 million dollars and then cut about 25 percent of its workforce.9,680Homes bought inone quarterQ3 2021, against 3,032 sold304Million written downPaid more than homeswould fetch25%Workforce reductionAs the business waswound downSource: Zillow Group, Q3 2021 shareholder letter · 2021
Figure 6.1.2 The forecast was a model problem. Buying thousands of homes on it, at speed, with the company’s capital, made it a business problem.

A forecast that misses is a model problem. A forecast that commits the company’s capital to thousands of homes in a single quarter, in a market moving faster than the model could follow, is a business problem. The difference between the two is the subject of this chapter.

Risk is a property of the use

The National Institute of Standards and Technology defines risk as a composite of how likely an event is and how severe its consequences are. It adds a point that matters more for leaders: AI systems are socio-technical. Their risks come from the technology together with how it is used, who operates it and the setting it is used in2.

That gives a working definition for an executive. AI risk is the potential business impact when an AI-enabled system produces, amplifies or acts on an outcome that is incorrect, unsafe, unauthorized or inappropriate. The error rate is one input. The consequence of the error is the other, and only the business can judge it.

Two questions side by side - how often it is wrong, which produces a score, and what happens when it is wrong, which produces a decision.MODEL QUESTIONHow often is it wrong?Produces a scoreBUSINESS QUESTIONWhat happens when itis wrong?Produces a decisionvs
Figure 6.1.3 The model question produces a score. The business question produces a decision, and that decision belongs to leaders.

Both questions matter. The mistake is to answer the first and assume you have answered the second. The goal is not zero risk, which no useful system and no human workforce achieves. It is risk that is understood, owned and controlled in proportion to the value at stake.

Risk can enter at every layer

A common mistake is to treat the model as the whole system. In practice the model is one layer of six, and risk can enter at any of them.

Six layers of an AI system from data at the bottom, through model, application, tools and workflow, to people and the organization at the top.People andorganizationTrust, habits, rolesWorkflowWhere the output is usedTools andintegrationWhat it can read and changeApplicationThe product around the modelModelWhat it learnedDataWhat it learned from and receivesCLOSERTOTHEBUSINESS
Figure 6.1.4 Many real failures involve more than one layer, and the upper layers are where the business feels them.

At the bottom, data decides what the system knows; a model built on past prices can miss a market that moves faster than it ever has, as Zillow found. The model is the next layer. The application wraps the model in a product, and a good model in a badly designed product is still a bad system. Tools and integration decide what the system can read and change, and every connection widens the surface: researchers have shown, in controlled tests, that instructions hidden in a web page or an email can hijack an AI assistant that reads them, a problem the chapter Security and AI Attacks covers3. The workflow decides where the output lands, and a correct output in the wrong process is still a bad outcome. At the top sit people and the organization: who trusts the output, who checks it and who is allowed to stop it.

The top layer is easy to underrate. One of the largest reviews of AI risk taxonomies, the MIT AI Risk Repository, catalogues 1,725 risks from 74 frameworks. Its authors found that human decisions cause nearly as many of those risks (38 percent) as AI systems themselves (42 percent); the rest have other or unclear causes. Both shares count risks named in the frameworks, not incidents or losses4. More than forty years ago, the psychologist Lisanne Bainbridge described the same irony in industrial automation: the more a system is automated, the harder the remaining human job of supervising it becomes5. Why oversight fails, and how to make it real, is the subject of Human Oversight and AI Incidents.

The map: six families of risk

There is no single authoritative list of AI risks, and you do not need one. NIST’s profile for generative AI names twelve risks that the technology creates or makes worse, from confabulation and data privacy to information security and the integrity of the supply chain6. The OWASP Top 10 for LLM applications ranks the security risks, led by prompt injection7. For an executive, six families cover the ground, and each has its own chapter in this module. The six are this course’s grouping, and most of NIST’s twelve fall into them: confabulation into accuracy and reliability, harmful bias or homogenization into bias, data privacy into privacy, information security into security, intellectual property and value chain and component integration into the IP and providers family, and human-AI configuration into operations and oversight. The rest, such as harmful or dangerous content, information integrity and environmental impacts, are content and societal harms that cut across the families.

A map of six AI risk families - accuracy and reliability, bias, privacy, security, intellectual property and providers, and operations, agents and oversight.Accuracy and reliabilityWrong, made up or inconsistentBias and fairnessUneven across groups of peoplePrivacy andconfidential dataInformation where it should not beSecurity and attacksManipulated or misusedIntellectual propertyand providersRights, contracts, dependencyOperations, agentsand oversightOutages, actions, people
Figure 6.1.5 Six families, one map. Each has its own chapter in this module; a single use case usually touches several.

The families are a map, not a filing system, because real failures rarely stay in one box. The Dutch childcare benefits scandal is a well-documented example. For years the Tax Administration used an automated system that designated benefit applications as risky, and nationality was one of its indicators. It had also kept data on the dual nationality of 1.4 million people long after the law required it to be deleted. In December 2021 the Dutch data protection authority fined the administration 2.75 million euros for discriminatory and unlawful processing8. By then the wider scandal had already brought down the government, which resigned in January 20219. One system touched bias, privacy, oversight and law at the same time, and the harm reached thousands of families before anyone with authority stopped it.

The first cluster: wrong, trusted, acted on

The family many leaders meet first is accuracy and reliability. It has three members worth telling apart. Accuracy is whether the output is correct. Hallucination, which NIST calls confabulation, is confidently stated content that nothing supports6. Reliability is whether the system behaves consistently enough for the process it serves. The next chapter treats all three in depth; here it is enough to see how they become business risk.

A chain from an AI output, to someone trusting it, to the business acting, to a consequence.AI outputWrong orunsupportedTrustedNobody checksActed onDecision ortransactionConsequenceCost or harmRisk = wrong output + unchecked trust + action
Figure 6.1.6 A wrong answer that nobody believes costs little. A wrong answer that is believed and acted on can cost a great deal.

The chain explains why the same error rate can be trivial or serious. Break any link, by checking the output, withholding trust or keeping the action reversible, and the risk shrinks. Remove every check, and a small error rate flows straight into the world.

The same capability, three different risks

Because risk lives in the use, the same capability can carry very different risk in different places. Consider one capability, reading documents and summarizing them, in three illustrative uses.

The same summarizing capability carries low risk for meeting notes, more for contract summaries and the most when it decides a claim or benefit.Use of the same summarizerWho acts on itIf it is wrongNotes from an internal meetingColleagues who were thereSomeone corrects itSummary of a contract fora lawyerA lawyer who reads the contractCaught, but time is lostSummary that decides a claimor benefitNobody before the decisionA person is harmed
Figure 6.1.7 Illustrative uses. The capability barely changes between the rows. The consequence changes completely.

Regulators reason the same way. The EU AI Act does not regulate models by how accurate they are; it classifies systems by what they are used for. Using AI to price life and health insurance for individuals is listed as high-risk, as is using it to decide eligibility for public benefits, which is exactly the kind of decision at the heart of the Dutch case10. How the law’s tiers fit together with your own risk appetite is covered in the module synthesis, Understanding AI Risk.

Four multipliers: undo, reach, notice, speed

The classic starting point, likelihood times impact, is useful and incomplete. For AI, four further questions multiply or shrink the impact of any error. The table shows three illustrative actions.

A draft email is reversible, narrow and easy to spot; a payment or contract is hard to undo, reaches far and may go unnoticed for weeks.ActionCan we undo it?How far does it reach?Would we notice?Draft emailOne readerYes, before sendingLetter to a customerPartlyOne customerOnly if they complainPayment or contractHardlyMoney, people, regulatorsPerhaps weeks later
Figure 6.1.8 Illustrative actions. Risk rises as actions become harder to undo, reach further and are harder to spot.

Reversibility asks whether the result can be undone. A draft can be deleted; a payment, a signed contract or a deleted record mostly cannot. Reach asks how many people, systems or decisions one error can touch before it is stopped: one reader, one customer or the whole book of business. Detectability asks whether anyone would notice. Some errors announce themselves; others, such as a customer who is overpaid, generate no complaint at all.

Speed is the fourth, and automation changes it most. A person makes one mistake at a time. A system connected to automation can repeat one mistake thousands of times before anyone looks, as the story below shows with refunds.

Autonomy ties the four together. The more the system can do on its own, the fewer people stand between an error and its consequence. This module uses one ladder for that, with five rungs from generate through recommend, plan and execute to operate autonomously; Agentic AI and Autonomous Actions explains it. When several multipliers stack up, irreversible, far-reaching, hard to notice and fast, the controls must be strongest.

Story: one assistant, two go-live choices

The following is an illustrative composite, not a documented case. Return to the dashboard that opened this chapter and suppose it belongs to a hotel group with resorts along one coast. A hurricane has forced several of them to close in peak season, and about 12,000 guests are owed refunds and compensation for cancelled stays in the next three weeks. The guest-services team has built an assistant that reads bookings, guest messages and receipts and proposes a refund. On 5,000 past refund cases, its proposal matched an experienced agent’s 98 percent of the time.

Two choices reach the head of guest services on the same Friday.

Choice A pays every refund under 5,000 automatically; Choice B pays only refunds under 200 automatically, re-checks 5 percent of them daily, sends the rest to staff as drafts and has a stop switch.CHOICE APay automatically from day oneEvery refund under 5,000 paid with no human look2% of 12,000 is about 240 wrong paymentsOverpayments draw no complaintsFast; errors paid and unseenCHOICE BDraft first, automate in stepsSlower; errors caught in time
Figure 6.1.9 Illustrative composite. Same assistant, same 98 percent. The two choices differ in what can be undone, how far an error reaches and whether anyone would notice.

Choice A is the team’s proposal: pay every refund under 5,000 automatically from the first day, because guests whose holidays a storm has ruined want their money and experienced agents are scarce. It is fast. But two percent of 12,000 is about 240 wrong refunds in three weeks, and every one of them is paid before a person looks at it. The underpayments will surface slowly, as complaints and perhaps as complaints to the consumer regulator. The overpayments will not surface at all, because nobody complains about being paid too much, and recovering money from guests after a storm is both hard and damaging to the brand. A published threshold also tells anyone inclined to inflate a request exactly where to aim.

Choice B uses the same assistant differently. Only refunds under 200 are paid automatically. Suppose, for illustration, that a third of the 12,000 claims fall below that line: about 4,000 automatic payments, of which 2 percent, about 80, will be wrong, each by less than 200. Each day a reviewer re-checks a random 5 percent of the automatic payments, about 200 over the three weeks, which is enough to see whether the error rate is holding. The other 8,000 or so refunds arrive as drafts that agents approve in batches, which still saves most of their time. One named person can switch automation off. As the samples build evidence, the automatic band widens, perhaps to 500.

Choice B is slower in its first days and costs agent hours. That is its price, and it should be stated openly. What it buys is visible in the multipliers: errors are caught while they can still be undone, reach is limited to a small band of refunds, and the daily sample makes the silent errors visible. The dashboard was identical for both choices. The risk was not, because the risk was never in the 98 percent. It was in what the business allowed the system to do with it.

What this means for leaders

The landscape yields four habits. First, ask the business question every time: what happens when it is wrong? A go-live page that answers only the model question is incomplete, however good its numbers. Second, look at the whole system: data, application, integrations, workflow and people, not only the model. Third, use the map to check which families a use case touches, knowing that many touch more than one. Fourth, size the multipliers, and put the strongest controls where errors are irreversible, far-reaching, quiet and fast. What you then do with a risk, whether you accept, reduce, transfer or avoid it, is the subject of the module synthesis.

Check yourself

  1. AI risk basically means hallucination.
  2. The same model can carry very different risk in two different workflows.
  3. A system that is 98 percent accurate is safe to automate fully.
  4. Human decisions cause a large share of AI risks.
  5. The EU AI Act classifies AI systems mainly by how accurate they are.
  6. Errors that nobody complains about carry no risk.

Reflection: read your own dashboard

What comes next

The map shows where AI risk lives. The family you are likely to meet first is accuracy: systems that are fluent, confident and sometimes wrong. The next chapter, Accuracy, Hallucination and Reliability, explains why that happens and how to design systems that detect, contain and recover from it.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

References

  1. Zillow Group. Q3 2021 Shareholder Letter (Form 8-K, Exhibit 99). U.S. Securities and Exchange Commission (EDGAR). 2021.
  2. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
  3. Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
  4. Peter Slattery et al. The AI Risk Repository: A Meta-Review, Database, and Taxonomy of Risks From Artificial Intelligence. Patterns (Cell Press); arXiv:2408.12622 v3. 2026.
  5. Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
  6. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
  7. OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
  8. Autoriteit Persoonsgegevens (Dutch Data Protection Authority). Tax Administration fined for discriminatory and unlawful data processing. Autoriteit Persoonsgegevens. 2021.
  9. France 24 (with AFP). Dutch government collectively resigns over childcare subsidies scandal. France 24. 2021.
  10. European Commission, AI Act Service Desk. AI Act, Annex III: High-risk AI systems referred to in Article 6(2). European Commission. 2024.
  11. U.S. Securities and Exchange Commission. In the Matter of Knight Capital Americas LLC, Release No. 34-70694. U.S. Securities and Exchange Commission. 2013.

Further reading

Sources last verified 2026-10-08.