Module 06 Synthesis — Understanding AI Risk
"Is the AI safe?" has no useful answer. The executive question is what could go wrong, how bad and how likely it is, what the controls leave behind, and whether a named owner accepts that residual risk. The law draws some of the lines first; appetite decides only the space it leaves.
After this chapter you can
- Describe AI risk as a property of the whole workflow, using one question from each Module 06 chapter.
- Size an inherent risk by impact, likelihood, exposure and control.
- Distinguish inherent risk from residual risk, and explain why acceptance must be informed and owned.
- Explain which risk tiers the EU AI Act fixes by law before internal appetite applies.
- Reach an explicit, owned decision - approve, approve with conditions, redesign or defer, or do not deploy.
Suppose a mid-sized pharmaceutical manufacturer has built an AI agent for its cold chain. The agent reads the temperature logs from every refrigerated shipment of vaccine, decides whether each pallet can be released to pharmacies or must be quarantined, and files the deviation record. For twelve weeks it ran in the shadow of the quality team and agreed with them on almost every shipment. The backlog that used to delay releases disappeared. At the steering meeting, the project lead asks one question: it passed every test, so can we call it safe and switch it on?
The scenario is invented, but the question is one many leadership teams ask, and it is the wrong one. Every AI system can fail, so the honest answer to “is it safe?” is always no, and the useful answer is never yes. Module 06 has spent ten chapters taking that question apart. This chapter puts the pieces back together into a single judgment that an executive can make, defend and own.
The question that replaces “is it safe?”
The NIST AI Risk Management Framework defines risk as a combination of two things: how likely an event is and how large its consequences would be1. That definition quietly retires the safety question. A system is not safe or unsafe. It carries risks of different sizes, some of which you can shrink and some of which you must decide to live with.
The executive question has five parts. What could go wrong? How bad could it be, and how likely? What controls exist? And is the risk that remains after those controls acceptable for this business, for this value, with a named person who accepts it? The goal is not zero risk, which no system and no human process achieves. The goal is risk that is visible, controlled and accepted on purpose.
Module 06 in one map
The AI Risk Landscape drew six families of risk and four multipliers. Each chapter after it left one question for every review: what happens when it is confidently wrong, who carries the errors, where the data travels, what a fooled system can access or send, whether we have rights to the input and output, what happens if the provider changes or fails, whether the work goes on with the AI switched off, which rung it sits on, and, at the center, who notices and who can stop it.
The map is a checklist only in the loosest sense. Its real use is to show that the risks interact. An instruction hidden in a supplier’s email, the indirect prompt injection that Greshake and colleagues demonstrated against working applications2 and that OWASP ranks first among LLM risks3, is a security problem. If the agent reading it can update records, it becomes an autonomy problem. If those records hold personal data, it becomes a privacy problem, and if nobody is watching, an oversight problem. So assess the workflow, not the model alone.
Size the inherent risk with four questions
Start with the use case as designed, before any control is applied. That is its inherent risk, and four questions size it.
Impact asks how bad the outcome could be for customers, patients, money, safety or reputation. Likelihood asks how often the failure or the attack could happen. Exposure asks how far one failure would spread before anyone stopped it, and it gathers the four multipliers from The AI Risk Landscape: whether the result can be undone, how far it reaches, whether anyone would notice and how fast it repeats. Control asks how well the design lets you prevent, detect and contain the failure. In the cold-chain scenario the impact is a spoiled vaccine reaching a patient, the likelihood looks low after the pilot, and the exposure is every pallet the agent releases on its own.
Two warnings keep this honest. First, low likelihood is not low risk. A rare failure with a severe impact can demand the strongest controls of anything in the portfolio, and the case later in this chapter began with an event that its operator said had not happened once in more than five million miles. Second, beware false precision. A score of 7.3 out of 10 looks exact and is not. NIST notes that AI risks are hard to measure reliably and that measurement in a test setting can differ from what happens in deployment1. Scores help you compare, rank and communicate. They do not remove the uncertainty, so keep human judgment in the decision.
From inherent to residual risk
Controls stand between the inherent risk and the business. Some reduce the chance of failure, some catch it early and some limit the damage when it happens. How organizations design and audit those controls is the subject of Module 07. For the decision, what matters is what they leave behind.
What remains is residual risk, and it is the only risk worth an executive’s signature. NIST’s framework lists the responses: mitigate it further, transfer part of it, avoid it, or accept it1. ISO 31000, the general risk-management standard that ISO/IEC 23894 applies to AI, describes acceptance as “retaining the risk by informed decision” [@iso-31000-2018; @iso-23894-2023]. The word that matters is informed. A risk that nobody examined has not been accepted; it has been ignored. Acceptance is explicit, it has conditions, and a named business owner signs for it, because the owner of the process is the owner of its risk.
Residual risk is also not fixed. It moves when the model is updated, the data changes, a tool is added or a new group of users arrives. Keeping it under review across a portfolio of systems is governance work, and Module 07 takes it up.
The rung decides how much control is enough
The same controls leave very different residual risk depending on how much the system is allowed to do. This module used one ladder for that, and Agentic AI and Autonomous Actions explains each rung.
On the lower rungs a person stands between the system and the world. That person is a control, but only a real one if they have the time, information and authority to disagree. Lisanne Bainbridge’s Ironies of Automation explained four decades ago why that control weakens: the more reliably a system works, the less practice its human monitors get at taking over4. Human Oversight and AI Incidents showed how automation bias erodes review in the same way.
The practical consequence is one of the most useful moves an executive has: lowering the rung is often the cheapest control available. The cold-chain agent could release only shipments whose logs are clean, at the Execute rung, and merely recommend a decision on any shipment with a temperature excursion, leaving the release to a quality specialist. The value drops a little. The residual risk drops a lot.
The law draws some lines before your appetite does
Risk appetite is how much risk an organization chooses to carry for a given value. A marketing team can tolerate more uncertainty in a draft than a quality team can in a release decision. But appetite does not decide everything, because some tiers are set by law before any internal discussion starts.
Under the EU AI Act, practices such as social scoring and emotion recognition in workplaces and schools are prohibited, and no business case moves that line. Uses listed in Annex III, such as hiring, credit, education and access to essential services, carry duties for risk management, logging, human oversight and conformity assessment. Telling people that they are dealing with AI, and labeling synthetic content, is a separate duty. In the United States there is no federal AI statute, but specific rules already apply, such as New York City’s bias-audit requirement for automated hiring tools, which Bias and Fairness covered. How an organization sorts each of its uses into legal and internal tiers is taught in AI Inventory and Risk Classification in Module 07.
Four outcomes, one owner
The decision, then, is not “AI, yes or no”. It is value weighed against residual risk, inside the lines the law has drawn, and it ends in one of four explicit outcomes.
A bare approval is often not the strongest answer. Approve with conditions usually is: approved, provided that the agent releases only clean shipments, that excursions go to a person, that its release rate is monitored weekly and that a manual process can take over within a shift. Redesign or defer fits a use case with value but no adequate control yet. Do not deploy fits when the law prohibits the use, when no control brings the residual risk within appetite, or when nobody will own it. NIST is blunt on the last point: where a system presents unacceptable risk, development and deployment should stop in a safe manner until the risk can be managed1.
Keep the review proportional. Internal summarization needs basic data controls and a light path. An agent that moves money or medicine needs limits, approvals, monitoring and a fallback. Before you sign, run one test drawn from the incident lifecycle in Human Oversight and AI Incidents: if this fails on a bad evening, can we detect it, stop it, recover and learn from it? If the answer is no at more than one step, the system is not ready for the rung it is asking for.
Story: the car that pulled over
The Cruise post-mortem is an unusually detailed public record of what happens when the whole chain is tested at once.
At about 9:29 p.m. on 2 October 2023, at Fifth and Market Streets in San Francisco, a human-driven car struck a pedestrian and threw her into the path of a Cruise robotaxi driving with no one on board. The robotaxi braked hard but hit her. Its collision-detection system, which could see little of her at the moment of impact, classified the collision as a side impact rather than a frontal one. That triggered a routine designed to make the car safer after a crash: pull over, out of traffic. Because the system did not register a person underneath, the car moved off again at up to 7.7 miles per hour and dragged her about 20 feet before it stopped5.
The technical failure was not in the car’s main task. It was in a safety control, the pullover, applied to a case its designers had not anticipated. Cruise’s own recall filing said a pullover was “not the desired post-collision response” in that situation, and a software change would have kept the car still5. Controls are designs too, and they carry their own residual risk.
The second failure was organizational. The independent review that Cruise and GM commissioned found that leadership knew about the dragging by the next morning, yet in the first briefings regulators were shown video over connections that were often poor, and were not told. Leaders were fixated on correcting the media story that the robotaxi had caused the crash. The review named poor leadership, mistakes in judgment, lack of coordination and an “us versus them” mentality with regulators, although it found no evidence of an intent to mislead5.
The consequences ran from a permit suspension and a recall to a 1.5 million dollar federal penalty for incomplete crash reports6 and a deferred prosecution agreement with a 500,000 dollar criminal fine for a false report to the regulator7. In December 2024 GM stopped funding Cruise’s robotaxi business8. Read it as a synthesis: a system at the top rung, a rare event of extreme impact, a control with an unexamined failure mode, an oversight and incident response that broke under pressure, and a legal line that turned a technical failure into a business-ending one.
What this means for leaders
Module 06 reduces to a habit of mind. Stop asking whether the AI is safe. Size the inherent risk with four questions, look at what the controls leave behind, check which lines the law has already drawn, choose the rung on purpose, and end every review with one of four outcomes and a name. Treat the incident response as part of the design, because the Cruise case shows that the response can cost more than the failure.
Check yourself
- If the model is accurate in testing, the AI system is safe to deploy.
- A failure that is very unlikely is a low risk.
- Residual risk is the risk that remains after controls, and it is what an executive decides on.
- An organization’s risk appetite decides which tier every AI use falls into.
- Moving a system down one rung of the autonomy ladder can be an effective control.
- In the Cruise case, the review found the costliest failures were in the car’s main driving task.
Reflection: sign one decision
What comes next
Making this judgment once, for one system, is the easy part. Making it consistently across hundreds of systems, with clear decision rights, policies and evidence, is governance. Module 07 begins with What Is AI Governance?, which asks who decides about AI, by which rules, and how an organization knows those rules are being followed.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
NYC Local Law 144 (automated employment decision tools) · US - New York City
NYC Local Law 144 of 2021; DCWP rules
An automated tool that substantially assists hiring or promotion decisions needs an independent bias audit within the past year, a published summary of results, and notice to candidates at least ten business days before use. Penalties USD 500 to 1,500 per violation.
- 2023-07-05 — Enforcement began
Last verified 2026-10-06 · official text
References
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
- OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
- Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
- Quinn Emanuel Urquhart & Sullivan. Report to the Boards of Directors of Cruise LLC, GM Cruise Holdings LLC and General Motors Holdings LLC regarding the October 2, 2023 accident and Cruise's responses (with Exponent technical root cause analysis as appendix). Cruise LLC (published report). 2024.
- National Highway Traffic Safety Administration. Consent Order between NHTSA and Cruise LLC (Standing General Order crash reporting). NHTSA. 2024.
- US Attorney's Office, Northern District of California. Cruise admits submitting a false report to influence a federal investigation and agrees to pay 500,000 dollars. US Department of Justice. 2024.
- General Motors. GM to refocus autonomous driving development on personal vehicles. GM Newsroom. 2024.
Further reading
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- ISO/IEC. ISO/IEC 23894:2023 Information technology - Artificial intelligence - Guidance on risk management. International Organization for Standardization. 2023.
- Quinn Emanuel Urquhart & Sullivan. Report to the Boards of Directors of Cruise LLC, GM Cruise Holdings LLC and General Motors Holdings LLC regarding the October 2, 2023 accident and Cruise's responses (with Exponent technical root cause analysis as appendix). Cruise LLC (published report). 2024.
Sources last verified 2026-10-08.