AI Evaluation and Approval Gates
A test result is evidence about the conditions it was measured in. Approval is a separate act: an owner with authority decides whether to proceed, under which conditions, and what would reopen the decision. Gates fail when the two are blurred, and when anyone treats a legal duty as a risk they can accept.
After this chapter you can
- Distinguish evaluation (evidence) from approval (the owned decision to proceed).
- Set requirements and acceptance thresholds before testing, and test in conditions like deployment.
- Check technical, business, operational and governance readiness, and confirm the approver is fit to decide.
- Use conditional approval, expiry and written reassessment triggers.
- Keep exceptions and residual-risk acceptance to internal standards; legal duties are met or the system stops.
In 2021, researchers at Michigan Medicine published an evaluation of a sepsis prediction model that was already running in many US hospitals. Its developer had reported an area under the curve, a standard measure of how well a model separates cases from non-cases, of between 0.76 and 0.83. The Michigan team measured it on 38,455 real hospitalizations and found 0.631.
The practical meaning was starker than the statistic. Of 2,552 patients who developed sepsis, the model missed 1,709, or 67 percent. It caught only 183 cases, 7 percent, that clinicians had not already spotted. Meanwhile it raised an alert on 18 percent of all hospitalizations, which is how alarm fatigue begins1.
The developer, Epic, disputed the method. It said the analysis “didn’t take into account the required tuning that should precede real-world deployment”, and that a higher alert threshold would cut false alarms2. In 2022 it released a revised model that each hospital tunes on its own historical data. A prospective study published in 2026, at four US health systems that had done that tuning, found far better discrimination, an area under the curve of 0.82 to 0.92, though most alerts were still false: only 13 to 26 percent flagged a patient who had sepsis3.
Read fairly, the story strengthens the point. Nobody lied. The developer’s number described the conditions in which it was measured, and its own answer was that the model needed local tuning first. The trouble starts when a number like that is allowed to stand in for a decision. Somebody, in every organization that switched the first model on, decided that the evidence was good enough for their patients. The question is how to make that decision deliberately.
The core idea
Two activities are routinely blurred. Evaluation is the structured work of finding out whether an AI system meets defined requirements, and what limits, risks and behavior remain. Its product is evidence. Approval is an organizational decision, based on that evidence and the known risks, about whether the system may move to the next stage and under what conditions. Its product is a decision, and a decision has an owner.
The NIST AI Risk Management Framework makes the same separation. Its measurement function is about testing and documenting; its management function begins with a determination of “whether its development or deployment should proceed”4. A score of 0.83 or 94 percent is an input to that determination, never the determination itself. A gate that only confirms a checklist was completed has skipped the decision while appearing to make it.
A real gate has real outcomes: approve, approve with conditions, send back for rework, or stop. “Let’s discuss it further”, with no owner and no next step, is not an outcome. It is a decision to drift.
Requirements come before tests
Evaluation without requirements measures what is easy rather than what matters. Before any testing, the owner and the specialists should agree what the system must do, what it must never do, who uses it, who is affected when it is wrong, and how wrong it may be. NIST asks organizations to determine and document their risk tolerance in advance for exactly this reason4.
The acceptance threshold is the most neglected part. “How good is good enough?” has no general answer, because different errors cost different things. In the sepsis case, a missed patient and a false alarm are not equal, and a single accuracy figure hides the balance between them. A threshold should therefore name the errors that matter for this use and the level of each that the organization can tolerate, and it should be written down before anyone sees the results.
Not all evidence weighs the same. Useful evidence is relevant to this use, representative of the people and cases the system will meet, recent, repeatable and traceable to the data and version tested. The higher the risk tier, the stronger and more independent the evidence should be. And each gate along the lifecycle asks its own decision question.
Test the conditions you will run in
A general benchmark reports how a system performed on someone else’s test. It does not show that the system suits your process, your data or your users, and treating it as a green light is benchmark theater. NIST’s wording is precise: performance should be “demonstrated for conditions similar to deployment setting(s)”4. The sepsis model is the textbook case. It was not a bad model in the abstract; it was an unproven one in Michigan.
The most useful question to bring to any evaluation review is a simple one: what important situations are we not testing? Edge cases, failure cases, deliberate attempts to trick the system and rare cases where one wrong answer outweighs a thousand right ones all belong in the plan. So does the real workflow, with real users under time pressure.
Two traps deserve a sentence each. Human oversight has to be evaluated, not merely declared: can reviewers actually override the system, and do they trust it too much? And a system that takes actions needs tests of its permissions, escalation and stop conditions, not only of its answers. The risks themselves were the subject of Module 06; at the gate, the question is only whether the evidence covers them.
Four kinds of readiness
A system can pass its technical tests and still be unready. Readiness has four parts, and strength in one never excuses weakness in another.
Technical readiness asks how often the system fails, how it fails, and whether failure can be detected and recovered from. Occasional errors can be acceptable when the workflow has effective safeguards. Business readiness asks whether the system solves the problem it was bought for, whether users accept it, and whether the value exceeds the full cost. Operational readiness asks whether support, monitoring, incident response and rollback exist on the day of launch; a successful prototype can still be operationally unready. Governance readiness asks whether the system is classified, has a named owner, and has the required controls actually working rather than merely planned.
Engineering supplies much of the evidence for the first and third. That evidence supports the decision; it is not the business approval. The accountable owner has to understand the value, the limitations, the controls and the conditions, because that person answers for the outcome.
Who may say yes
Which body approves which tier is a routing question with its own answer. What matters at the gate is whether the person saying yes is fit to say it. An approver needs three things at once: accountability for the outcome, authority over the resources and the risk, and enough information to understand what is being accepted. Do not assign approval to someone who cannot understand or accept the relevant risk, and never to the person who did the assessment.
The international management-system standard for AI, ISO/IEC 42001, builds this in: the AI risk treatment plan, and the acceptance of any residual risk, must be approved by designated management5. Residual risk is what remains after controls and legal duties are met. Risk-management guidance for AI treats retaining a risk as an informed decision6: a positive act with a name and a date attached, not the absence of an objection.
Gates fail in two directions. A weak gate waves risk through. A heavy gate delays everything, catches little and teaches teams to go around it. The way to tell them apart is to measure the gate itself: time to decision, rework rate, exceptions granted, backlog and, most telling, incidents in systems the gate approved. Routine checks such as required documents, approved providers, data classification and expiry dates can be automated, so that human attention goes to judgment.
Conditions, expiry and reassessment triggers
High risk is not an automatic no. Conditions are how governance lets an organization learn safely. An approval can cap the user group, limit the data, require human approval of each action, add monitoring, and expire after a fixed period, for example ninety days, at which point the owner must come back with live evidence.
Every approval should say in writing what would reopen it. These reassessment triggers are part of the decision, not an afterthought: a change of purpose, data, model or provider, more autonomy, a new user group, a breach of a monitored threshold, or a serious incident. The trigger list attached to each approval is what sends a system back to the gate without anyone having to argue that a change “feels big”. The law already writes some triggers for you. Under the EU AI Act, a substantial modification of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning7.
The approval record holds it together. It states what was approved, for what purpose, at what risk tier, on what evidence, under what conditions, by whom, when, and when or on what trigger it must be reviewed.
Residual risk is not a waiver of the law
Sometimes the evidence shows that a requirement is not fully met. There are three honest responses: meet it, approve an exception, or do not proceed. Which of them is available depends on where the requirement comes from.
If the requirement is one of your own standards, a documented exception with an owner, a rationale, compensating controls and an expiry is legitimate; how exceptions are managed over time belongs to the next chapter. If the requirement comes from law, there is nothing to approve. No committee and no risk owner can accept the risk of not complying, and a voluntary framework such as the NIST AI RMF is yours to tailor in a way that a statute is not.
Story: Amsterdam’s Smart Check
In 2019, the welfare department of the City of Amsterdam began building a model to help decide which applications for social assistance deserved a closer investigation. The aim was to investigate fewer people, and to treat them more fairly than the existing process did. The city tried to do everything the responsible-AI literature recommends. The model, called Smart Check, used 15 characteristics and left out direct demographic inputs such as gender, nationality and age. It was trained on about 3,400 past investigations. Consultants were hired, academics were consulted, and a council of welfare recipients was briefed10.
The first gate was a laboratory one. A test in May 2022 found the model was nearly twice as likely to wrongly flag an applicant with a non-Western nationality as one with a Western nationality. The team reweighted the training data to reduce that disparity before going further11. By the standard of most organizations, the evaluation had been done, honestly and well.
The second gate was live. From June to August 2023 the city ran the model on all incoming applications, so the pilot was limited in time, not in the people it touched11. Nearly 1,600 applications went through it10. The model flagged more people than intended, not fewer. The bias had not gone; it had reversed, and the model now wrongly flagged Dutch nationals and women more often, along with applicants who had children. It also did no better than the caseworkers it was meant to help10. In late November 2023 the city announced that it would shelve the pilot11; the alderman responsible for social affairs, Rutger Groot Wassink, told the city council he could not justify a system that might be discriminating10.
Read as a gate story, three lessons stand out. First, a careful laboratory evaluation is still evidence about the laboratory; only a live run produced evidence about the city. Second, the conditional step worked as a gate should, though not by being small: every applicant in those three months was scored. What contained the harm was that the run was limited in time and reversible. It had an end, the evidence was reviewed, and a named official with authority made a decision he could explain in public. Third, evaluation could not have answered the question raised in March 2022 by the Participation Council, the city’s panel of welfare recipients, which said the experiment touched citizens’ fundamental rights and should be discontinued10. Whether a benefit is worth a risk is an approval question. A system like this, which profiles benefit applicants, would likely be high-risk under the EU AI Act once the Annex III obligations apply; confirm the classification of any comparable system with counsel.
What this means for leaders
The executive’s job at a gate is not to read the test report more carefully than the engineers did. It is to ask the questions a test report cannot answer. Insist that requirements and thresholds are written down before testing starts. Ask what was not tested. Check that all four kinds of readiness are covered, not only the technical one. Prefer a conditional yes with an expiry and written triggers over an unconditional yes or a reflexive no. And hold the legal line: a risk owner can accept residual risk, never the risk of breaking the law.
Check yourself
- A strong benchmark score is enough to approve a system.
- Acceptance thresholds should be set before the test results are known.
- High-risk AI should never be approved.
- Engineering sign-off on the test results is the business approval.
- A risk owner can accept the risk of not meeting a legal requirement.
- An approval should state what would reopen it.
Reflection: the last yes you gave
What comes next
A gate is a moment. Between gates, something has to keep the conditions honest and notice when a trigger fires: the policies and standards that set the rules, the monitoring that watches the running system, and the audit that checks independently that all of it works. The next chapter, AI Policies, Standards, Monitoring and Audit, builds that durable control system.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
References
- Andrew Wong, Erkin Otles, John P. Donnelly and others. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 181(8), 1065-1070. 2021.
- Fierce Healthcare. Epic's widely used sepsis prediction model falls short among Michigan Medicine patients. Fierce Healthcare. 2021.
- Andrew Wong, D. Currey, M. Schwinne and others. Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open, 9(2), e260181. 2026.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- ISO/IEC. ISO/IEC 42001:2023 Information technology - Artificial intelligence - Management system. International Organization for Standardization. 2023.
- ISO/IEC. ISO/IEC 23894:2023 Information technology - Artificial intelligence - Guidance on risk management. International Organization for Standardization. 2023.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
- European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
- European Parliament and Council of the European Union. Regulation (EU) 2016/679 (General Data Protection Regulation). Official Journal of the European Union. 2016.
- Eileen Guo, Gabriel Geiger and Justin-Casimir Braun. Inside Amsterdam's high-stakes experiment to create fair welfare AI. MIT Technology Review (with Lighthouse Reports and Trouw). 2025.
- Lighthouse Reports. How we investigated Amsterdam's attempt to build a 'fair' fraud detection model. Lighthouse Reports. 2025.
Further reading
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- ISO/IEC. ISO/IEC 42001:2023 Information technology - Artificial intelligence - Management system. International Organization for Standardization. 2023.
- Andrew Wong, Erkin Otles, John P. Donnelly and others. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 181(8), 1065-1070. 2021.
- Eileen Guo, Gabriel Geiger and Justin-Casimir Braun. Inside Amsterdam's high-stakes experiment to create fair welfare AI. MIT Technology Review (with Lighthouse Reports and Trouw). 2025.
Sources last verified 2026-10-08.