AI Academy · Book
Executives & Directors · Module 07 · Chapter 011

AI Policies, Standards, Monitoring and Audit

In one 2026 survey, most large organizations had some form of AI governance; few leaders said it worked. A policy says what must be true. Standards make it testable, controls make it happen, monitoring shows management whether it is happening, and an independent audit says whether that picture can be trusted. Exceptions and findings are where the system proves itself, or quietly fails.

≈ 16 min read

After this chapter you can

  • Explain how a standard turns a policy into a testable minimum, and why each standard implies a test.
  • Distinguish preventive, detective and corrective controls, and test a control for both design and operating effectiveness.
  • Monitor control health with thresholds, owners and deadlines, alongside system health.
  • Run exceptions as recorded, conditioned and expiring decisions rather than informal skips.
  • Distinguish management monitoring from independent audit, and close findings only when the changed control has been tested.

In May 2026 the American Arbitration Association published a survey of 500 senior legal and executive leaders at large organizations in the United States and Canada. Almost nine in ten, 87 percent, said their organization had some form of AI governance in place. Before reading on, guess how many said that governance was operating effectively.

The answer was 22 percent. The same share, 22 percent, were very confident they could produce evidence of their governance decisions if a regulator or an auditor asked for it. Only a third had defined how to escalate when an AI system misbehaves1.

In a 2026 survey 87 percent of large organizations had some AI governance, but only 22 percent said it worked and 22 percent were confident they could prove it.87%Have AIgovernanceIn some form22%Say it worksOperating effectively22%Could prove itVery confident theycould show evidence33%Know howto escalateWhen AI misbehavesSource: American Arbitration Association survey of 500 leaders · May 2026
Figure 7.11.1 Governance on paper is common. Governance that works, and can show that it works, is rare.

These are self-reported figures from a single survey, not audited results, and they should be read that way. But they describe a gap that many executives will recognize. The documents exist. What is missing is the machinery that turns a document into behavior, notices when behavior drifts, and lets someone outside the work confirm that it is all real.

The core idea

A policy says what must be true. A standard makes that requirement specific enough to test. A control is the mechanism that makes it happen. Monitoring shows management, continuously, whether it is happening. Audit tells the board, independently, whether management’s picture can be trusted. And every finding feeds back into the policy, the standard or the control.

Governance runs as a loop from policy to standard, control, monitoring and audit, with findings feeding improvement back into the policy.PolicyWhat must be trueStandardThe testableminimumControlPrevent, detect,correctMonitoringManagement's viewnowAuditIndependentassuranceImproveFindings changethe loopControl loopIntent to evidence
Figure 7.11.2 Governance is a loop, not a document. Each stage produces something the next one can check.

AI Policy vs AI Governance traced a single requirement from principle to evidence and showed that every link needs an owner. This chapter is about the machinery that keeps those links working day after day: how standards are written, what kinds of control exist, how a control is tested, what management watches, how exceptions are handled without becoming a second, unwritten policy, and what independent audit adds. ISO/IEC 42001, the international standard for AI management systems, is built on the same loop. It asks for a documented AI policy, then for monitoring and measurement, internal audit, management review and corrective action2.

Standards turn a policy into something you can test

A policy is written to last. It states intent in words broad enough to survive new tools and new teams: high-risk AI systems must be evaluated before use. That durability is also its weakness. Nobody can check whether “evaluated” has happened, because nobody has said what it means.

A standard says what it means. It sets the minimum that every team must meet, in terms concrete enough that a third party could check it.

A policy states intent, a standard sets a checkable minimum, and each standard implies a test anyone could run.TopicThe policy saysThe standard saysThe testEvaluationEvaluate high-risk AI before useTest on local data; record resultsin the registerFind the record for threelive systemsDataUse only approved dataNamed data classes; no restricteddata in promptsSample logs for restricted fieldsOversightKeep humans in controlReviewer can stop the system;overrides loggedPull last month's override log
Figure 7.11.3 If you cannot write the test, the standard is not yet a standard.

Three properties make a standard useful. It is specific: it names the evidence that shows compliance. It is reusable: a small library of standards for data use, evaluation, human oversight, third-party AI, logging and the permissions of AI agents means each new project inherits its requirements instead of renegotiating them. And it is stable: the procedures that carry out a standard, such as which form to fill in or which tool runs the tests, can change every quarter without the standard changing at all.

Policies and standards also need governing themselves. Each has an owner, an approver, a version, and a date by which it will be reviewed. ISO/IEC 42001 lists the review of the AI policy at planned intervals among its reference controls2. A standard that nobody has reviewed since the last generation of tools is a standard that teams have already started to work around.

Controls prevent, detect and correct

A control is any mechanism that makes an unwanted outcome less likely or less harmful. Controls come in three kinds, and a sound system uses all three.

Preventive controls stop failures, detective controls find them and corrective controls repair them, with a common baseline and more controls as risk rises.PreventApproval gates, accesslimits, blocked dataDetectLogs, alerts, sampling,user reportsCorrectRollback, revokeaccess, retireA baseline for every system, more as risk rises
Figure 7.11.4 Prevention fails sometimes. Detection and correction decide how long a failure lasts and how far it spreads.

Preventive controls stop the unwanted event: a release pipeline that refuses to deploy without an approval record, an access rule that keeps an agent away from the payments system, a filter that blocks restricted data from leaving the network. Detective controls find what prevention missed: logs, alerts, periodic sampling of outputs, and channels for users to report problems. Corrective controls restore an acceptable state: rolling back a release, revoking access, switching a feature off, retiring a system. The COSO internal control framework, which is widely used by finance and audit teams, asks organizations to choose a deliberate mix of preventive and detective, manual and automated control activities3.

Two design rules follow. First, set a baseline that applies to every AI system, such as a named owner, a risk tier, approved data use, access control, logging and an incident route, and add controls as the tier rises. The tiers themselves are set in AI Inventory and Risk Classification. Second, automate routine controls wherever you can. A manual check that people skip under deadline pressure is a control in name only, as the story later in this chapter shows.

A control can fail twice: in design and in operation

A useful idea executives can borrow from financial audit is that a control has two separate ways to fail. The standard that governs audits of internal control at US listed companies makes the distinction precisely. Design effectiveness asks whether the control, if it is operated as prescribed by competent people, would actually meet its objective. Operating effectiveness asks whether it is in fact operating as designed. Auditors test the second by inquiry, observation, inspection of records and, where it matters, re-performing the control themselves4.

A two-by-two of design against operation - only a control that is well designed and actually operated is effective.YesNoDesignmeets theobjectiveNoOperates as designed · YesPaper controlRight on paper, skipped in practiceEffectiveKeep testing itNo controlNeither designed nor runBusy but blindRuns every day, misses the risk
Figure 7.11.5 A control must be both well designed and actually operated. Test each separately.

Take one requirement and test it both ways. The policy says no high-risk system goes live without approval. The control is a release pipeline that blocks deployment without an approval record. The design test asks whether the block covers every route to production, including the vendor-hosted tool a business team can configure without touching the pipeline at all. The operating test tries a controlled release without approval and samples last quarter’s releases to see whether any went through. A control can pass one test and fail the other, and either failure can go unnoticed for years if nobody runs the test.

For AI used in financial reporting, none of this is optional. The Sarbanes-Oxley Act requires management to assess internal control over financial reporting each year, with auditor attestation for larger filers, and AI used in the close, in reconciliations or in reporting falls inside those controls.

Monitor the controls, not only the system

AI Lifecycle Governance covered monitoring the system itself: its accuracy, drift, cost and business outcomes. This chapter adds a second layer that is easy to miss. Management also has to watch whether the controls are working. How many high-risk systems are running on a current approval? How many exceptions are past their expiry date? How many alerts were acknowledged within the agreed time? How many audit actions are overdue? How much of the AI estate is in the inventory at all?

COSO calls this ongoing evaluation, built into normal operations, as distinct from the separate evaluations that an independent function carries out from time to time3. The NIST AI Risk Management Framework asks organizations to plan ongoing monitoring and periodic review of the risk management process itself, with roles and review frequencies defined5.

Monitoring thresholds rise from normal to warning, alert and incident, and each level has a named owner and a deadline.NormalLogged and trendedWarningThe owner looks this weekAlertThe owner acts by a set timeIncidentEscalate, contain, report
Figure 7.11.6 Every level names who acts and by when. A red tile on a dashboard is not a response.

Every indicator needs thresholds, and every threshold needs an owner and a clock. A dashboard can show red for weeks and still look like control to the people who glance at it. If an alert goes to a shared mailbox, nobody has been asked to act. The test of monitoring is not how many indicators you track. It is how long a red indicator stays red.

Exceptions: controlled flexibility or a second policy

No policy anticipates every situation. A team needs a tool the standard has not yet approved; a deadline means an evaluation will finish a week after launch. A mature system plans for this. An immature one pretends it will not happen, and the exceptions happen anyway, informally, where nobody can see them.

A recorded, conditioned and expiring exception is controlled flexibility; an informal open-ended one becomes a hidden second policy.ExceptionrequestedRecorded, conditioned, expiringControlled flexibilityInformal and open-endedA hidden second policyEVERY EXCEPTION RECORDSThe ruleRationaleCompensating controlsOwnerApproverExpiry
Figure 7.11.7 An exception is a decision with an end date. Without one, it becomes the rule.

A controlled exception has a short, fixed anatomy. It names the rule being departed from and the reason. It states the extra risk and the compensating controls that will hold it down, such as a human review of every output until the evaluation is complete. It has an owner who answers for it and an approver who is not the person asking. It has an expiry date, and when the date arrives the exception lapses unless someone makes a fresh decision; it does not renew with a click. All exceptions sit in one register that the governance office can see, and expired ones show up on the control dashboard like any other failed control.

Two neighbors handle the edges. When a business chooses to live with a known residual risk rather than depart from a rule, that is risk acceptance, covered in AI Evaluation and Approval Gates. When the same exception is requested again and again, that is a signal that the policy itself no longer fits, as AI Policy vs AI Governance argued. The most dangerous exception is the one nobody recorded: a step that was simply skipped, once, and then again.

Audit is separate, and closes findings in fact

Monitoring and audit are often confused, and the confusion weakens both. Monitoring is management watching its own controls. Audit is a separate evaluation by people who do not run those controls, reporting to the board.

Monitoring is management's continuous view of what is happening; audit is an independent, risk-based judgment on whether the control system can be trusted.MonitoringAuditAsksWhat is happening now?Can we trust the control system?Done byManagement and second-line specialistsInternal audit, independent of the workRhythmContinuous or frequentPeriodic and risk-basedProducesAlerts and actionsFindings, owners and verified fixes
Figure 7.11.8 Monitoring is management’s view. Audit is a second, independent look at whether that view is accurate.

AI Roles, Ownership and Accountability placed internal audit in the third line of the IIA’s Three Lines Model, providing independent and objective assurance6. The audit profession’s own standards show why that independence has to be protected in practice. Internal auditors must not assess activities they were responsible for, and their objectivity is presumed impaired if they give assurance on something they ran within the previous twelve months7. That is the practical reason not to put internal audit on the approval panel for AI systems. An auditor who approved a system cannot later judge, with any credibility, whether the approval process worked. Audit can advise on how a control is designed; management must own and operate it.

Audit needs a trail to follow. For any material AI system, the organization should be able to reconstruct what was decided, by whom, on what evidence, under which version of which standard, which exceptions applied, and what has changed since. If the trail is captured as the work happens, in the release pipeline, the inventory and the exception register, an audit is a matter of sampling. If it has to be assembled afterwards, the audit mostly measures how good people are at remembering.

The last step is where programs often fall short. An audit finding is closed when the control has changed and the change has been tested, not when a plan or a new document has been written. The IIA’s standards require internal auditors to confirm that management has actually implemented agreed actions, with risk-based follow-up and a tracking system. Where management has not acted by the agreed date, the chief audit executive must decide whether management, “by delay or inaction”, has accepted a risk beyond the organization’s tolerance7. ISO/IEC 42001 makes the same demand of the organization itself: nonconformities lead to corrective action whose effectiveness is reviewed2. Overdue findings are not an administrative backlog. They are decisions to live with a known weakness, made by default.

Story: the pitch that passed every check but one

One of the most candid recent post-mortems of a control failure involving AI comes from a national technology magazine, a publication whose staff report on AI for a living and whose newsroom includes a dedicated team of fact-checkers. In August 2025 it published an unusually frank account of how it had been deceived8.

A magazine published a fabricated freelance story in May 2025, was alerted by its payments process, retracted it and later published a post-mortem admitting it skipped its own checks.7 Apr 2025Pitch arrivesA first-timefreelancecontributor7 MayStorypublishedEdited; noalarms raisedDays laterPayments flagWriter cannot be setup for paymentMayRetractedAI detectors hadsaid human-written21 AugPublicpost-mortemFact-check andsenior edit skipped
Figure 7.11.9 The editorial controls existed and were skipped. The control that caught the problem sat in finance.

On 7 April an editor received a pitch about couples who hold their weddings in online games. It was, in the magazine’s words, the kind of story it might have built in a lab. The editor assigned it, the writer responded promptly to edits, and the story ran on 7 May. Over the following days the writer could not provide enough information to be set up in the payments system and insisted on being paid through an online payment app or by check. Now suspicious, an editor ran the article through two third-party AI-detection tools. Both said it was probably written by a human. A closer look at the story’s details and further correspondence showed it was an AI fabrication, and after more checking by the head of the research desk the magazine retracted it8. Other publications removed pieces under the same byline9.

The magazine’s own diagnosis is the instructive part. The story “did not go through a proper fact-check process or get a top edit from a more senior editor”, and first-time contributors “should generally get both”8. Read that through this chapter’s layers. The policy was sound: publish only verified journalism. The standard existed: first-time contributors get a fact-check and a senior edit. The control was skipped, without any record that it had been, which makes it an unrecorded exception. The detective control that worked was not an editorial one at all; it was the finance process that verifies who is being paid. The AI-detection tools, which looked like a control, gave false reassurance: a design failure. And the remediation, by the magazine’s account, was “steps to ensure this doesn’t happen again”, which it did not specify.

An auditor reviewing the episode would not ask for a new policy. The question is whether the senior edit and fact-check for first-time contributors are now enforced by the workflow, for example by a publishing system that will not release a first-time byline without both sign-offs, and whether anyone has tested that block. That is the difference between a closed finding and a closed file.

What this means for leaders

The survey that opened this chapter suggests that, for large organizations, the question is no longer whether they have AI governance. The better question is whether it operates, and whether they could show that it does. Leaders can push that question in four ways. Ask for standards that come with their tests. Ask for control health, not only system health, on the dashboards you see. Treat the exception register as a list of decisions with dates, not a list of favors. And measure audit by verified fixes, not by findings raised.

Check yourself

  1. A clear, published AI policy is enough for an organization to show its AI is governed.
  2. A control can be well designed and still fail.
  3. Monitoring dashboards make independent audit unnecessary.
  4. Internal audit should sit on the approval panel for high-risk AI systems.
  5. A well-run exception has an approver other than the requester and an expiry date.
  6. An audit finding can be closed once management has written a remediation plan.

Reflection: the control nobody tested

What comes next

This chapter completes the control and oversight mechanisms of AI governance: inventory and classification, lifecycle controls, approval gates, and now the standards, monitoring, exceptions and audit that keep them honest. The final chapter of the module, Module 07 Synthesis — Governing AI at Scale, puts all of these pieces together into a single executive governance framework that can grow with the number of AI systems you run.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

Sarbanes-Oxley Act (internal control over financial reporting) · US - federal (SEC-registered companies)

Sarbanes-Oxley Act of 2002, Sections 302 and 404 (Public Law 107-204)

Section 302 requires the CEO and CFO to certify each periodic report and the effectiveness of disclosure controls. Section 404 requires management to assess internal control over financial reporting each year, with auditor attestation for larger filers. AI used in close, reconciliation or reporting processes falls inside these controls: its outputs need the same evidence, review and change control as any other step that affects the financial statements.

  • 2002-07-30 — Signed into law

Last verified 2026-10-08 · official text

References

  1. American Arbitration Association. Most Organizations Have AI Governance; Few Say It Works in Practice, New American Arbitration Association Survey Finds. American Arbitration Association. 2026.
  2. ISO/IEC. ISO/IEC 42001:2023 Information technology - Artificial intelligence - Management system. International Organization for Standardization. 2023.
  3. Committee of Sponsoring Organizations of the Treadway Commission (COSO). Internal Control - Integrated Framework (2013). COSO. 2013.
  4. Public Company Accounting Oversight Board. AS 2201: An Audit of Internal Control Over Financial Reporting That Is Integrated with An Audit of Financial Statements. PCAOB. 2007.
  5. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
  6. The Institute of Internal Auditors. The IIA's Three Lines Model: An update of the Three Lines of Defense. The IIA. 2020.
  7. The Institute of Internal Auditors. Global Internal Audit Standards. The IIA. 2024.
  8. WIRED. How WIRED Got Rolled by an AI Freelancer. Conde Nast. 2025.
  9. Charlotte Tobitt. Wired and Business Insider remove 'AI-written' freelance articles. Press Gazette. 2025.

Further reading

Sources last verified 2026-10-08.