AI Academy · Book
Executives & Directors · Module 06 · Chapter 002

Accuracy, Hallucination and Reliability

AI systems will sometimes be fluent, confident and wrong, and a high average score can hide exactly where. Reliability is not a number the model brings with it. It is designed around the model: measure it on your real work, ground it, check what can be checked, and let it say "I don't know".

≈ 16 min read

After this chapter you can

  • Explain what a hallucination is and why fluent answers can be invented.
  • Explain why accuracy-only scoring rewards guessing, and ask for right, wrong and declined rates.
  • Test whether an accuracy figure represents the real work, slice by slice.
  • Distinguish grounded answers from true ones, and confidence from calibration.
  • Place automatic checks and abstention where the model cannot talk its way past them.

In July 2025, Australia’s Department of Employment and Workplace Relations published an independent assurance review of the IT system it uses to apply automated penalties in the welfare system. The review came from Deloitte Australia, under a contract worth 440,000 Australian dollars. Weeks later a Sydney University researcher of health and welfare law, Chris Rudge, told journalists the report was “full of fabricated references”. It cited academic work that does not exist and quoted a federal court judgment with words the judge never wrote. A corrected version followed, this time disclosing that a generative AI tool built on Azure OpenAI had been used in drafting. Deloitte agreed to repay the final installment of its fee. The department said the findings and recommendations had not changed1.

Nothing about the original report looked wrong. It was long, structured, footnoted and written in the confident register of professional assurance. That is the point. The errors survived drafting, review and publication because they looked exactly like the correct parts. An answer that says “I don’t know” invites someone to check. An answer that is fluent, specific and wrong slips past the check.

Fluent, wrong and expected

The AI Risk Landscape showed that risk depends on what the business allows a system to do with its answers. This chapter looks inside the first family on that map: why AI systems produce wrong answers, why those answers are so convincing, and how to build a system that catches them.

Start from a settled fact. Every AI system will sometimes be wrong, and no vendor can promise otherwise. The useful question is not whether the model errs but whether the system around it makes important errors less likely, detects them when they happen and stops them before they reach a customer, a payment or a decision. NIST’s AI Risk Management Framework names “valid and reliable” as a property of trustworthy AI, and defines reliability as the ability to perform as required, without failure, for a given time and under given conditions2. Conditions are the key word. Reliability belongs to a system doing a job, not to a model in a lab.

A chain from question to sources, model, checks and action; reliability belongs to the whole chain, not the model.QuestionClear orambiguousSourcesRetrieved,current,readableModelWrites theanswerChecksRules andsourcetestsActionPayment ordecisionA good model in a weak chain still fails.
Figure 6.2.1 Errors can enter at every link, and every link is also a place to catch them. Reliability is a property of the chain.

Read the chain from left to right and you can see where a wrong answer is born. The question may be ambiguous. The system may retrieve the wrong document, an old version or a page it cannot read. The model may fill the gap with something plausible. And a missing check may let the result flow straight into action. Each link is also a place where the error can be caught, which is why the rest of this chapter walks the chain.

What a hallucination is

A hallucination is output that presents unsupported, invented or incorrect content as if it were valid: an invented fact, a citation to a paper that does not exist, a policy clause that was never written. NIST’s profile for generative AI prefers the word confabulation and lists it among the risks unique to, or made worse by, generative AI3. Researchers distinguish two kinds. An intrinsic hallucination contradicts the source the system was given, such as a summary that reverses a contract term. An extrinsic one adds content that the source cannot support, such as the fabricated court quote4.

The cause lies in what a language model does. It generates likely sequences of words from patterns learned in training. It is not a database of verified facts, and nothing in the generation step checks the words against the world. Plausible language and verified truth usually coincide, which is why these systems are useful. When they come apart, the output keeps its fluency.

The visible qualities of a fluent answer prove nothing about the hidden requirements of a correct one.WHAT THE ANSWER SHOWSGood grammar · Clear structure ·Confident tone · Specific detailWHAT CORRECTNESS NEEDSA real sourceThe current versionNumbers that add upThe right context
Figure 6.2.2 Nothing above the waterline proves anything below it. Fluency is not evidence.

This is why OWASP’s list of the main security risks in language-model applications pairs misinformation with overreliance: people trust credible-sounding output, skip the verification and build it into decisions5. A typo warns the reader. A fluent, specific, wrong answer does the opposite.

Why models guess

If models can be uncertain, why do they not say so? A 2026 paper in Nature gives an uncomfortable answer: we taught them not to. Adam Tauman Kalai and colleagues show that the way models are scored rewards guessing. Most of the benchmarks they reviewed grade an answer as right or wrong, and “I don’t know” earns the same zero as a wrong answer. Under that rule, a model that always guesses beats one that admits uncertainty, just as a student who guesses on every multiple-choice question outscores one who leaves blanks6.

The paper’s own example makes the trade visible. On SimpleQA, a set of short factual questions, two models from the same developer scored almost the same on accuracy. They were very different on errors.

On SimpleQA one model was 16 percent correct and 21 percent wrong; the other 21 percent correct and 77 percent wrong.CorrectWrongDeclined to answerModel that abstains16%20.8%63.2%100%Model that guesses20.6%76.8%100%Source: Kalai et al., Nature · 2026
Figure 6.2.3 Similar accuracy, very different reliability. The model that guesses is right a little more often and wrong far more often.

The model that guesses is right on 20.6 percent of questions and wrong on 76.8 percent. The model that abstains is right on 16.0 percent and wrong on only 20.8 percent, because it declines the rest. A leaderboard that reports accuracy alone ranks the guesser first. For an enterprise process, the second model is far easier to make reliable, because its failures announce themselves. The authors propose scoring that states the penalty for an error up front, so that abstaining becomes the rational choice when stakes are high6.

The practical lesson is to ask for three numbers, not one: how often the system is right, how often it is wrong, and how often it declines. Accuracy alone tells you very little about the second.

Accurate on what?

The second trap is the single accuracy figure. Accuracy is how often a system produces the right result for the task being measured, and the task changes what “right” means.

Accuracy means something different for classifying, extracting, answering, summarizing and acting, and each needs its own check.TaskAccurate meansA check that fitsClassify a documentRight categorySample against a reviewerExtract a valueRight field and numberTotals add up; value is in the sourceAnswer apolicy questionRight for our current policyCited version is currentSummarizeFaithful, nothing addedCompare claims with sourceTake an actionTask done correctly and safelyLimits and reversal
Figure 6.2.4 There is no single accuracy number for AI. Each task needs its own definition of right and its own check.

Even for one task, an average can mislead. NIST’s framework says accuracy measurements should always be paired with realistic test sets that represent the conditions of expected use, and that results can be broken down by data segment2. Medical-imaging researchers gave the failure a name, hidden stratification: models with strong overall scores performed poorly on important subgroups of patients that nobody had separated out in training or testing7. The average was true. It was not about the cases that mattered.

Two questions catch much of this. First, who wrote the test, and does its mix of questions match how people will really use the system? Second, what is the score for each important slice of work, not only overall? A test that under-samples one process cannot reveal concentrated failures in it. The story later in this chapter shows how.

Grounded is not the same as true

Many enterprise assistants first retrieve documents and then answer from them. RAG and Enterprise Knowledge — Executive Mental Model explained how this retrieval-augmented design works and why it beats retraining8. For reliability, two properties need to be kept apart. Factuality asks whether an answer is true. Groundedness asks whether it is supported by the source the system was supposed to use. Ask for the company’s travel reimbursement limit and common practice is irrelevant: the only right answer is the one in the current policy.

Grounding helps, and it does not end the problem. The retriever can fetch the wrong document, an outdated version or a scanned page whose text it cannot read, and the model will still write a coherent answer. The legal research tools in What Exactly Is Artificial Intelligence? showed that retrieval narrows the gap without closing it9. One large test of everyday assistants points the same way: in 2025, journalists at public broadcasters in 18 countries checked 2,709 news answers from four assistants.

Nearly a third of those answers, 31 percent, had serious sourcing problems: claims the cited source did not support, no source at all, or sourcing that could not be verified10. A citation is something to check, not proof. It can point to the wrong source, an old version or a real document that does not contain the claim.

Confidence, consistency and the right to say “I don’t know”

Many systems report a confidence score. It helps only if it is calibrated, meaning that answers given with 90 percent confidence are right about 90 percent of the time. That cannot be assumed. Modern neural networks are often overconfident even as they become more accurate11. Before anyone relies on a confidence threshold, test it on your own work.

Consistency is a second signal. Ask a generative system the same question twice and the answers can differ. That is a nuisance where one answer is required, and it is also useful information: when repeated answers disagree in meaning, the system is probably making something up. Researchers have turned this into a hallucination detector that samples several answers and measures how far they agree12. That is a result on research benchmarks; it has not yet been shown to cut errors in business deployments, so treat it as a promising capability. Where a task needs one fixed answer, such as a tax calculation, a date check or an ID match, give it to rules or ordinary software and keep the model for drafting, summarizing and explaining.

An underrated reliability feature is the right to decline.

A reliable assistant answers with a source when evidence is clear, asks when the question is ambiguous, and abstains or escalates when evidence is missing.Clear question,readable currentevidence?YesEvidence foundAnswer with source?Question ambiguousAsk to clarifyNoEvidence missingAbstain or escalate
Figure 6.2.5 A reliable assistant has three moves, not one. Asking and declining are features, not failures.

“What is the approval limit?” Approval for what, for whom, in which region? A design that must always answer will guess. A design that can answer, ask or escalate will sometimes frustrate users and will far more often keep a wrong answer out of the process. The Kalai results put a number on the trade, and it favors the system that knows when to stop.

Checks the model cannot talk its way past

One of the strongest reliability patterns is also the simplest: let the model do the flexible work and let something else verify it. Take an illustrative invoice. An assistant reads it and reports a total of 125,000; a rule confirms that the line items add up to 125,000. An assistant quotes a value from a specification; a second step confirms that the value appears in the cited document and that the document is the current version. These checks are cheap and cannot be argued with, because they do not read the answer’s tone.

The model reads, drafts and explains; deterministic checks confirm numbers, quotations and source versions, each with an owner and fallback.The model doesReads messy documentsDrafts and summarizesExplains in plain wordsChecks verifyNumbers add upQuoted value is in the sourceSource is the current versionEvery check needs an owner and a fallback.
Figure 6.2.6 Flexible work goes to the model; verification goes to rules, lookups and people with authority.

People are checks too, but only real ones. A reviewer adds reliability only with the evidence, the time, the expertise and the authority to override. Otherwise review becomes a signature on whatever the system produced. Why that happens, and how to design oversight that works, is the subject of Human Oversight and AI Incidents. For this chapter, the point is narrower: place automatic checks wherever an error can be detected mechanically, and save human attention for what only judgment can catch.

Story: a quality engineer’s ordinary day

The following is an illustrative composite, not a documented case. A global manufacturer of industrial equipment runs plants on three continents. Last year its engineering function launched an assistant that answers questions about internal procedures and supplier specifications. Before launch, a team at headquarters wrote 200 test questions and checked the answers. The assistant got 194 right: 97 percent. Plants adopted it quickly.

Follow one supplier-quality engineer through one day at one plant.

A quality engineer's day built on confident assistant answers looked ordinary from the first delivery to the final report.07:10Castings arriveAsks for thehardness range08:30Batch releasedResults inside thequoted range11:00Answers a buyerCoating thickness froma spec15:30Signs the reportA normal day
Figure 6.2.7 Illustrative composite. Every step looked routine. The assistant’s answers were clear, specific and cited the right documents.

At ten past seven a shipment of castings arrives. She asks the assistant for the hardness range in the supplier’s specification. The answer is crisp, gives a range and cites the specification by number and section. Her inspection results fall inside it, and at half past eight she releases the batch to the line. Before lunch a buyer asks about a coating thickness and she answers the same way. At half past three she signs the receiving report. Nothing hesitated and nothing warned her. Engineers in other plants had days just like it.

Three weeks later, an engineer at a sister plant opens the specification itself to settle a dispute and finds a different range. The tables in many supplier specifications were scanned images. The assistant’s document pipeline could read the text around them but not the numbers inside them, so the model received a section heading and no values. It did what models are scored to do. It produced a plausible range typical of that material and cited the section it had been given. The citation was real. The number was not in it.

Why did 97 percent not reveal this? Because the test barely sampled the work that failed.

Table lookups were 10 of 200 test questions but 30 percent of plant questions, so a 97 percent test score meant about 87 percent in use.Question typeTest questionsWrongAccuracyShare ofplant questionsProcedures andnarrative text190299%70%Values from tables10460%30%Overall200697%About 87% in use
Figure 6.2.8 Illustrative figures. The signal was in the test, but averaged away. Weighted by real plant use, 97 percent becomes about 87.

The headquarters team wrote its questions from procedures and narrative text, which it knew well. Only 10 of the 200 asked for a value from a table, and 4 of those 10 were wrong. With 2 other errors, the total came to 6 wrong and 194 right. The four table failures looked like scattered noise. (These figures are part of the illustration.) On the plant floor, though, about 30 percent of questions asked for a value from a table. Weight the test results by real use, 70 percent at 99 percent and 30 percent at 60 percent, and the system was right on about 87 percent of plant questions, with more than nine in ten of its errors in exactly the work the plants relied on most.

Was the model inaccurate, or was the system unreliable? Look at the chain. The question was clear. The source was the right document but unreadable. The model guessed instead of declining. No check confirmed that the quoted range appeared in the cited section. And the evaluation did not represent the work. The fixes follow the same chain: a test set drawn from logged plant questions and reported by slice, a rule that the quoted value must appear in the source text, abstention when a section cannot be read, and an owner for the specification library. None of them requires a better model. All of them are management decisions, and the errors belonged to the organization that ran the system.

What this means for leaders

Four habits follow. First, ask “accurate on what?” before accepting any score. Insist on test questions that match real use and on results broken down by the slices of work that matter. Second, ask for three numbers: right, wrong and declined. A system that declines honestly is easier to make reliable than one that guesses well. Third, put checks where the model cannot talk past them: arithmetic, quotations, source versions and limits. Fourth, reward the system and the people for saying “I don’t know”, and design the escalation path before launch rather than after the first incident.

Check yourself

  1. If the assistant cites a source, the answer is supported by that source.
  2. Grounding answers in your own documents ends hallucination.
  3. A model with slightly lower accuracy can be the more reliable choice.
  4. A system can score 97 percent in testing and fail one important process badly.
  5. A high confidence score means the answer is probably right.
  6. An assistant that asks a clarifying question has failed.

Reflection: find the untested slice

What comes next

A system can be accurate and still be unfair. Errors that are rare overall can fall disproportionately on one group of people, in the same way the table lookups hid inside the 97 percent. The next chapter, Bias and Fairness, looks at where bias comes from and why apparently neutral systems can produce unequal outcomes.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

EU Product Liability Directive (revised) · EU

Directive (EU) 2024/2853

No-fault liability now explicitly covers software, including AI systems and SaaS, and updates or the lack of security updates. Easier proof for claimants with complex products.

  • 2026-12-09 — Applies to products placed on the market from this date

Last verified 2026-10-06 · official text

References

  1. The Associated Press. Deloitte to partially refund Australian government for report with apparent AI-generated errors. Associated Press (via U.S. News & World Report). 2025.
  2. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
  3. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
  4. Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.
  5. OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
  6. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang. Evaluating large language models for accuracy incentivizes hallucinations. Nature 653(8116), 1047-1051. 2026.
  7. Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro and Christopher Re. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL '20), 151-159. 2020.
  8. Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
  9. Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies 22(2), 216-242. 2025.
  10. European Broadcasting Union and BBC. News Integrity in AI Assistants: An international PSM study. European Broadcasting Union. 2025.
  11. Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Q. Weinberger. On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning, PMLR 70. 2017.
  12. Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn and Yarin Gal. Detecting hallucinations in large language models using semantic entropy. Nature 630, 625-630. 2024.
  13. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 15: Accuracy, robustness and cybersecurity. Official Journal of the European Union. 2024.

Further reading

Sources last verified 2026-10-08.