AI Academy · Book
Executives & Directors · Module 05 · Chapter 012

AI in Document and Data Processing

Document AI pays off when information trapped in paper, scans and messy files becomes checked data that a real process uses. Reading the page is the easy part. The value is decided by the chain around the reader: validation, routing by error cost, traceability and the people who handle exceptions.

≈ 14 min read

After this chapter you can

  • Distinguish character recognition from document understanding.
  • Describe the processing chain from ingest to workflow action.
  • Assign each check to AI, rules, records or people.
  • Set automation thresholds from the cost of errors, not from an automation target.
  • Choose a first document process by volume and error cost, against a measured baseline.

In the summer of 2013, a German computer scientist named David Kriesel was looking at scanned copies of construction plans and noticed something odd. On the copy, one room was labeled at about 22 square meters, while the room next to it, visibly larger, was labeled 14. In a cost table, 65 had become 85 and 60 had become 80. The originals were correct. The office scanner had changed the numbers1.

Before reading on, make a prediction. How long had scanners of that type been quietly rewriting digits before someone noticed?

The answer is about eight years. The cause was a compression method that saved space by spotting characters that looked alike and storing one shape for all of them. When it judged that a 6 looked enough like an 8, it printed an 8. Xerox confirmed the defect within ten days, said it could occur in all compression modes, and began shipping patches to hundreds of thousands of devices on 22 August 20131.

Office scanners silently swapped digits in scanned documents for about eight years; a 65 became an 85 on a clean-looking copy.8 yearsBefore anyone noticedDigits swapped with no warning65 → 85One cost-table entryThe copy looked perfectly cleanSource: D. Kriesel, with Xerox's confirmation · 2013
Figure 5.12.1 A machine that reads documents can be confidently wrong, and nothing on the page tells you so.

Be clear about what this was: a compression bug in ordinary scanning software, not artificial intelligence. No model was involved. It is still a useful warning for anyone about to put AI to work on documents. A machine that reads produced output that looked flawless, gave no sign of doubt, and was wrong in exactly the places that mattered. What caught it was not the machine. It was a person who noticed that the numbers did not make sense against the drawing.

From trapped information to reliable inputs

Organizations run on information that never fully entered a system. Contracts, delivery notes, certificates, claims, application packs, supplier letters and spreadsheets carry the facts that finance, procurement, legal and operations depend on. A document is also more than its words. It has tables, signatures, stamps, layout and context: the total sits at the bottom right for a reason, and a handwritten note in the margin can override the printed terms.

The executive problem is not that there are too many files. It is that high-volume work is held back by information that a person must read and retype before any system can use it. Scanning alone does not fix that. A searchable file is easier to find, but someone still has to open it, read it and key the values into the system that acts on them.

The idea of this chapter fits in one sentence. The value of document AI is not extracting text. It is turning information trapped in documents and messy data into checked inputs that a business process can act on, with a trace back to the source and a person in the loop where errors are costly.

Characters versus meaning

The first reaction in many leadership teams is that this was solved years ago: “We already have OCR.” Optical character recognition is useful, and it answers one question: which characters are visible on this page? The result is a page of text.

OCR answers which characters are on the page and yields text; document AI answers what the document means and yields a usable record.OCR ASKSWhich characters are onthe page?A page of textDOCUMENT AI ASKSWhat does it mean, and whatdo we need?A usable recordvs
Figure 5.12.2 Character recognition produces text. Document understanding produces a record a process can use.

Document understanding asks a different question: what does this document mean, and what does the business need from it? For a supplier invoice, the business does not need a page of characters. It needs the supplier, the invoice number, the date, each line item, the tax, the total, the purchase order and the payment terms, each in the right field. The step change came when models learned text and page layout together, so that position on the page, table structure and labels inform what a value is2. Multimodal language models now take that further.

Two capabilities matter most in practice. Classification routes each item: an invoice to accounts payable, a claim to claims, a change order to the project team. Table extraction keeps rows together, so that a quantity stays attached to its price instead of dissolving into a stream of words. Both are where character recognition alone breaks down, and both are where modern models score well in published benchmarks. That is a capability; how it holds on your own documents, and what it saves, has to be measured in your process. The IRS story later in this chapter is one public record of what a deployment actually delivered.

They also bring a new failure. A language model that cannot read a smudged figure does not leave a blank. It can produce a plausible value, the same failure the scanner showed in 2013, now with far more fluency. Researchers call this hallucination, and risk frameworks call it confabulation [@ji-2023-hallucination; @nist-ai-600-1-2024]. That is why the rest of the design matters more than the reader.

The processing chain

Value does not appear when a value is extracted. It appears when the extracted value enters the system that acts on it.

Documents flow through ingest, extract, validate, and act and learn; value appears only when the output enters a workflow.IngestKeep source,owner,permissionsExtractFields, tables,amountsValidateClassify andcheckAct and learnPost, route,escalateValue appears when the output enters a workflow.
Figure 5.12.3 Four steps turn a document into an action. Much of the engineering, and much of the risk, sits after extraction.

Ingest keeps the facts about the document: where it came from, when it arrived, its identifier, its owner and who may see it. Extract pulls out the fields, tables, names and amounts. Validate asks whether the item is what it claims to be and whether the values hold up: does the total equal the sum of the lines, is the supplier known, does the purchase order exist? Act and learn posts, routes or escalates, archives the source and feeds corrections back so the same mistake is caught next time.

Invoice processing is the simplest picture of the whole chain. The system classifies the document, extracts the fields, then performs a three-way match of purchase order, goods receipt and invoice. A match can go straight through. A mismatch goes to a person. Contracts follow the same pattern: extract parties, dates, amounts, renewal terms and unusual clauses, then send them for legal review, with the AI highlighting differences from the standard rather than drawing legal conclusions. A document chatbot that answers questions about a file can be handy. The large gains come when the output lands in the finance system, the payment run or the case system.

Give each job to the simplest reliable tool

It is tempting to send every document through the most capable model available. It is rarely the best design. As AI vs Automation argued, rules and learned models fail in different ways, and good systems use each where it is strongest. In document processing the division of labor is unusually clear.

A hybrid design gives ambiguity to AI, fixed checks to rules, verification to records and exceptions to people.PartBest atInvoice exampleAIAmbiguityUnfamiliar layoutsRulesFixed checksLimits, duplicatesRecordsVerifying factsPurchase-order matchPeopleExceptionsLarge mismatches
Figure 5.12.4 Each part of the design does the job it does best. The model is one part, not the whole system.

AI earns its place on ambiguity: unfamiliar layouts, free-text terms, handwriting, a supplier who puts the total somewhere odd. Rules handle fixed checks: is the date valid, is the amount under the limit, is a required field present, has this invoice number been paid before? A rule answers the same way every time, costs almost nothing to run and can be read by an auditor. Records verify facts against the systems of record: does this supplier exist, were these goods received, is this the bank account we hold on file? People handle the exceptions and the judgment calls.

The record check deserves emphasis because it is the one teams skip. A model can read a bank account number perfectly from a forged letter. Only a comparison with the account you already hold, and a call-back when it differs, catches the fraud. Ask of each step: what is the simplest technology that is reliable enough here?

Confidence picks the route; error cost sets the threshold

Good document systems attach a confidence level to each extracted value, and the confidence decides the route.

High-confidence extractions post automatically under rules, medium ones are checked against records, and low ones go to a person.How confident isthe extraction?High confidencePost automatically; rulesstill applyMedium confidenceCheck against recordsLow confidencePerson reviews thesource page
Figure 5.12.5 Confidence chooses the route. Where to draw the lines between routes is a business decision.

Two cautions apply. First, a confidence score is only useful if it is calibrated: a system that says it is 95 percent sure should be right about 95 percent of the time. Modern neural networks are often overconfident, so calibration has to be tested on your own documents, as Accuracy, Hallucination and Reliability explains in Module 063.

Second, errors are not equal, and that is what should set the thresholds. Sending a correct invoice to review costs a few minutes of someone’s time. Paying a wrong one can cost far more, and the money comes back slowly if at all. So the design that automates the most is not automatically the cheapest. In one illustrative comparison, an accounts payable design that automated less, with rules and record checks outside the model, cost less than half as much as one that chased the highest straight-through rate, because far fewer wrong payments got through. Test the shape with your own volumes and error costs.

The automation rate was never the right target. The objective is the maximum safe and economically valuable automation, and the threshold that achieves it belongs to the process owner who bears the cost of errors, not to the data team.

Every value must trace back to its page

For any value that moves money or changes a decision, someone will eventually ask: where did this come from?

Every important extracted value should trace to its source document, page, time, model version, checks and approver.ValueOne extractedinvoice totalDocumentPageTimeModelChecksApprover
Figure 5.12.6 Provenance turns a clean-looking number into a defensible one.

The answer is called provenance: which document, which spot on which page, when it arrived, which model and version extracted it, which checks it passed, and who approved it, if anyone did. With provenance, a reviewer can check a doubtful value in seconds by looking at the highlighted spot on the original page, and an auditor or an unhappy supplier can follow a payment back to its source.

Without provenance, the system produces numbers that look clean and cannot be defended. The 2013 scanner is the warning here too. An organization that had scanned its paper and then destroyed the originals would have had no way to tell which digits had changed. Keep the source, keep the link, and make the link easy to follow.

Documents also carry personal, financial, health and employee data, and access, purpose and retention rules still apply after extraction. Privacy and Confidential Data, in Module 06, takes that up.

Story: the reader was never the bottleneck

The United States Internal Revenue Service offers a documented, public post-mortem of trying to move a document-heavy process onto machines.

In March 2022, the National Taxpayer Advocate described how the agency handled paper tax returns. Each digit on every paper return was keyed into IRS systems by an employee, several hundred digits for a moderately complex return. The backlog stood at nearly 15 million returns. Employees made transcription errors on about 22 percent of paper returns. And 50 to 60 percent of the paper returns had been prepared with tax software: the data had existed in digital form, been printed, mailed and typed in again4.

In August 2023 the Treasury committed that, by the 2025 filing season, the IRS would digitally process all paper-filed tax and information returns, up to 76 million documents a year5. In February 2026 the Treasury Inspector General for Tax Administration reported how it went6.

The IRS aimed to digitize all paper returns by 2025 but digitized 7 percent in pilots and 5 percent in the 2025 season.Goal for the 2025filing season100%Pilots, 2023 to 20247%2025 filing season, to May5%Source: TIGTA report 2026-408-003 · 2026
Figure 5.12.7 The pilots proved paper returns could be read and processed. The chain around the reader could not keep up.

The pilots worked. Contractors scanned returns, extracted the data and sent it to the IRS for processing, and the inspector general found that they “successfully proved” paper returns could be digitized and processed. But they covered 3.8 million of the 53.3 million paper forms received from February 2023 to December 2024, 7 percent. In the 2025 filing season, to May, contractors scanned 517,000 of the 9.8 million paper forms received, 5 percent. Both shares count the same three forms (Forms 940, 941 and 1040) against what arrived on paper in each period; neither is a share of all returns filed. An in-house system was stopped in April 2025 after nearly 61 million dollars had been spent, before it had scanned a single return6.

The reasons are instructive because none of them is about reading accuracy. The in-house system waited ten months for approval, and its contract had to be cancelled and re-awarded. Scanning volume was limited by people: staff to open mail, prepare and scan documents and quality-review the images. One contractor estimated it needed about 600 people in that pipeline, and each needed a background clearance that took four to five weeks. In September 2025 the IRS selected four contractors, with awards totaling 2.3 billion dollars through 20306.

The economics show why it was worth the effort. Paper individual returns cost 43 times more to process than electronic ones, and in the 2025 season they consumed 72 percent of processing costs while making up 6 percent of returns6.

Three lessons travel well. First, the reader is the smallest part of the system: intake, review capacity, contracts and sequencing decided the outcome. Second, measure the baseline before you judge the machine. A human process with errors on more than one return in five is not the gold standard a pilot must match. Third, the cheapest document is the one you never print. Half the paper returns started as digital data; the best design captures data where it is born.

What this means for leaders

The first decision is where to start. Two questions place any document process on a simple grid: how much volume is there, and what does an error cost?

Start document AI where volume is high and errors are cheap; keep a person approving where errors are costly.HighLowCost ofan errorLowVolume · HighAssist the expertSpecialist decidesHuman in the loopPerson approvesLeave for laterLittle gainStart herePeople monitor
Figure 5.12.8 Start where volume is high and errors are affordable; keep a person approving where errors are costly.

High volume with a low error cost is the natural first candidate: classifying incoming mail, extracting fields from routine forms, posting small matched invoices. Here the system processes and people monitor and override, a pattern called human on the loop. Monitoring is harder than it sounds. Lisanne Bainbridge showed in 1983 that people asked to watch an automated process lose the practice and attention they need to step in when it fails, so sample and review the automated flow on purpose7.

High volume with a high error cost is often where much of the money sits, and much of the risk. Use human in the loop: the system extracts and recommends, a person approves, then the process runs. Low volume with a high error cost is about helping the specialist read faster, not replacing the review. A document agent that receives, extracts, validates and escalates on its own is a later question, covered in AI Agents and Intelligent Workflows.

Check yourself

  1. OCR and AI document processing are the same thing.
  2. If average extraction accuracy is high, straight-through processing is safe.
  3. For amount limits, date formats and duplicate checks, rules usually beat AI.
  4. A machine-read document that looks clean can still contain wrong values.
  5. In the IRS program, reading accuracy was what held paperless processing back.
  6. Once documents are scanned and searchable, the processing problem is solved.

Reflection: find the retyped data

What comes next

Turning documents into reliable data gives the organization something to reason with. The next chapter, AI in Prediction, Forecasting and Optimization, looks at what AI does with that data: forecasting what is likely to happen, and deciding what to do about it.

References

  1. David Kriesel. Xerox scanners/photocopiers randomly alter numbers in scanned documents. dkriesel.com (blog, with dated updates). 2013.
  2. Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei and Ming Zhou. LayoutLM: Pre-training of Text and Layout for Document Image Understanding. Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2020). 2020.
  3. Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Q. Weinberger. On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning, PMLR 70. 2017.
  4. Taxpayer Advocate Service, Internal Revenue Service. Getting Rid of the Kryptonite: The IRS Should Quickly Implement Scanning Technology to Process Paper Tax Returns. National Taxpayer Advocate blog. 2022.
  5. U.S. Department of the Treasury. IRS Launches Paperless Processing Initiative. U.S. Department of the Treasury, press release JY1666. 2023.
  6. Treasury Inspector General for Tax Administration. The IRS Has Made Limited Progress Achieving Paperless Processing. TIGTA, Report Number 2026-408-003. 2026.
  7. Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.

Further reading

Sources last verified 2026-10-10.