AI in Prediction, Forecasting and Optimization
AI has made good predictions cheap, and the best of them now arrive as ranges, not single numbers. The value still depends on the decision they feed. Leaders who ask for the range, price both directions of error and check whether a question needs a prediction or a cause get the benefit; the rest get a sharper number and the same plan.
After this chapter you can
- Distinguish prediction, forecasting and optimization, and the decision each one feeds.
- Ask for a forecast as a range and explain average error in plain words.
- Price the cost of being wrong in each direction and lean the plan toward the cheaper mistake.
- Tell a decision that needs only a prediction from one that needs evidence of cause, and ask for the test.
- Set the written objective and the hard and soft limits an optimization must respect.
In December 2024, Nature published a weather model from Google DeepMind called GenCast. It does not produce one forecast. It produces a whole set of possible fifteen-day futures for the planet, and each one takes about eight minutes on a single AI chip. Tested against the ensemble of the European Centre for Medium-Range Weather Forecasts, long the benchmark, GenCast was more skillful on 97.2 percent of 1,320 measures, and better on extreme weather and the tracks of tropical cyclones1. In July 2025 the European Centre put its own AI ensemble into daily operation: 51 forecasts per run, each using about a thousand times less energy than the physics-based system2. The two results are different kinds of evidence. GenCast’s lead is a research result, scored against past weather; the European Centre’s ensemble is a deployment, now run every day alongside the physics-based system.
The detail that matters for a leader is not the speed. It is the shape of the answer. Leading forecasting systems now say “here is the spread of what may happen”, not “here is what will happen”. Yet in many companies the forecast that reaches the executive committee is still a single number: next quarter’s demand, the cash position in March, the call volume on Monday. Ask what the chance is of landing 10 percent above it, or what it costs if the number is wrong, and the room usually goes quiet.
A forecast is an input, not a decision
The economists Ajay Agrawal, Joshua Gans and Avi Goldfarb describe AI as a fall in the cost of prediction, and the module opener and Finding Strategic AI Opportunities made the consequence plain: a prediction pays only when it changes what someone does3. This chapter is about the craft of making that happen across the enterprise, in finance, HR, sales, service, maintenance and pricing alike.
Three habits do most of the work. Ask for the range, because every forecast is uncertain and a single number hides how uncertain. Price the errors, because being wrong in one direction often costs more than being wrong in the other. And ask whether the question needs a prediction or a cause, because a pattern that forecasts well can still mislead you about what an action will change. When the decision is complex enough to hand to an optimizer, a fourth habit follows: write down the objective and the limits, because the system will pursue exactly what you wrote.
Three jobs before the decision
Three words are used as if they meant the same thing. They are three jobs.
Prediction estimates something we do not yet know about a case: will this customer leave, will this pump fail this month, is this transaction fraudulent? The honest answer is a probability, not a verdict. Forecasting estimates a quantity over time: calls next week, cash at quarter end, rooms booked on a festival weekend. Optimization searches for the best action given a forecast, an objective and limits: which roster meets the service target at lowest cost, which price for each room tonight.
AI can help with all three, and the boundaries matter because each one fails differently. A churn prediction can be accurate and useless if nothing is done with it. A forecast can be right on average and wrong on the days that count. An optimizer can find a brilliant answer to the wrong question. The supply-chain versions of these jobs, from the bullwhip to route planning, belong to AI in Operations and Supply Chain; the habits below apply everywhere.
Ask for the range
Every forecast is a distribution of possible outcomes, and the single number on a slide is usually its middle4. A useful forecast says both: the central figure and the span within which reality is expected to fall, for example “22, and probably between 14 and 30”. Forecasters call that span a prediction interval. Executives can simply call it the range.
Consider an illustration. Suppose a hotel group forecasts, for each night of a busy week, how many booked guests will not turn up. The central forecast for Friday is 22 rooms. The range runs from 14 to 30, and it is wider on Friday and Saturday than midweek because weekend leisure guests are less predictable.
You will also hear accuracy reported as an average error, often as a percentage; the acronym MAPE, mean absolute percentage error, is the common one. In plain words it says: on a typical day, the forecast misses by about this much. That is worth knowing. But an average blends the quiet days with the days that matter, so a model can improve its average while getting worse on the rare, expensive days. Ask to see the error on those days separately.
The second question is whether the range itself can be trusted. If a team says it is 80 percent confident, reality should land inside the range about eight times in ten; Accuracy, Hallucination and Reliability in Module 06 calls this calibration. And because customers, markets and weather move, a forecast that was well calibrated last year can quietly drift, as How Machine Learning Learns explained. Checking the range against reality is a standing duty, not a launch task.
Price both directions of error
Once the range is on the table, the question most dashboards never answer becomes unavoidable: what does it cost when the forecast is wrong, in each direction?
Return to the hotel. If it overbooks by too few rooms, some rooms sit empty; suppose each costs 150 in lost revenue. If it overbooks by too many, a guest with a reservation is turned away. The hotel pays for a room elsewhere, transport and a gesture of apology, and risks a loyal customer: suppose 600. Walking a guest costs four times as much as an empty room.
That ratio tells you how to use the range. Planning to the central forecast of 22 would mean walking guests on about half of all Fridays. The better rule is to plan so that the expensive mistake happens only about one time in five, because 150 divided by the sum of 150 and 600 is one fifth. In this illustration that means overbooking by about 17 rooms, not 22. The plan deliberately leans toward the cheaper mistake.
The same logic runs through the enterprise. A blood bank short of a rare type and an office over-ordering paper are not the same mistake. Missing a failing turbine bearing costs far more than an unneeded inspection. Flagging a loyal customer as a fraud risk costs differently from missing a fraudster. AI in Document and Data Processing used the same reasoning to set confidence thresholds. Notice who has to supply the numbers: not the data science team, but the business owners who know what a walked guest or a stopped line really costs. Pricing the errors is leadership work.
Prediction or cause?
Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan and Ziad Obermeyer drew the distinction with two decisions about rain5. Whether to take an umbrella depends only on predicting the weather; your umbrella will not change it. Whether to do a rain dance depends on a causal question: does dancing bring rain? Many business decisions are umbrellas. Which machine to inspect first, how many agents to schedule and how many rooms to overbook need only a good prediction. Others are rain dances. Whether a retention offer keeps customers, whether a price cut grows volume and whether a training program reduces errors are questions about what an action will change.
A predictive model is built to find patterns, and a pattern can predict well for reasons that make it dangerous as a guide to action. One well-known pneumonia model learned that asthma lowered the risk of death, because asthmatic patients were sent straight to intensive care; used as a rule for who could go home, it would have sent some of the highest-risk patients away6. The remedy is a test. Observational studies once linked hormone therapy with 40 to 50 percent less heart disease7, yet a randomized trial found about 29 percent more8, because the women who chose the therapy were healthier to begin with.
In business the test is usually simpler: give the offer to a random group, hold it back from a similar group, and compare. Baselines, Metrics and Measurement in Module 03 covers how to design that comparison. The habit to keep here is a question: when a model’s pattern is about to become an intervention, ask for the test before the rollout.
Optimization does what you wrote down
Optimization is the step from knowing to doing. Given a forecast, it searches for the plan that best meets an objective without breaking the limits. Route planning, the classic example, belongs to AI in Operations and Supply Chain. The executive lesson is simpler and applies to every optimizer: best according to what?
The objective has to be written in terms a system can compute. “Maximize room revenue minus the cost of walked guests” is an instruction. “Delight our guests” is not. Hard limits must never break: safety, the law, signed contracts, physical capacity. Soft limits can bend at a price: a service target, staff preferences, a cap on overtime.
The catch is that an optimizer finds the best answer to the question you gave it, not the question you meant. In the illustrative hotel, if the objective prices every walked guest the same, the system will turn away a member of its top loyalty tier as readily as a first-time visitor. Nothing malfunctioned. The value the leadership team cared about was never written down. Writing the objective and the limits, and reviewing them when the plan surprises people, is where much of the value of optimization is won or lost.
Story: the forty-nine-foot forecast
Put yourself in Grand Forks, North Dakota, in the spring of 1997, and decide before you read on.
On 27 February the National Weather Service issued its outlook for the Red River at Grand Forks: a crest of 47.5 feet if no more precipitation fell, and 49 feet with normal precipitation. Forty-nine feet was close to the record set in 1979. The outlook gave two numbers and no probability9. Suppose your city’s flood defenses can be raised to 52 feet, three feet above the outlook. Raising them higher, preparing an evacuation and urging residents to buy flood insurance would all cost money, time and public nerves. Is three feet of margin enough?
Two facts were available to anyone who asked. The Weather Service later explained that an outlook assuming normal precipitation sits roughly at the middle of the possibilities: about a 50 percent chance of being equaled or exceeded9. And the political scientist Roger Pielke found that at East Grand Forks, across the side of the river, these outlooks had missed the actual crest by about 3.5 feet on average, and by more than 10 percent in 5 of 12 years10. A three-foot margin was smaller than the average miss.
What happened next is a matter of record. The forecasts were revised upward only in mid-April, as the river rose. It crested at about 54 feet on 21 and 22 April, overtopped the defenses and flooded most of Grand Forks and East Grand Forks. Damage around the two cities came to about 3.6 billion dollars of the 4 billion for the whole event, by the Weather Service’s estimate; no deaths were directly attributed to the flooding9. Pielke reports a post-flood survey in which 95 percent of Grand Forks respondents knew of flood insurance, yet 79.6 percent said the forecasts had led them to conclude flood insurance was unnecessary, and found that people had read the two numbers as a range, as a maximum or as exact10.
Now price the errors the way this chapter suggests. Suppose, conservatively, there was only a one-in-ten chance that the river would beat a three-foot margin. With 3.6 billion of damage at stake, that is an expected loss of 360 million dollars riding on the margin. The record suggests the chance was higher: with an even chance of exceeding 49 feet and past misses averaging 3.5 feet, a crest above 52 feet was plausibly nearer one in four. Planning for a bigger miss would not have prevented all of that damage, but even part of it is likely to exceed the cost of more sandbags, an earlier evacuation plan and a public warning to insure. The decision did not need a better forecast. It needed the range, the record of past misses and a priced cost of being short.
The Weather Service’s own assessment drew the same lesson: it needed to improve how it estimates and conveys the uncertainty of its outlooks, through probabilistic forecasts built from many simulated seasons9. That is exactly what AI ensembles such as GenCast now make cheap. The cost of producing a range has collapsed. The cost of ignoring one has not.
What this means for leaders
Start by refusing single numbers. Any forecast that informs a material decision should arrive with a range, a record of how often reality has fallen outside it, and its error shown on the expensive days, not only on average. This is a cheap request, and it changes the conversation from “is the number right?” to “what do we do across the range?”
Then make error pricing a business task. The cost of running short and the cost of over-preparing should be set by the people who own the consequences, written down and revisited. Where a model’s pattern is about to become an action, an offer, a price or a program, require a test with a held-back group before the full rollout. And where an optimizer is involved, treat the objective and the limits as a policy document that leadership signs, not a technical setting.
Check yourself
- A more accurate forecast automatically creates value.
- Lower average error always means a better forecast for the business.
- If running short costs four times as much as over-preparing, plan so you run short only about one time in five.
- A model that accurately predicts an outcome tells you what will happen if you intervene.
- In 1997, the 49-foot Red River outlook was roughly a midpoint, not a ceiling.
- An optimizer finds the best answer only within the objective and limits it is given.
Reflection: find the single number
What comes next
Throughout this module, AI has predicted, extracted and recommended, and a person has acted. The next chapter, AI Agents and Intelligent Workflows, asks what changes when the system can act on the forecast itself: coordinating steps, using tools and taking bounded actions inside a workflow.
References
- Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet and colleagues (Google DeepMind). Probabilistic weather forecasting with machine learning. Nature. 2024.
- European Centre for Medium-Range Weather Forecasts (ECMWF). ECMWF's ensemble AI forecasts become operational. ECMWF. 2025.
- Ajay Agrawal, Joshua Gans and Avi Goldfarb. Prediction Machines: The Simple Economics of Artificial Intelligence. Harvard Business Review Press. 2018.
- Rob J. Hyndman and George Athanasopoulos. Forecasting: Principles and Practice, 3rd edition. OTexts. 2021.
- Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan and Ziad Obermeyer. Prediction Policy Problems. American Economic Review 105(5), 491-495. 2015.
- Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm and Noemie Elhadad. Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission. Proceedings of KDD 2015, ACM. 2015.
- JoAnn E. Manson and colleagues for the WHI Investigators. Estrogen plus Progestin and the Risk of Coronary Heart Disease. New England Journal of Medicine 349, 523-534. 2003.
- Writing Group for the Women's Health Initiative Investigators (Rossouw et al.). Risks and Benefits of Estrogen Plus Progestin in Healthy Postmenopausal Women: Principal Results From the Women's Health Initiative Randomized Controlled Trial. JAMA 288(3), 321-333. 2002.
- NOAA National Weather Service. Service Assessment and Hydraulic Analysis: Red River of the North 1997 Floods. U.S. Department of Commerce, NOAA. 1998.
- Roger A. Pielke Jr. Who Decides? Forecasts and Responsibilities in the 1997 Red River Flood. Applied Behavioral Science Review 7(2), 83-101. 1999.
Further reading
- Rob J. Hyndman and George Athanasopoulos. Forecasting: Principles and Practice, 3rd edition. OTexts. 2021.
- Philip E. Tetlock and Dan Gardner. Superforecasting: The Art and Science of Prediction. Crown. 2015.
- Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect. Basic Books. 2018.
Sources last verified 2026-10-10.