AI in Operations and Supply Chain
Operational AI ends in a physical move: a reorder, a repair, a route, a part rejected on the line. A signal creates value only when someone is allowed to act on it in time, inside a chain that does not amplify it. Leaders own the decision clock, the override design and the limits of authority, and they score the result in downtime, stock and cash rather than in model accuracy.
After this chapter you can
- Explain why operational AI creates value only where a physical action changes, and who must be allowed to take it.
- Recognize operations as coupled decisions, and how forecasting from orders and batching amplify swings (the bullwhip).
- Match AI systems to the clock of the decision they serve, and shorten that clock where the money is.
- Score maintenance and logistics AI in downtime, miles, stock and cash, counting the work false alarms create.
- Design overrides and authority tiers by physical consequence and reversibility.
When executives at Procter & Gamble looked closely at the order patterns for Pampers, one of their best-selling products, they found something odd. Babies use diapers at a fairly steady rate, and sales in retail stores did fluctuate, but not by much. The orders that distributors sent to P&G were another matter: the executives were surprised by how much they swung. And when they looked at P&G’s own orders for materials from suppliers such as 3M, the swings were greater still. Hewlett-Packard found the same shape in one of its printers: modest swings in a reseller’s sales, bigger swings in the reseller’s orders, and the biggest in the orders from the printer division to the company’s own chip division1.
Nobody in that chain lacked a forecast. Each tier forecast diligently, from the orders that arrived at its door. Hau Lee, V. Padmanabhan and Seungjin Whang named the result the bullwhip effect, and their diagnosis still reads as a warning for anyone buying AI for operations: the problem was not the quality of any one prediction. It was how decisions were connected, how often they were made, and what information each decision maker was allowed to see.
The value is in the physical move
Most of the AI this course has discussed so far ends in words: a draft, an answer, a summary. Operational AI ends in the physical world. A truck leaves or waits. A part is ordered or not. A machine is stopped tonight or runs until it breaks. Stock sits in one warehouse while the shortage is in another. That changes how value is created, and it changes what can go wrong.
The general distinction between predicting, forecasting and optimizing has its own chapter, AI in Prediction, Forecasting and Optimization, and AI in Finance already made the point that a forecast is judged by the decision it changes. Operations adds three hard conditions. The signal has to reach someone who is allowed to act. It has to arrive in time for the action to matter. And the action has to happen inside a chain of other decisions that can either absorb it or amplify it.
Organizations tend to spend on the first two steps of this loop, because that is where the technology is. The loop often breaks at the third and fourth, because that is where the organization is. The executive test for any operational AI proposal is therefore blunt: which physical action will happen differently, who will take it, and how fast?
Operations is a set of coupled decisions
It helps to see operations not as a list of AI tools but as a family of decisions. Planning decides what demand to expect and how supply should follow. Purchasing decides what to buy, from whom and when. Production decides what to make and in what sequence. Maintenance decides when to intervene on an asset. Quality decides what passes inspection, and camera-based inspection can in principle check every unit on a line rather than a sample, a capability whose payoff depends on what the line does with each rejection. Logistics decides routes, loads and delivery slots.
The word that matters is coupled. Each of these decisions feeds the next, so a local improvement can make the whole worse. Lee and his colleagues traced the bullwhip to four causes, and every one of them is a decision rule, not a data problem1. Two of them should worry any buyer of AI forecasting. If each tier feeds its model the orders it receives rather than what end customers actually consume, a faster, more responsive model will pass on the swings faster. And if orders are still batched once a month, a daily forecast changes nothing until the batch is placed.
The remedies Lee and his colleagues proposed were organizational: share sell-through data with suppliers, order in smaller and more frequent batches, stabilize prices, and allocate scarce supply by past sales rather than by what customers claim to need. AI makes several of those remedies cheaper to run. It does not make them happen.
Every decision has a clock
Operational decisions run on very different clocks. A vibration reading on a pump may need an answer in minutes. A delivery route is re-sequenced within the day. Replenishment runs daily or weekly. The production plan is often fixed a week or a month ahead. Supplier contracts and capacity move over quarters. An AI system is worth what the fastest decision it can actually change is worth.
Suppose, as an illustration, that a logistics team receives an early and correct AI warning that a storm will close a key port on Friday. The alert lands in a shared inbox on Tuesday. On Wednesday everyone has seen it and no one owns it. On Thursday someone asks who may reroute the containers, and the answer involves three approvals and a carrier contract nobody has read. On Friday the port closes and the containers wait. The prediction worked. The decision clock did not.
Two practical consequences follow. First, real-time systems are worth their cost only where the decision is real-time; a weekly replenishment decision does not need streaming data. Second, the cheapest improvement is often to shorten the decision clock itself: move from monthly to weekly ordering, give a planner authority to transfer stock between warehouses, pre-approve the alternative carrier. Those are management decisions, and they are often worth more than the next model upgrade. Route optimization shows the scale of small daily moves: UPS expected its ORION planner, which chooses the order of a driver’s stops and lets drivers depart from the planned sequence where they know better, to cut about 100 million miles a year at full deployment, an expectation rather than an audited result2.
Maintenance: score the downtime, not the model
Predictive maintenance replaces two old habits: servicing on a fixed calendar, and running equipment until it fails followed by an emergency repair. Condition-based maintenance uses sensor data, such as vibration, temperature and oil analysis, to estimate when an asset is likely to fail and to choose when to intervene3. The promise is real. A US Department of Energy guide estimated that a working predictive program saves 8 to 12 percent over preventive maintenance alone, and cited industry surveys reporting 35 to 45 percent less downtime and 70 to 75 percent fewer breakdowns4. Those are survey averages from 2010, before modern machine learning; treat them as the size of the prize, not as a forecast for your plant.
The trap is to measure the model instead of the maintenance. Every alert sends a technician, and many stop a machine. A model that catches more failures by raising more alarms can cost more than it saves.
The figure is an illustration, not plant data, and so is its arithmetic. Suppose each failure caught early avoids 40,000 of unplanned downtime, and each false alarm costs 7,500 in technician time and a planned stop. A sensitive setting flags 100 machines in a quarter and catches 20 real failures: 800,000 saved, minus 80 false alarms costing 600,000, leaves 200,000. A stricter setting flags 40 machines and catches 16: 640,000 saved, minus 24 false alarms costing 180,000, leaves 460,000. The second model catches fewer failures and creates more than twice the value. The general logic of unequal error costs belongs to AI in Prediction, Forecasting and Optimization; the operational lesson is to report downtime, maintenance cost, technician hours and lost production, and to count the work every false alarm creates. Teams that are sent to too many healthy machines also stop trusting the alerts, which is a cost no dashboard shows.
Authority grows with physical consequence
Because operational AI moves physical things, who may act matters more here than in most functions. The useful test is reversibility: can the action be undone before it costs anything? A purchase order can usually be cancelled the same day. A container that has sailed, a batch that has been released or a line that has been stopped cannot.
At the bottom, AI acts within limits on routine, reversible actions, and every action is logged. In the middle, it recommends and a named person approves. At the top, people decide, protected by deterministic safeguards such as interlocks and certified control systems that do not depend on a model being right. The tiers are not fixed: an action can move down a tier once its record is good, and should move up after an incident. The general design of human oversight is covered in Module 06; here the point is that in operations the tier is set by physics, not by enthusiasm.
Story: the stores that ignored the buyers
An automobile spare-parts retailer had given its merchants, the buyers who decide what each store stocks, a data-driven tool. It drew on local car registrations, weather and each product’s sales history to recommend stocking decisions. Like many firms, the retailer let the merchants override the tool, on the reasonable grounds that they might know things the tool did not. Saravanan Kesavan and Tarun Kushwaha ran a field experiment to find out what that discretion was worth. For twelve months, across more than 30,000 products, the merchants’ overrides were carried out in some stores and ignored in a random selection of others8.
Overall, the stores that honored the overrides were 5.77 percent less profitable than the stores that followed the tool to the letter8. That sounds like a verdict against human judgment, and it is not. The effect reversed with the product’s life cycle. For growth-stage products, the overrides drove more than 23 percent greater profitability; for mature and declining products, the merchants did worse than the tool98. The explanation is intuitive once stated: for a new part, the tool has little history to learn from, and an experienced buyer who knows which car models are arriving in the area does have an edge. For a part that has sold steadily for years, the tool has seen everything the buyer has, and the buyer’s adjustments mostly add noise. The two figures measure different things: the 5.77 percent compares all products in the two groups of stores, while the 23 percent covers only the growth-stage subset.
A later experiment at a spare-parts retailer, by the same two researchers and Dayton Steele, tested a different design. Instead of overriding the tool’s output, staff could adjust the inputs to the forecast that fed an inventory algorithm. Those adjustments raised profitability by 4.92 percent on average compared with automation alone, and the authors note that forecast performance need not translate into profit performance10.
Read the two results together and the lesson for an operations leader is not “trust the algorithm” or “trust your people”. It is that the override is a design decision. Where may people intervene: on the inputs, where they add what the model cannot see, or on the final answer, where they tend to second-guess what it already knows? For which products, suppliers or situations? And is each override recorded and scored in profit, so the organization learns whose judgment adds value and where? An override button is not an override policy.
What this means for leaders
Operations is where AI’s predictions turn into trucks, parts and machine hours, and where a good signal is most easily wasted. The leadership work sits in four decisions. Map the decisions that the AI will feed, and check that improving one will not amplify swings in the next. Set the decision clock deliberately, and shorten it where the money is. Design the override: where people may intervene, for which cases, and how each intervention is scored. And set authority by physical consequence, writing down what the system may do alone, what it may only recommend, and what stays with people and hard safeguards. Through all of it, insist on an operational scoreboard: service level, stockouts, expired or obsolete stock, downtime, maintenance cost, miles, fuel and working capital.
Check yourself
- The bullwhip effect is mainly caused by poor forecasting models.
- A maintenance model that catches more failures always creates more value.
- Real-time AI is worth its cost only where the decision itself is made in real time.
- In a randomized field experiment, buyers’ overrides of a stocking tool lowered profit overall but raised it for new products.
- The EU Machinery Regulation, applying from January 2027, covers machinery safety components that use AI.
Reflection: follow one signal
What comes next
Operations shows AI at the physical core of the enterprise, where a prediction becomes a truck, a part or a stopped machine. The next chapter turns to a function where AI helps build the systems that run the business itself, and where the same question follows us: what actually changes when the AI gets better? That chapter is AI in Software Engineering.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU Machinery Regulation · EU
Regulation (EU) 2023/1230
Safety rules for machinery, covering self-evolving behaviour and safety components that use AI. Relevant to robots and autonomous equipment.
- 2027-01-20 — Applies
Last verified 2026-10-06 · official text
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Hau L. Lee, V. Padmanabhan and Seungjin Whang. The Bullwhip Effect in Supply Chains. Sloan Management Review 38(3), 93-102. 1997.
- Transport Topics. UPS Routing Program ORION Helps Drivers Trim Miles, Reduce Costs. Transport Topics. 2016.
- Andrew K. S. Jardine, Daming Lin and Dragan Banjevic. A review on machinery diagnostics and prognostics implementing condition-based maintenance. Mechanical Systems and Signal Processing 20(7), 1483-1510. 2006.
- G. P. Sullivan, R. Pugh, A. P. Melendez and W. D. Hunt. Operations & Maintenance Best Practices: A Guide to Achieving Operational Efficiency, Release 3.0. Pacific Northwest National Laboratory for the US Department of Energy, Federal Energy Management Program. 2010.
- European Parliament and Council of the European Union. Regulation (EU) 2023/1230 on machinery. Official Journal of the European Union, L 165 (corrigendum OJ L 169, 4 July 2023). 2023.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
- European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
- Benjamin Kessler. In an Algorithmic Workplace, How Can Humans Excel?. George Mason University News. 2021.
- Saravanan Kesavan and Tarun Kushwaha. Field Experiment on the Profit Implications of Merchants' Discretionary Power to Override Data-Driven Decision-Making Tools. Management Science 66(11), 5182-5190. 2020.
- Saravanan Kesavan, Tarun Kushwaha and Dayton Steele. Profit Implications of Judgmental Adjustments to Forecast Inputs: Evidence from a Large-Scale Field Experiment. Management Science 72(1), 119-127. 2026.
Further reading
- Hau L. Lee, V. Padmanabhan and Seungjin Whang. The Bullwhip Effect in Supply Chains. Sloan Management Review 38(3), 93-102. 1997.
- Saravanan Kesavan and Tarun Kushwaha. Field Experiment on the Profit Implications of Merchants' Discretionary Power to Override Data-Driven Decision-Making Tools. Management Science 66(11), 5182-5190. 2020.
- G. P. Sullivan, R. Pugh, A. P. Melendez and W. D. Hunt. Operations & Maintenance Best Practices: A Guide to Achieving Operational Efficiency, Release 3.0. Pacific Northwest National Laboratory for the US Department of Energy, Federal Energy Management Program. 2010.
Sources last verified 2026-10-08.