AI Use-Case Discovery and Design
When AI projects fail, a weak model is seldom the main reason; more often, nobody pinned down the problem. Good use cases are found by mapping real work task by task, giving each task to a rule, to AI or to a person, and passing the design through four screens before money moves.
After this chapter you can
- Turn a solution-shaped AI proposal into a problem statement with a user, a measure and four named roles.
- Map a workflow as it is and locate where time and errors are lost.
- Test suitability task by task, choosing a rule, AI or a person for each.
- Design the split of work and a one-page brief with a baseline and a stop rule.
- Apply four pass-or-fail screens before funding the smallest test.
In 2024, researchers at RAND interviewed 65 experienced data scientists and engineers about why artificial intelligence projects fail. The question matters because, by the estimates RAND cites, more than 80 percent of AI projects fail, about twice the rate of information technology projects that do not involve AI. You might expect the main causes to be technical: weak models, poor infrastructure, problems too hard for the technology. Those were on the list. The cause interviewees described as the most common was different. The people involved misunderstood, or miscommunicated, what problem the AI was supposed to solve1.
That is the contradiction at the heart of this chapter. The part of an AI project that looks hardest, the technology, is less often where it breaks. The part that looks easiest, saying clearly what problem you are solving and for whom, is where it most often goes wrong. Use-case discovery and design is the discipline that fixes that.
Ask whether AI should, not whether it can
Almost any AI demonstration can prove that AI can do something. A fluent draft, a tidy summary, a plausible forecast. That proof is cheap, and it answers the wrong question. The executive question is whether AI should do this work, in this workflow, for this outcome, more cheaply and safely than the alternatives.
Discovery sits between two pieces of work covered elsewhere in the course. Finding Strategic AI Opportunities identified where the business creates value and where it is stuck. What Does AI Value Actually Mean? showed how to state the value you expect as a testable hypothesis. Discovery and design is the step in between. It takes an opportunity, such as “our quoting is too slow”, and turns it into a designed use case: a named user, a mapped workflow, a clear split of work between rules, AI and people, a measure with a baseline and a decision about whether to proceed.
The method has five moves, and the rest of this chapter takes them in order. Write the problem before the solution. Map the work as it is. Test each task: rule, AI or person. Design who does what. Then put the design on one page and pass it through four screens.
Write the problem before the solution
Many AI proposals arrive as solutions. “We need an AI assistant.” “Let’s build a chatbot.” “Can we use a language model for our contracts?” Each names a technology and leaves the problem implied. The first discipline is to send the proposal back with one question: to solve what?
The difference shows up in a single sentence, here an illustrative one. “Use AI to summarize customer calls” is a mechanism. “Cut the time account managers spend preparing customer follow-ups from about an hour to about fifteen minutes, without lowering follow-up quality” is a problem. The second names a user, a piece of work, a measure and a guardrail. It also leaves open whether a summary is even the right mechanism. Perhaps the time goes into finding the right contract, not into reading the call notes.
A well-written problem statement also names four roles, which are often four different people. Someone performs the work, someone receives the result, someone owns the workflow and someone owns the business outcome. In a quoting team, the analyst performs the work, the customer receives the quote, the desk manager owns the workflow and the commercial director owns the win rate. If you cannot name all four, you do not yet know who will judge whether the use case worked, and nobody will be able to stop it if it does not.
Map the work as it is
Before anyone designs an AI solution, someone has to map how the work happens today: not the process in the manual, but the work as people actually do it. That means watching it, timing it and asking where it waits. The map should break the workflow into tasks, and mark where the time and the errors go.
The map changes the conversation in two ways. First, it shows where the friction really is. In many knowledge workflows, the time is lost in waiting, searching, re-keying and handoffs, not in the skilled step that the proposal wanted to accelerate. Second, it shows the difference between speeding up a task and improving the workflow. A team that uses AI to draft one reply may save, say, two minutes. A team that redesigns the sequence may remove a whole loop of manual work.
That second point is one of the oldest lessons in process improvement, and it predates AI by decades.
Michael Hammer’s summary, in the title of his 1990 article, was “don’t automate, obliterate”. The modern evidence points the same way: as The Enterprise AI Use-Case Landscape showed, McKinsey’s 2025 survey found workflow redesign to be the practice most linked to an earnings impact from generative AI, and few organizations had done it3.
Test each task: rule, AI or person
A mapped workflow is a list of tasks, and the most expensive mistake in discovery is to treat the whole list as an AI problem. Each task deserves its own answer. Some should be done by a fixed rule, some by AI, some by a person and some should be removed altogether.
Start with the cheapest question: is the task a fixed rule? “If the margin is above the floor and the customer is in good standing, approve the quote” is deterministic. Ordinary software does it predictably and cheaply, every time, with nothing to evaluate. Google’s guidance to its own engineers makes the point bluntly. Rule number one is “don’t be afraid to launch a product without machine learning”, and its author notes that if machine learning would give a 100 percent boost, a simple heuristic often gets you 50 percent of the way4. The US National Institute of Standards and Technology builds the same check into its AI risk framework: mapping an AI system’s context should include whether non-AI or non-technology alternatives would serve better5.
Now change the task. “Read a customer’s free-text request, with its attachments, and turn it into a structured order” is not a rule. The input varies, the language is ambiguous and the answer depends on context. That is the kind of work where AI can add value, though whether it does in your workflow is a capability to test, not a result to assume. Four more questions decide whether it should. Is the input varied and information-heavy? Language, documents and judgment favor AI; fixed formats favor rules. Can the output be checked cheaply? If a person can verify a draft in seconds, AI is a good fit, but if checking takes as long as doing the work, the gain disappears. Is the data or context available? AI cannot read what it is not given. And what does an error cost? A wrong draft that a person catches costs seconds. A wrong payment, price or rejection that nobody catches costs much more.
The answers will differ for tasks that look alike. As AI Is Changing Everything showed, AI’s capability is jagged, strong on some tasks and harmful on others that seem similar6. That is why the test runs task by task, not use case by use case. Rule number three in the same Google guidance completes the picture: once a hand-built rule becomes too complex to maintain, that is the moment to choose machine learning4. Rules first is a starting point, not a dogma.
Design who does what
Once some tasks are AI candidates, design the future workflow and write down the split of work explicitly. For each task, what does AI do, what does a person do, and what is AI never allowed to touch? This is part of the design, not a control bolted on at the end.
Three design choices matter more than executives usually expect. The first is the difference between reading and writing. An AI that reads records and drafts a recommendation is one level of risk. An AI that can change a record, send a message or commit money is another, and the design should say which it is.
The second is where the human sits. When the consequences are meaningful, AI recommends and a person decides. When actions are small, bounded and reversible, AI may act while a person monitors. Human Oversight and AI Incidents treats these oversight models in depth, and AI Agents and Intelligent Workflows covers AI that acts across several steps. At the design stage, the requirement is simply to choose and write it down.
The third is what the design leaves the person to do. Lisanne Bainbridge’s classic paper on the ironies of automation warned that when designers automate what they can, people are left with the hardest tasks and the least practice at them7. A good design gives people the judgment calls and the evidence they need to make them. The pattern works: in a study of more than 5,000 customer-support agents, an AI tool that suggested responses, which agents could use, edit or ignore, raised the number of issues resolved per hour by 15 percent on average8.
Put the design on one page
A use case is ready for a funding decision when it fits on one page. The page is the artifact that discovery produces, and the test of whether discovery happened at all.
Two boxes deserve emphasis. The measure needs a baseline, taken before anything changes, or no later result will mean anything; Baselines, Metrics and Measurement explains how to set one. And the smallest test needs a stop rule written in advance: the result that would make you halt or redesign. The discipline of controlled experiments applies here. Decide the metric and the comparison before you look at the data, not after9. A team that writes its stop rule first is testing an idea. A team that writes it afterwards is defending one.
Four screens before money moves
The finished brief passes through four screens. Each is a pass or fail, not a score.
Value asks whether there is a prize that the business owner, not the AI team, would recognize. Feasibility asks whether the data, skills and systems exist now, not whether a vendor says they could. Risk asks whether the worst plausible error is acceptable given the human role designed in; Module 06 covers how to assess it. Strategic fit asks whether the use case serves something the enterprise already cares about. A brief that fails one screen is not a failure. It is a redesign, a later candidate or a clear no, and a clear no is cheap.
The screens matter because they catch, before the pilot, the reasons pilots die after it. Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value10. That is a prediction, not a measurement, but every reason on its list is one of the four screens, checked too late. Finding Strategic AI Opportunities compares opportunities against each other, and Prioritizing the AI Portfolio ranks the survivors. The screens here answer a narrower question: is this one design ready to test? They come before any pilot; the four gates a pilot must pass on its way to scale come later, in From AI Pilot to Production to Scale.
Story: the quote desk
This story is an illustrative composite, drawn from common patterns in freight forwarding. The company and the numbers are invented.
Suppose you run a mid-sized freight forwarder. Your sales team is losing spot-freight bids because quotes take too long: customers want a price the same day, and your desk averages about a day and a half. The commercial director brings you a proposal: an AI pricing engine that learns from past quotes and generates prices automatically. The demo is impressive.
Before you decide, the head of operations asks for two weeks to map the quote workflow. Her map shows four things. Requests arrive as free-text emails with attachments, and analysts re-key dimensions, weights and delivery terms by hand, with occasional costly errors. Every quote above a modest size then waits in a manager’s approval queue, although the manager approves almost all of them by checking one thing: whether the margin is above the floor. Some quotes wait for carrier rates. And the pricing itself, the step the AI engine targets, takes analysts a small share of the time, and their judgment on market rates is the desk’s real asset.
You now have three options. Fund the AI pricing engine as proposed. Skip AI entirely and fix the approval queue. Or split the work: a rule for approvals, AI for reading the requests, people for pricing. Which would you choose?
The company chose the split. The approval check became a rule: quotes above the margin floor go straight out, and only exceptions reach the manager. AI reads each incoming request and drafts a structured order, which the analyst confirms in seconds before pricing it. Pricing stays with the analysts. The team measured turnaround and win rate against the baseline it took during mapping, and wrote a stop rule in advance.
Suppose the results look like the figure: turnaround falls from 36 hours to 16. The largest gain came from a rule, which no AI proposal had mentioned. AI earned its place on the one task that was varied, language-heavy and cheap to check. The pricing engine was not rejected; it was parked until a brief could say which task it would change and how the analysts would check it. Carrier waiting time remained, and that was a contracting problem, not an AI one.
What this means for leaders
Discovery and design are leadership work, not technical preparation. The questions that decide whether an AI investment pays off are about problems, workflows and people, and they are cheapest to answer before anything is built. Send solution-shaped proposals back for a problem statement. Pay for mapping before you pay for models; two weeks of watching the work is often the best-value spending in an AI program. Expect a good design to give some tasks to rules and some to people, and treat that as a sign of rigor, not a lack of ambition. Insist on the one-page brief with a baseline and a stop rule. And hold every brief to the four screens, so that the reasons pilots fail are found while they still cost little.
Check yourself
- Most AI projects fail because the models are not good enough.
- If a workflow has a lot of friction, it is a good AI opportunity.
- Suitability should be tested task by task, not for the use case as a whole.
- A design that gives the final decision to a person is not really an AI use case.
- The stop rule for a pilot should be written before the pilot starts.
- A use case that fails one of the four screens should be dropped for good.
Reflection: map one workflow
What comes next
The method is the same everywhere: problem, map, task split, brief, screens. What changes from function to function is the work, the data and the cost of an error. The next chapter, AI in Finance, applies the method to one of the most data-rich and control-heavy functions in any enterprise, where the line between drafting an analysis and authorizing a payment matters a great deal.
References
- James Ryseff, Brandon F. De Bruhl and Sydne J. Newberry. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI. RAND Corporation, RR-A2680-1. 2024.
- Michael Hammer. Reengineering Work: Don't Automate, Obliterate. Harvard Business Review 68(4), July-August 1990. 1990.
- McKinsey & Company (QuantumBlack). The state of AI: How organizations are rewiring to capture value. McKinsey & Company. 2025.
- Martin Zinkevich. Rules of Machine Learning: Best Practices for ML Engineering. Google for Developers. 2017.
- National Institute of Standards and Technology. AI RMF Playbook: Map. NIST Trustworthy and Responsible AI Resource Center. 2023.
- Fabrizio Dell'Acqua et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, HBS Working Paper 24-013. Harvard Business School. 2023.
- Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond. Generative AI at Work. The Quarterly Journal of Economics 140(2). 2025.
- Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
- Gartner. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025. Gartner Newsroom. 2024.
Further reading
- Michael Hammer. Reengineering Work: Don't Automate, Obliterate. Harvard Business Review 68(4), July-August 1990. 1990.
- James Ryseff, Brandon F. De Bruhl and Sydne J. Newberry. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI. RAND Corporation, RR-A2680-1. 2024.
- Martin Zinkevich. Rules of Machine Learning: Best Practices for ML Engineering. Google for Developers. 2017.
Sources last verified 2026-10-08.