AI in Product Development
AI has made product ideas, concepts and prototypes almost free. The scarce resource is now evidence about which of them customers will actually use and pay for. Leaders who use AI to learn faster, and who keep the choice of what to build with accountable people, get better products; leaders who use it only to build faster get the wrong product sooner.
After this chapter you can
- Distinguish product discovery (what to build, for whom, how we know) from delivery (building it correctly).
- Explain why cheap ideas and prototypes move the bottleneck to testing, given that at Microsoft most tested ideas failed.
- Join what customers say with what they do before treating feedback analysis as a roadmap.
- Use AI to diverge and people to converge, and rank product claims on the evidence ladder, including simulated respondents.
- Own the question and proxy that product AI optimizes, as the Netflix Prize post-mortem shows.
In the summer of 2023, four innovation professors at Wharton and Cornell ran a contest they had run many times before, with one new entrant. For years they had asked students in their product-design courses to invent a new physical product for college students that would retail for under 50 dollars. This time they gave the same brief to GPT-4. One person working with the model produced 200 ideas in about 15 minutes; a person working alone typically manages about five. Then the researchers pooled 200 student ideas with 200 from the machine and asked consumers to rate how likely they were to buy each one. The average student idea scored a purchase probability of 40 percent; the average GPT-4 idea scored 47 percent. Of the 40 best-rated ideas, 35 came from the machine1.
The authors titled their paper Ideas are Dimes a Dozen, and the title is the lesson. Read the last tile again. The ideas were judged by what people said they would buy in an online survey. Nobody manufactured the products, put them on a shelf or watched who paid. The study shows, convincingly, that generating options is no longer the hard part of product development. It says nothing about the hard part that remains: finding out which option customers will actually use, pay for and come back to.
The scarce resource has moved
Product development and software engineering are often blurred, and the difference matters more now. Software engineering asks how to build something correctly; AI in Software Engineering showed how AI speeds that up and where the gains stall. Product development asks a prior question: what should we build, for whom, and how do we know it creates value? Marty Cagan, whose book is the working manual for many product teams, calls the second activity discovery and the first delivery, and argues that discovery exists to retire four risks before money is spent: whether customers will value it, whether they can use it, whether it can be built, and whether it works for the business2.
Many product ideas fail that test, even at companies with strong product teams. When Microsoft ran controlled experiments on its own ideas, only about one third improved the metric they were designed to improve. In mature products such as Bing, the success rate was lower still, around 10 to 20 percent34. These were ideas that experienced people believed in enough to build and test.
Put the two findings side by side. AI can multiply the number of plausible ideas and prototypes many times over, and at Microsoft most tested ideas did not work. If the capacity to test them stays the same, the queue simply moves: the team drowns in promising options and still has to guess. The strategic opportunity is not to build faster. It is to learn faster which ideas deserve to be built.
Where AI enters the learning loop
Product development is a loop rather than a line. Signals come in, the team frames a problem, generates options, tests the most promising ones and decides what to do next. Each turn of the loop should leave the next decision a little less of a guess.
AI now helps at every point. It reads and clusters thousands of reviews, support conversations, survey comments and sales notes. It sharpens a vague request into a problem statement with questions attached. It drafts concepts, user flows and clickable prototypes in hours. It proposes experiment designs and works through the results. These are capabilities. Apart from the idea-quality studies later in this chapter, published evidence that they improve product results is still thin. What it does not do is close the loop. Deciding whether to invest, change course or stop is a judgment about uncertain evidence, strategy and trade-offs, and someone has to own the consequences. The rest of this chapter takes the loop one step at a time, because the way AI helps, and the way it misleads, is different at each.
Listen to what customers say, and watch what they do
Customer feedback analysis is one of the most accessible first uses of AI in product work. The volume is large, the task is easy to check, and a person reviews the output before anything is decided. It is also where an expensive mistake often begins, because feedback has two built-in biases. It comes from the customers who choose to write, not from the ones who quietly leave. And customers usually describe solutions rather than problems. A request for “more filters” or “an export button” is a customer’s guess at a fix for something that is getting in their way.
Clayton Christensen and his colleagues put the underlying point in terms of jobs to be done: customers hire a product to make progress in a particular circumstance, and data about who they are or what they ask for is not the same as understanding that job5. The practical consequence is simple. What customers say has to be joined with what they do: where journeys are abandoned, which workarounds people use, which features are opened once and never again, and who stops buying.
A good AI summary of feedback is therefore an input to a question, not an answer. A ranked list of the most-requested features is a list of hypotheses about problems. Before any of them is funded, a product leader should ask what the behavioral data says about the same customers, and who is missing from the sample.
Diverge, then converge
Generation is where AI is most impressive, and the evidence on idea quality is encouraging, though it measures ideas, not market results. The Wharton study that opened this chapter is one example. A larger field experiment at Procter & Gamble found that an individual professional working with AI produced solutions that expert judges rated at least as highly as those of a pair of colleagues working without it6.
Two cautions keep these results in proportion. First, they measure the quality of ideas, judged by consumers in a survey or by experts. Neither measures a product succeeding in the market. Second, AI’s ideas tend to cluster. A follow-up study on the same college-product task found that pools of ideas generated by GPT-4 were less diverse than pools generated by groups of people, although better prompting narrowed the gap7. A team that lets the model generate every option may explore a smaller part of the space than it believes.
The useful pattern is to diverge with AI and converge with people. Let the model widen the set of options quickly, push it deliberately for variety, and then have a small group with customer knowledge and commercial judgment choose the few worth testing. A polished prototype from that process proves that an idea can look right. It does not prove that customers understand it, want it, can use it, or that it solves their problem.
Simulated customers are a pre-test, not a verdict
A newer temptation is to skip the customer altogether. If a model can answer survey questions in the voice of a typical buyer, why not ask it what people will pay? James Brand, Ayelet Israeli and Donald Ngwe tested exactly this, comparing willingness-to-pay estimates from language-model “respondents” with estimates from studies of real people. The model’s answers were sometimes comparable, but often inaccurate, and in some cases pointed in the wrong direction altogether, valuing positively what people valued negatively, or the reverse. Fine-tuning the model on earlier human survey data improved alignment for features within the same product category and population. It did not help for new product categories or for differences between customer segments8.
That is a precise result, and it points to a precise use. Simulated respondents are useful for drafting and stress-testing a survey or a concept before real customers see it. They are weakest exactly where product teams most want help: genuinely new products and differences between segments. The authors conclude that they are a supplement to human studies, not a substitute.
The ladder is a useful habit for any product review. AI makes the bottom two rungs very cheap to reach, which is good, because it lets teams discard weak ideas early. It also makes it easy to present bottom-rung evidence with top-rung confidence. When someone shows you a beautiful prototype and enthusiastic survey numbers, the right question is which rung the claim stands on, and what it would take to climb one more.
Write the hypothesis before you look
Experiments are how a product team climbs the ladder: concept tests, usability tests, pricing tests and controlled online experiments. AI helps here too. It can suggest the primary metric, the segments to check, the guardrails and the confounders a team has missed, and it can analyze results in minutes.
That speed creates a specific danger. When analysis is nearly free, a team can slice the results a hundred ways and find a story in noise. The defense is old and simple: write down the hypothesis, the metric and the result that would change your mind before the test starts. Baselines, Metrics and Measurement covers the measurement side of this in depth, including the pre-agreed decision rule.
A useful test for any experiment report is to ask the team what result would have made them stop. If nobody can say, the experiment confirmed a decision that had already been made.
Story: the million-dollar answer to an old question
An instructive post-mortem comes from a media company that ran one of the best-known AI contests.
In October 2006 Netflix, then mainly a DVD-by-mail business, offered 1 million dollars to anyone who could improve the accuracy of its recommendation system by 10 percent. The company knew that “good recommendations” was hard to score, so it chose a proxy: how accurately an algorithm predicted the star rating a member would give a film9. The contest drew more than 51,000 contestants from 186 countries10.
The contest worked as designed. After a year, a team reached an 8.43 percent improvement, using a blend of 107 algorithms and more than 2,000 hours of work. Netflix took the two strongest components, rebuilt them to handle more than 5 billion ratings instead of the contest’s 100 million, and put them into production. In September 2009 a merged team called BellKor’s Pragmatic Chaos crossed the 10 percent line and won the prize910.
Then came the part that makes it a post-mortem. In 2012 two of Netflix’s personalization leaders explained that the company had evaluated the winning methods offline and that the extra accuracy did not seem to justify the engineering effort needed to bring them into production. The bigger reason was that the business had moved. Netflix launched streaming in 2007, one year into the contest. With streaming, members sampled several titles in one sitting, and the company could see what was watched to the end and what was abandoned. The data that mattered was no longer mainly what members said about a film afterwards. It was what they did9. The company also kept offline accuracy in its place: it tracked how well such measures predicted gains in live A/B tests and, because the link was imperfect, used them only to choose which ideas to test11.
In 2017 Netflix replaced its five-star ratings with a thumbs up or down. In a 2016 test with hundreds of thousands of members, thumbs drew 200 percent more ratings than stars. Todd Yellin, its vice president of product, explained that members gave documentaries high ratings but watched lighter films far more often: “What you do versus what you say you like are different things”12. By 2012 the company already reported that 75 percent of what people watched came from some form of recommendation9.
Netflix has said that the prize paid for itself in algorithms, attention and talent, so this is not a story of failure11. It is a story about the question. The model builders did exactly what they were asked. Choosing the question, a proxy fixed in 2006 for a DVD business, was product work, and the product moved under it. Three lessons carry over to any team using AI to build products. AI optimizes the question it is given, so deciding the question is the product leader’s job. A proxy needs a review date, because products and customers change. And what customers do is stronger evidence than what they say, whether they say it in a rating, a survey or a feature request.
What this means for leaders
Four decisions follow, and each belongs to leadership rather than to the product team alone.
Fund learning capacity, not only building capacity. If AI multiplies the ideas and prototypes a team can produce, the constraint moves to research time, access to behavioral data and the ability to run experiments. A budget that buys generation tools but not testing capacity buys a longer queue of untested ideas.
Ask which rung every claim stands on. Make the evidence ladder part of funding decisions. A model’s prediction or a survey can justify a cheap test. Only behavior in a test can justify a significant build.
Diverge with AI, converge with people. Use AI to widen options and to break down silos between technical and commercial staff. Keep the choice of the few bets with a small, varied group of people who know the customer and own the result. Where AI-generated options should become AI-native products, the question moves to AI-Native Products and Business Models in Module 10; how to sequence the bets is the work of Module 09, and the risks of personalization and automated decisions are covered in Module 06.
Own the question and its proxy. Every AI-assisted product effort optimizes a measure someone chose. Name the owner of that choice, and set a date to ask whether it still matches what customers do.
Check yourself
- AI-generated prototypes mean a team can skip product discovery.
- In a Wharton study, GPT-4 produced most of the best-rated new product ideas.
- Language-model respondents reliably estimate what customers will pay for a new product.
- A highly requested feature can be a solution in disguise.
- Netflix deployed the full algorithm that won its 1-million-dollar prize.
Reflection: find the old question
What comes next
This chapter and the ones before it looked at AI in the functions that serve customers and build products. The next one turns to protecting the enterprise, where AI strengthens the defenders and also hands attackers new tools. The next chapter is AI in Cybersecurity.
References
- Karan Girotra, Lennart Meincke, Christian Terwiesch and Karl T. Ulrich. Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation. Wharton Mack Institute / SSRN working paper 4526071. 2023.
- Marty Cagan. Inspired: How to Create Tech Products Customers Love, 2nd edition. Wiley. 2017.
- Ron Kohavi, Thomas Crook and Roger Longbotham. Online Experimentation at Microsoft. Microsoft (Third Workshop on Data Mining Case Studies and Practice Prize). 2009.
- Ron Kohavi, Diane Tang and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. 2020.
- Clayton M. Christensen, Taddy Hall, Karen Dillon and David S. Duncan. Know Your Customers' Jobs to Be Done. Harvard Business Review. 2016.
- Fabrizio Dell'Acqua, Charles Ayoubi, Hila Lifshitz, Raffaella Sadun, Ethan Mollick, Lilach Mollick, Yi Han, Jeff Goldman, Hari Nair, Stewart Taub and Karim R. Lakhani. The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. National Bureau of Economic Research (Working Paper 33641); later published in Organization Science (doi 10.1287/orsc.2025.20702). 2025.
- Lennart Meincke, Ethan R. Mollick and Christian Terwiesch. Prompting Diverse Ideas: Increasing AI Idea Variance. arXiv 2402.01727. 2024.
- James Brand, Ayelet Israeli and Donald Ngwe. Using LLMs for Market Research. Harvard Business School Working Paper 23-062 (revised April 30, 2026; first circulated 2023 as "Using GPT for Market Research"). 2026.
- Xavier Amatriain and Justin Basilico. Netflix Recommendations: Beyond the 5 stars (Part 1). Netflix Technology Blog. 2012.
- KDnuggets. BellKor Pragmatic Chaos wins Netflix prize by a few minutes. KDnuggets News 09:18. 2009.
- Xavier Amatriain and Justin Basilico. Netflix Recommendations: Beyond the 5 stars (Part 2). Netflix Technology Blog. 2012.
- Janko Roettgers. Netflix Replacing Star Ratings With Thumbs Ups and Thumbs Downs. Variety. 2017.
Further reading
- Karan Girotra, Lennart Meincke, Christian Terwiesch and Karl T. Ulrich. Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation. Wharton Mack Institute / SSRN working paper 4526071. 2023.
- Fabrizio Dell'Acqua, Charles Ayoubi, Hila Lifshitz, Raffaella Sadun, Ethan Mollick, Lilach Mollick, Yi Han, Jeff Goldman, Hari Nair, Stewart Taub and Karim R. Lakhani. The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. National Bureau of Economic Research (Working Paper 33641); later published in Organization Science (doi 10.1287/orsc.2025.20702). 2025.
- Teresa Torres. Continuous Discovery Habits: Discover Products that Create Customer Value and Business Value. Product Talk LLC. 2021.
Sources last verified 2026-10-08.