Proprietary Data, AI Moats and Differentiation
Owning data that nobody else has is a head start, not a moat. Data defends a position only when more of it keeps making the product better for customers, in ways a rival cannot buy, and only when you have the right to learn from it. The honest test is the shape of the improvement curve: does it keep rising, or has it already flattened?
After this chapter you can
- Distinguish a data asset that gives a head start from a data moat that keeps a lead.
- Separate network effects, data scale effects and within-user learning in a data claim.
- Apply the honest test - is product quality still rising with more data? - using seven questions.
- Explain why recording outcomes inside the work keeps a learning curve rising.
- Recognize how use rights and customer trust limit what a data asset can defend.
In May 2019, two investors at Andreessen Horowitz, Martin Casado and Peter Lauten, published a curve that should sit in front of every board that hears the words “data moat”. It came from companies building automated assistants that answer customer requests. As the assistants were fed more and more real conversations, the share of requests they could handle rose quickly at first, then bent, and then flattened near about 40 percent. Past that point, more data bought almost nothing1.
Their conclusion ran against the standard pitch, in which more data makes a better product that brings more customers, forever. On the curve, each new piece of data was worth less than the one before, while the unusual cases that would have helped were getting harder and more expensive to collect. The first firm to gather a large pile had a head start. A rival with a smaller pile would reach the same plateau a little later and then stand level with it.
Many organizations hold data that nobody else has: transaction histories, maintenance logs, customer records, years of operating decisions. Many of them describe it as their AI moat. Sometimes they are right. This chapter gives you the test that tells you which case you are in.
Unique data is a head start, not yet a moat
The previous chapter, AI and Competitive Advantage, argued that an advantage needs a gap in cost or customer value and a guard that stops rivals closing it, and it described the AI flywheel through which use improves the offer23. Proprietary data is the guard most often claimed. It is also the one most often overstated. The economists Andrei Hagiu and Julian Wright, after studying how firms actually compete on customer data, concluded that most leaders overestimate how much protection it gives, because they treat data-driven learning as if it were as strong and lasting as a true network effect, and it usually is not4.
The distinction that matters is between an asset and a moat. An asset is valuable today. A moat keeps you ahead tomorrow, because using it makes it better in ways a competitor with money and a good model still cannot reproduce.
That makes the board-level question narrow and answerable. Not “how much data do we have?” but “does more of it keep making the product better for the people who use it, and could a rival get the same improvement another way?”
Scale effects are not network effects
Much of the confusion comes from one phrase, data network effect, which bundles three different mechanisms.
A classic network effect exists when a product becomes more valuable to each user because other users join. A telephone network or a marketplace works this way: a buyer gains directly from every extra seller. That value does not depend on anyone learning anything.
A data scale effect is weaker. More users generate more data, and the data may improve the product for everyone. But the users gain nothing from one another directly; they gain only through whatever the system learns, and learning runs into diminishing returns. Casado and Lauten argue that most of what gets called a data network effect is really this1.
Within-user learning is weaker still. Hagiu and Wright use the smart thermostat as their example: it learns a household’s preferred temperatures within days. That makes the device better for that household, but it does almost nothing for the next customer, and it stops improving once the routine is learned4. It can still raise switching costs, which is a real benefit, but it is not a data moat.
The honest test: plateau or rising curve
The cleanest way to judge a data claim is to draw the curve: product quality, as customers experience it, against the amount of data behind it. Two shapes are common; the chart below sketches them with illustrative numbers.
On the first curve, the lead is real but temporary. A competitor that starts later with less data reaches the same plateau and then competes on price, service or distribution, as if the data did not exist. On the second curve, the leader keeps pulling away, because each round of new data still teaches the system something it did not know.
Rising curves do exist, and they have conditions. Studies of Amazon’s forecasting and Yahoo’s search engine point the same way: more data kept paying only where it covered a long tail of rare cases, a world that keeps changing, and learning that spreads from one user to the others56. Where those conditions are missing, expect the flat shape.
A useful image is the difference between an ore seam and a spring. A seam is rich at first, and each further load holds less ore and more rock; a neighbor who starts digging later reaches the same rock. A spring keeps flowing because fresh rain keeps falling on ground that you hold. The question for every claimed data moat is which one it is.
Seven questions before you call it a moat
Hagiu and Wright turn the curve into seven questions, and they make a useful checklist for a board paper4. Each answer pushes a claimed moat toward the seam or toward the spring.
Three of the seven deserve extra attention, because they are the ones most often skipped. Shelf life cuts both ways: data that goes stale fast weakens the value of an old archive, but it rewards a firm that keeps collecting fresh data and acting on it quickly. Imitation is about the improvement, not the data: if a rival can see what your product now does well and copy it by other means, your data protected nothing. And reach is the difference between a data scale effect and within-user learning: if one customer’s data helps only that customer, the moat is a switching cost at best.
Two conditions sit underneath all seven: the data must be usable in practice, as Data Strategy for AI showed, and you must hold the right to learn from it, the subject of a later section.
What keeps a curve rising: outcomes captured in the work
When a data curve keeps rising, you will often find the same ingredient underneath: the system learns what happened after each decision. A forecast is followed by actual demand; a recommendation by whether it was accepted; a repair by whether it held. Without that outcome, more activity data only makes the pile bigger. With it, each decision becomes a lesson, and the lessons include exactly the unusual cases that a flat curve lacks.
Two practical points follow. First, outcomes are captured in the workflow, not in the data team: the crew, clerk or salesperson closing the task records the result, so the workflow has to make that easy. Second, outcome capture can be undone. A change of contractor, system or priority can cut the loop without anyone noticing, and the story later in this chapter shows how fast the curve then falls.
The right to learn: use rights and trust
A data asset you are not allowed to learn from is not a usable moat. It helps to think of rights as layers. Holding data, because you need it to run a service, is the bottom layer. Using it to run and improve your own operations sits above that. Training AI systems on it often needs a further basis that nobody asked for when the data was collected. Sharing or selling it needs the most.
Rights can also move against you, and machine data is where they move most.
Data a manufacturer assumed was its own may now reach a rival through the customer. Trust works the same way. Customers share data with organizations they expect to use it fairly. Use it in ways they did not expect, and the flow that feeds the curve can dry up, which is a cost no balance sheet shows until it is paid.
Story: Flint’s lead pipes, two seasons, one city
There are no competitors in a city’s pipe-replacement program, which makes it an unusually clean test of what data is worth. Nothing changed between the two seasons in this story except whether the learning loop ran.
In 2016, after the water crisis, Flint, Michigan, had to find and replace the lead and galvanized steel service lines feeding thousands of homes. Its own records were incomplete, and every excavation cost money: a few hundred for a quick hydro-excavation, 2,500 to 5,000 for traditional digging7. Two researchers then at the University of Michigan, Eric Schwartz and Jacob Abernethy, built a model that estimated, for each home, the probability of a hazardous line from its age, value, location and old city records. Crucially, it was designed to learn from the result of every excavation, and the team later proposed choosing some digs partly to learn from them89.
The loop worked. In the last months of 2016, about nine in ten of the city’s traditional excavations found a hazardous line8. By late 2017, the hit rate was still above 80 percent7.
In 2018 the city made a different choice. A new contractor took over, and residents told the mayor, in effect, that crews had checked a neighbor’s house and skipped theirs. Wanting to leave nobody behind, the city dug every house on selected blocks across all wards, set the model’s targeting aside, and stopped sending 2018 results to the researchers, so the model stopped learning in December 201778. The residents’ concern was legitimate; a program that seems to skip people loses public trust. But the cost of cutting the loop was large. By mid-August, the city’s hit rate was 19.7 percent, below the 31.4 percent that digging at random among eligible homes would have produced. In Ward 5, where the model expected the most remaining lead, 163 digs found it 95.7 percent of the time; in Ward 4, where it expected the least, 702 digs found it 2.4 percent of the time8.
In February 2019 the city agreed, in a settlement with residents’ groups, to return to the model and to prioritize up to 5,200 addresses it identified10. The contractor in 2018 had the same budget and even a map of the model’s predictions. What it did not have was the habit of recording what each dig found and feeding it back. That habit, not the archive and not the algorithm, was the asset no newcomer could buy. And the reason the loop was cut, public trust, is part of the same lesson: a learning system that the people it serves do not trust will not keep running.
What this means for leaders
The first discipline is vocabulary. When a business case says “data moat” or “data network effect”, ask which of the three mechanisms it means and ask to see the curve. If the honest answer is a data scale effect that will plateau, value it as a head start and plan to win on something else once rivals catch up. A correct label stops the organization overpaying for a guard it does not have.
The second is to invest in the spring rather than the seam. The money that wins is rarely the money spent on storing more history. It is the money spent on recording outcomes inside the work, shortening the time between a lesson and a better product, and reaching the unusual cases that keep the curve rising.
The third is to protect the loop as an asset. Name an owner, and treat any change of contractor, system or process that could cut outcome capture as a risk to the moat itself. And put the right to learn on the same list: what you may use data for, and what customers expect, are part of whether the moat exists.
Check yourself
- The organization with the most data in a market will keep the lead.
- Many claimed data network effects are really data scale effects.
- A smart thermostat that learns one household’s habits builds a strong data moat.
- A rising data curve needs conditions such as a long tail of rare cases and learning that spreads across users.
- If we already hold customer data, we may train AI on it.
- In Flint, the 2018 drop in hit rate came from losing the data archive.
Reflection: seam or spring?
What comes next
With advantage now tested, one chapter remains. Module 04 Synthesis — The Executive AI Strategy brings the module’s choices together into one strategy a leadership team can defend.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
EU Data Act · EU
Regulation (EU) 2023/2854
Users of connected products and related services have a right to access the data their use generates and to share it with third parties of their choice; data holders must make it available on fair, reasonable and non-discriminatory terms. Unfair data-sharing terms imposed on another business are not binding, and cloud providers must make switching possible. It limits how far machine-generated data from customers' use can serve as an exclusive moat. The EU Digital Omnibus proposals may amend parts of it.
- 2025-09-12 — Applies (most provisions)
- 2026-09-12 — Access-by-design duty for newly placed connected products
- 2027-01-12 — Cloud switching charges abolished
Last verified 2026-10-09 · official text
California privacy rules on automated decisions (CCPA regulations) · US - California
California Consumer Privacy Act; CPPA regulations on ADMT, risk assessments and cybersecurity audits (approved by OAL Sept 2025)
The most concrete US privacy rule on AI. Businesses that use automated decision-making technology to make a significant decision about a California resident (finance or lending, housing, education, employment or pay, healthcare) must give notice before use, offer an opt-out unless an exception applies, and answer access requests. Processing that poses significant privacy risk needs a documented risk assessment. There is no comprehensive federal privacy statute; about 20 states have their own laws, and California's is the reference point.
- 2026-01-01 — Updated CCPA regulations take effect; risk-assessment duty applies to new high-risk processing
- 2027-01-01 — ADMT duties for significant decisions: pre-use notice, opt-out (with exceptions) and access (some firm alerts cite enforcement from 1 Apr 2027)
- 2028-04-01 — Attestation of 2026-2027 risk assessments due to the CPPA; cybersecurity audits phase in 2028-2030 by revenue
Last verified 2026-10-06
References
- Martin Casado and Peter Lauten. The Empty Promise of Data Moats. Andreessen Horowitz. 2019.
- Richard Rumelt. Good Strategy/Bad Strategy: The Difference and Why It Matters. Crown Business. 2011.
- Marco Iansiti and Karim R. Lakhani. Competing in the Age of AI: Strategy and Leadership When Algorithms and Networks Run the World. Harvard Business Review Press. 2020.
- Andrei Hagiu and Julian Wright. When Data Creates Competitive Advantage...And When It Doesn't. Harvard Business Review 98 (1), January-February 2020. 2020.
- Patrick Bajari, Victor Chernozhukov, Ali Hortacsu and Junichi Suzuki. The Impact of Big Data on Firm Performance: An Empirical Investigation. AEA Papers and Proceedings 109, 33-37. 2019.
- Maximilian Schaefer and Geza Sapi. Complementarities in learning from data: Insights from general search. Information Economics and Policy 65, 101063. 2023.
- Alexis C. Madrigal. How a Feel-Good AI Story Went Wrong in Flint. The Atlantic (republished by Route Fifty). 2019.
- Eric M. Schwartz. Declaration of Eric M. Schwartz, Concerned Pastors for Social Action v. Khouri, No. 16-10277 (E.D. Mich.), ECF No. 203-4. United States District Court for the Eastern District of Michigan (filed by NRDC). 2018.
- Jacob Abernethy, Alex Chojnacki, Arya Farahi, Eric Schwartz and Jared Webb. ActiveRemediation: The Search for Lead Pipes in Flint, Michigan. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 5-14. 2018.
- MSW Management. Flint to Revert Back to Data-Driven Predictive Model for Lead Pipe Replacements. MSW Management. 2019.
Further reading
- Andrei Hagiu and Julian Wright. When Data Creates Competitive Advantage...And When It Doesn't. Harvard Business Review 98 (1), January-February 2020. 2020.
- Martin Casado and Peter Lauten. The Empty Promise of Data Moats. Andreessen Horowitz. 2019.
- Marco Iansiti and Karim R. Lakhani. Competing in the Age of AI: Strategy and Leadership When Algorithms and Networks Run the World. Harvard Business Review Press. 2020.
- Jacob Abernethy, Alex Chojnacki, Arya Farahi, Eric Schwartz and Jared Webb. ActiveRemediation: The Search for Lead Pipes in Flint, Michigan. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 5-14. 2018.
Sources last verified 2026-10-08.