Data Strategy for AI
Almost every organization has data. Far fewer have data that a particular AI system can reach, trust, understand and lawfully use. A data strategy for AI starts from the outcomes the AI strategy has already chosen, finds the data gap that blocks each one, and fixes that gap first instead of trying to make everything ready at once.
After this chapter you can
- Explain why usable enterprise data, not model access, is usually the binding constraint on AI.
- Connect a data strategy for AI to the existing enterprise data strategy instead of duplicating it.
- Distinguish data that exists from data that is usable, and explain why quality is fitness for a specific use.
- Rate the data behind a chosen AI opportunity as green, amber or red and choose the gap to fund first.
- Assign business ownership, evaluation data and permitted purpose to the data sets a chosen opportunity depends on.
Before reading on, make a prediction. Over two years, the data-quality researchers Tadhg Nagle, Thomas Redman and David Sammon asked 75 executives to run a simple test. Each one gathered the last 100 records their own team had created, chose 10 to 15 attributes that mattered, and marked every obvious error. A record counted only if it had no errors at all. The researchers set the bar for “acceptable” at 97 clean records out of 100. What share of the scores do you think cleared it? Write your number down.
The answer was 3 percent. That figure counts teams’ scores. Counted record by record instead, 47 percent of newly created records had at least one critical error on average1.
The sample was small and self-selected, so treat the exact figures with care. The direction is hard to argue with. Every one of those teams had data. It lived in their systems and fed their reports. Very little of it would have passed if an AI system had been asked to act on it at scale.
Having data is not enough
The argument of this chapter fits in one sentence. For enterprise AI, having data is not enough. The advantage comes from usable data, for the outcomes the strategy has already chosen.
Usable means four things at once: the right data, with the context that gives it meaning, reachable by the right people and systems, under the right controls. Notice what the sentence leaves out. It does not say build a data lake, or train a model on everything the company owns, or make all enterprise data perfect before starting. Those are programs. A strategy chooses.
The reason data has become the binding constraint is simple. Capable models can now be bought, rented or swapped, and your competitors have the same access you do. What they cannot buy is your operating history, what it means and who is allowed to use it. Marco Iansiti and Karim Lakhani, describing the firms that run on AI, put the data pipeline first among the four components of what they call the “AI factory”, ahead of the algorithms themselves2. The model is the visible part. The data is the part that has to be built.
Most large organizations already have an enterprise data strategy: platforms, analytics, master data and governance, usually run by a chief data officer. A data strategy for AI should not compete with it. It adds the requirements AI brings, such as company knowledge held in documents, the context that makes records interpretable, the data used to test AI systems, and rules for what an AI may reach. Starting a separate “AI data platform” by default duplicates the foundation and splits the people who know it best.
Start from the outcome, not the data estate
One of the most expensive questions in this field is “How do we make all our data AI-ready?” It sounds responsible, and it turns into a multi-year modernization program with no customer.
A better question runs in the other direction, starting from an opportunity the AI strategy has already chosen. How those opportunities are found and ranked is the work of Finding Strategic AI Opportunities; this chapter takes the shortlist as given.
The list that results is short, specific and fundable, which is the difference between a strategy and a program.
Starting from the outcome also tells you how good the data has to be. That turns out to be the most useful thing a data strategy can know.
Data that exists is not data you can use
Status reports about data usually stop at existence: the data is in a system, it shows up in a dashboard. For an AI system, existence is only the visible tip. Whether the data is usable depends on questions that rarely appear in a status report.
Access asks whether the systems that need the data can reach it, through a supported route rather than an export someone runs by hand. Quality asks whether it is accurate and complete enough. Freshness asks whether it is current enough for this use: a weekly load is fine for planning and useless for a same-day decision. Context asks whether a system, or a new employee, can tell what a value means. Lineage asks where the data came from and what happened to it on the way. Permissions ask whether it may be used for this purpose, by this person. Ownership asks who is accountable when any of the others fail.
Quality deserves a closer look, because it is where well-meant programs go wrong. In a study still cited three decades later, Richard Wang and Diane Strong asked data users what quality meant to them. Their answer was that quality is fitness for use, and that its dimensions fall into four groups: intrinsic qualities such as accuracy, contextual ones such as relevance and timeliness, representational ones such as interpretability, and accessibility3. The contextual group is the one that matters most here. The same data set can be good enough for one purpose and unacceptable for another. For example, customer addresses that are 90 percent correct may be fine for regional demand forecasting and dangerous for automated contract notices.
That is why “clean all the data first” is the wrong instruction. Set the minimum bar the chosen use case needs, test against it, and improve the data where the value requires it. Why quality matters more than sheer volume in how models learn is covered in Data, Models and Compute; the strategic point is that the use case, not a general standard, sets the bar.
Rate the gap: green, amber or red
Once each chosen opportunity has a list of the data it needs, every item on the list gets a simple rating. The scale is deliberately coarse: its job is to make the constraint visible to an executive committee.
Two rules keep the rating honest. First, rate against the use, not in general: “the customer master is amber for the service assistant” is useful, while “the customer master is amber” starts an argument. Second, a red rating is not always a funding request. Sometimes the cheapest response is to change the scope, for example by launching first where the data is already good enough, so that the gap shrinks to something the team can fix.
The output is a short list of gaps, each tied to the value it blocks. Funding the gap that blocks the most value first is the core discipline of a data strategy for AI.
Records, documents and the test set
Enterprise data comes in three broad shapes: structured records such as transactions and sensor readings, logs and events, and unstructured material such as contracts, manuals, emails and recordings. Prediction and optimization systems still run mostly on the first two. Generative AI has made the third far more valuable, because so much company knowledge lives in documents that no system could use well before.
Documents bring their own version of the iceberg. A policy manual exists in five versions, and only one is current. A contract was amended by a letter stored somewhere else. For generative AI, company knowledge usually reaches the model through retrieval, where the system finds relevant material and passes it to the model with the question; RAG and Enterprise Knowledge explains how that works. The strategic consequence is that retrieval can only find what is findable, current and clearly authoritative. Deciding which version of each document counts is data work, and someone has to own it.
One more data set belongs in every AI data strategy, and it is often the one missing: evaluation data. These are representative questions or cases, the answers or outcomes a good system should produce, and the edge cases and known failures it must handle. Without them, nobody can show that a system works, compare a cheaper model with the current one, or notice when quality slips. An evaluation set is a strategic asset that improves with every incident. Treat it like one, with a named owner and a budget.
Owners, products and permissions
Data becomes usable when someone is accountable for keeping it usable. The data-management profession has said this for years: the DAMA-DMBOK, the field’s standard reference, treats ownership and stewardship of data domains as foundations of governance, not extras4. Zhamak Dehghani’s data mesh pushed the idea further, arguing that the business domains that create data should own it and serve it to others as a product5.
You do not have to adopt the whole architecture to use the core idea. Treat each data set that a chosen AI opportunity depends on as a product: a named owner in the business, defined users, stated quality expectations, documentation and a lifecycle. Set standards centrally and leave meaning and ownership in the domains. The AI team consumes these products; it should not quietly become the owner of every domain’s data because it was the first to need it. The shared tooling that serves many teams was the subject of AI Platform Strategy.
Permissions complete the picture. The fact that a system holds data is not permission to use it for a new purpose, and an AI assistant must not show people information merely because it can reach it. How retrieval respects each user’s rights is covered in Privacy and Confidential Data. At the strategic level, purpose is the question to settle early, because it can turn a green data set red overnight.
Story: three years of sensor data at a copper mill
Take the position of a mining executive in 2018. Freeport-McMoRan, one of the world’s largest copper producers, wanted more copper from its mines in North and South America. The conventional route was to build capacity: more crushing equipment, haul trucks and large shovels. By the company’s later estimate, adding about 90,000 tonnes of copper a year that way would have cost 1.5 to 2 billion US dollars8.
At its Bagdad mine in Arizona, the concentrating mill had been collecting data for years from thousands of sensors. Imagine three proposals on your desk. The first funds the capacity expansion. The second launches a company-wide program to clean, consolidate and modernize plant data before attempting any analytics. The third picks one mill and one outcome, throughput, and tests whether the data the mill already has can support it. Decide before you read on.
Freeport chose the third route. From 2018, metallurgists, engineers and operators at Bagdad worked with data scientists on three years of data from the mill’s sensors8. The data did not need to be made perfect first. It needed context, and the people who ran the mill supplied it. The analysis showed that the mill was processing seven distinct types of ore, and that its standard control settings did not match the properties of all of them. Adjusting settings for each ore type, including the acidity in the flotation tanks, raised Bagdad’s output by about 9,000 tonnes of copper a year9.
Then Freeport scaled what had worked. In its results for the end of 2019, it announced a rollout across its North and South American operations, with estimated extra production of about 100 million pounds of copper in 2021 and about 200 million pounds in 2022, for capital of roughly 200 million US dollars10. Its chief executive, Richard Adkerson, called 200 million pounds an aspirational goal, to be reached “with very little capital investment”8. The Bagdad gain is a reported result. The rollout figures are the company’s own targets at the time, set against a conventional expansion costing several times as much; this chapter found no published figure for what the rollout itself delivered.
Three lessons carry beyond mining. The data already existed; what made it usable was a chosen outcome and the context that only operators and metallurgists could add. The scope was one mill, which kept the data question small enough to answer. And the investment in data followed the value: the rollout was funded after one site had shown what the data was worth. Not every data estate is as tidy as a mill’s sensor history, and documents scattered across ten systems are harder. The sequence still applies.
What this means for leaders
A data strategy for AI is not a separate plan sitting beside the AI strategy. It is the AI strategy’s list of chosen opportunities, translated into the data each one needs and the gap that blocks it. That makes it short, specific and fundable. Leaders should expect to see the gaps ranked by the value they block, a named business owner for every data set that matters, an evaluation set for every serious AI system, and the purpose of each use settled before the build begins. They should be wary of any plan whose first milestone is “all data ready”.
Check yourself
- We need a complete, modernized data platform before we can start with AI.
- A data set can be good enough for one use and unacceptable for another.
- If a system already holds the data, an AI system may use it.
- In a two-year study of 75 executives, only 3 percent of teams’ data-quality scores met the acceptable bar.
- The AI team should own the data that its systems use.
- Freeport’s copper gains at Bagdad came from cleaning all its plant data before any analysis began.
Reflection: what is below your waterline?
What comes next
Data and platforms do not build, run or govern anything on their own. Every step in this chapter, from rating a gap to owning an evaluation set, depends on people with the right skills in the right places. The next chapter, AI Talent and Capability Strategy, asks which capabilities to build inside the organization, which to source, and how leaders should plan the workforce their AI strategy needs.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
References
- Tadhg Nagle, Thomas C. Redman and David Sammon. Only 3% of Companies' Data Meets Basic Quality Standards. Harvard Business Review. 2017.
- Marco Iansiti and Karim R. Lakhani. Competing in the Age of AI: Strategy and Leadership When Algorithms and Networks Run the World. Harvard Business Review Press. 2020.
- Richard Y. Wang and Diane M. Strong. Beyond Accuracy: What Data Quality Means to Data Consumers. Journal of Management Information Systems 12(4), 5-33. 1996.
- DAMA International. DAMA-DMBOK: Data Management Body of Knowledge, 2nd edition. Technics Publications. 2017.
- Zhamak Dehghani. Data Mesh: Delivering Data-Driven Value at Scale. O'Reilly Media. 2022.
- European Parliament and Council of the European Union. Regulation (EU) 2016/679 (General Data Protection Regulation). Official Journal of the European Union. 2016.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
- International Copper Association. Freeport McMoRan Looks to the Future with Artificial Intelligence. International Copper Association. 2020.
- MetalMiner. Freeport-McMoRan Uses AI to Optimize Production at Arizona Copper Mine. MetalMiner (summarizing the Financial Times, Neil Hume). 2019.
- International Mining. Freeport to invest in data science, AI programs at North and South America mines. International Mining. 2020.
Further reading
- Zhamak Dehghani. Data Mesh: Delivering Data-Driven Value at Scale. O'Reilly Media. 2022.
- Richard Y. Wang and Diane M. Strong. Beyond Accuracy: What Data Quality Means to Data Consumers. Journal of Management Information Systems 12(4), 5-33. 1996.
Sources last verified 2026-10-10.