AI Academy · Book
Executives & Directors · Module 06 · Chapter 004

Privacy and Confidential Data

AI makes information easier to find, combine and pass on, and that is exactly what makes it a privacy risk. An approved platform is still a path that data travels along. Leaders who follow that path, and insist on minimum data, a defined purpose, minimum access and controlled retention, get the value without the leak.

≈ 16 min read

After this chapter you can

  • Trace where sensitive information travels through an AI system, from prompt and hidden context to output, logs and third parties.
  • Distinguish privacy risk from confidentiality risk.
  • Explain why retrieval must respect existing access rights, and why encryption alone does not.
  • Separate inference, logging and training when questioning a provider.
  • Name the GDPR duties and the US position, led by California's rules on automated decisions, that apply when AI uses personal data.

Before reading on, make a guess. Of all the data that employees copy, paste or upload into AI tools at work, what share is sensitive: source code, customer records, financial or HR information? One in twenty? One in ten?

Cyberhaven, a data-security company, tracks how data moves inside its customers’ organizations. In its 2025 report it found that 34.8 percent of the corporate data employees shared with AI tools was sensitive, up from 10.7 percent two years earlier1. Its 2026 report put the figure at 39.7 percent of all data movements into AI tools, and found that the average employee entered sensitive data into an AI tool about once every three days2.

The share of data employees put into AI tools that is sensitive rose from 10.7 percent in 2023 to 34.8 percent in 2025 and 39.7 percent in 2026.202310.7%202534.8%202639.7%Source: Cyberhaven AI Adoption and Risk Reports, 2025 and 2026 · 2026
Figure 6.4.1 Share of data going into AI tools that is sensitive, in one vendor’s customer telemetry. The 2026 bar counts data movements, the earlier bars data volume, so read the direction, not the exact rise.

These are one vendor’s measurements of its own customers, not a random sample of the economy, and the 2026 figure counts data movements rather than volume. Treat the exact numbers with care. The direction is harder to argue with. Employees are not mainly typing trivia into AI tools. They are giving them the information that does the work, and the work runs on sensitive information.

Follow the information

Many leaders, asked about AI privacy, ask whether the tool is secure or approved. Those are good questions, and they are not enough. The more useful question is where the information goes.

Information flows from prompt and context through the model to output, logs, and people and systems, with privacy risk at every step.Promptplus hiddencontextModelprocessesitOutputan answeror an actionLogsand historyPeopleand othersystemsWhat can it see? Who receives it? Where does it go? How long does it stay?
Figure 6.4.2 The AI data path. Privacy risk exists before, during and after the model answers, not only inside the model.

Information enters as a prompt, often with far more context attached than the user typed. The model processes it. An answer comes out, and the answer, the prompt and the documents behind them are often kept in logs and conversation history. Then the answer reaches people and other systems: a colleague, an email, a ticket, another vendor’s tool.

Each step is a place where personal or confidential information can reach someone it should not. That is why the US National Institute of Standards and Technology lists “privacy-enhanced” among the characteristics of trustworthy AI3, and why its profile for generative AI names data privacy as one of the risks that is new or worse with this technology, including the leakage and unauthorized disclosure of personal information4. Approval of a platform gives you stronger contracts and better controls. It does not tell you what data went in, why, who can see the answer, where it is stored or how long it stays.

Two risks that overlap

Privacy and confidentiality are often used as if they meant the same thing. They overlap, and the difference matters for who owns the problem.

Privacy is about people. It concerns how information relating to an identifiable person, a customer, an employee, a reader or a patient, is collected, used, shared, kept and protected. Confidentiality is about access. It concerns information that must not reach unauthorized people or organizations, whether or not anyone is named in it.

A two-by-two grid separating personal data from confidential information, with examples of each combination.YesNoPersonaldataNoConfidential information · YesPrivacy riskAttendee list from a public eventBoth risksHR investigation fileLow riskPublished press releaseConfidentiality riskStrategy, source code
Figure 6.4.3 An AI system can create privacy risk without a single secret, and confidentiality risk without a single name.

An attendee list from a public event is hardly a secret, but it is personal data and privacy rules apply. Unreleased strategy or source code contains no personal data and can still do great commercial damage if it leaks. An HR investigation file is both. Legal categories differ by country, so classify by sensitivity, business impact and legal requirement, and let the privacy team and the security team each own their half.

Use only the data the task needs

The cheapest privacy control is a question asked early: does this use actually need this data?

A funnel narrowing from everything available, to data with a defined purpose, to the fields the task needs, to data sent after masking.Everything the model could useData with a defined purposeFields the task needsSent after maskingLess data means less exposure
Figure 6.4.4 Narrow the data at each stage before it reaches the model. Every field removed is one that cannot leak.

Start from everything the model could technically use, which is almost always far more than the task requires. Narrow it to data with a defined, legitimate purpose. Customer support records do not become training data or marketing analytics simply because they are convenient. Narrow again to the fields the task needs. An assistant that answers questions about order status needs the order status, not the full customer profile, payment details and contact history. Then mask or remove what is left over before it is sent.

Some information never belongs in an ordinary prompt: passwords, private keys, access tokens and database credentials. People paste them by accident, usually while asking for help with a broken script. Secret scanning, redaction and data loss prevention tools exist to catch exactly this, and they are worth their modest cost.

European law makes this principle a duty rather than a preference. The GDPR requires a lawful basis and a defined purpose for processing personal data, and data protection by design and by default5. The sensible default anywhere is minimum data, minimum access and minimum retention, expanded only when the business case justifies it.

The model sees more than the prompt

When an employee types a short question, they see one line. The model may receive a great deal more.

Above the waterline, one short question; below it, chat history, retrieved documents, user profile, system instructions, database results and tool outputs.WHAT THE USER TYPEDOne short questionWHAT THE MODEL MAY RECEIVEChat historyRetrieved documentsUser profileSystem instructionsDatabase resultsTool outputs
Figure 6.4.5 A one-line question can move a large amount of enterprise information onto the AI path, most of it unseen.

Modern assistants attach chat history, retrieved documents, the user’s profile, hidden system instructions, database results and the outputs of other tools. None of it was typed, and most of it the user never sees. A question that looks small can open the whole file room. The user may not know that happened. The organization still needs to know what entered the path.

Agents raise the stakes, because an agent can not only see information but send it onward, by email, into tickets or into other systems. The question grows from “what can it see?” to “where can it send it, and what can it do?” How to bound an agent’s permissions is the subject of Agentic AI and Autonomous Actions, later in this module.

Retrieval must respect the rights people already have

Many enterprise assistants work by retrieval: the user asks, the system searches a body of documents, and the model answers from what it finds6. This design reduces how much data is sent and keeps answers current. It also creates a common privacy failure in enterprise AI.

Retrieval first checks whether the user already has access; authorized documents can inform the answer, unauthorized ones are blocked before reaching the model.Does this useralready haveaccess?YesYes - same rights asthe personDocument can informthe answerNoNo - not authorizedBlocked before the model
Figure 6.4.6 The decisive check happens before the model sees anything. An assistant should never have more access than the person asking.

If the search index ignores the permissions on the original files, the assistant can faithfully quote a document the person asking was never allowed to open. Nothing was invented. The model did its job well. The failure is one of authorization, and the security community now lists it among the top risks for these applications, under sensitive information disclosure and weaknesses in the stores that hold documents for retrieval7.

Notice what encryption does here: nothing. Encryption protects data in storage and in transit from outsiders. It does not decide whether the right person is asking for the right purpose. Authorization does. As Data Strategy for AI put it in Module 04, permissions have to travel with the data; retrieval is where that promise is tested.

Inference, logs and training are three questions

Leaders often ask a provider one question: “Do you train on our data?” It matters, but it is one of three.

Inference sends data through the model to an answer, logs store it, and training builds it into the model; each needs a different question.StageWhere the data goesWhat to askInferenceThrough the model to an answerWhere is it processed?Logs and historyInto stored recordsWho can read them, and for how long?TrainingInto the model itselfIs our data used, and can we opt out?
Figure 6.4.7 Three different destinations for the same data, each with its own question for the provider and the contract.

Inference is data passing through the model to produce an answer. Ask where that processing happens, because location can differ by region, service and plan. Logs and history are stored records of prompts, answers, documents and tool calls. They help with debugging and audit, and they are also a new store of sensitive data. Ask who can read them, how long they are kept and whether they can be deleted. Training is data shaping the model itself. Whether your data is used depends on the specific service, configuration and contract, and one provider’s practice tells you nothing about another’s.

Training deserves care because it is hard to undo. Europe’s data protection authorities have said that a model trained on personal data cannot automatically be treated as anonymous; it is anonymous only if it is very unlikely that the people in its training data can be identified or their data extracted through queries9. Once personal data shapes a model, deleting the original file does not take it out.

The chain is also longer than one vendor. Cloud hosting, search indexes, monitoring tools and plugins can each hold copies. Verify each answer against current documentation and the contract, and assume that deleting data in one place does not remove every copy.

Shadow AI is a design problem, not only a discipline problem

Employees adopt AI faster than policy can keep up. When the approved route is slow or missing, they use whatever is closest: a personal account, a free tool, a phone. That is shadow AI. Cyberhaven found that about a third of ChatGPT use at work in its customer base ran through personal accounts2.

A 2026 breach-cost study, as summarized by a law firm, reported shadow AI incidents in 43 percent of breached organizations, up from 20 percent a year earlier10.

There are two tempting responses, and both fail. A blanket ban feels safe, but people who need to get work done move to tools and devices where the organization sees nothing. No rules at all means data leaks quietly into whatever tool is nearest. What works is in between: approved tools that are genuinely easy to use, clear rules about which data may go into which tool for which purpose, and training that explains why. When the safe route is also the fast route, shadow AI shrinks. How to keep an inventory of the AI systems in use is covered in AI Inventory and Risk Classification, in Module 07.

Personal data brings legal duties, in Europe and the United States

When the data path carries personal data, the law follows it. The duties are different on each side of the Atlantic, and an executive should know the shape of both.

EU duties under the GDPR - lawful purpose, transparency, privacy by design and impact assessments - beside the US position of state laws led by California's 2026 and 2027 rules.EU: GDPRLawful basis and purpose (Arts. 5-6)Tell people how data is used (Arts. 13-14)Privacy by design (Art. 25)DPIA before high-risk use (Art. 35)US: no federal privacy lawAbout 20 state laws; California leadsRisk assessments for risky processing (2026)ADMT notice, opt-out and access (2027)
Figure 6.4.8 Europe has one regulation with broad duties. The United States has state laws, with California’s rules on automated decisions the reference point.

In the European Union, the GDPR has applied since 25 May 2018. For AI on personal data, four duties matter most. You need a lawful basis and a defined purpose. You must tell people how their data is used, including in AI processing. You must build data protection into the design and the defaults. And you must carry out a data protection impact assessment, a DPIA, before processing that is likely to be high risk. The UK regulator’s guidance, written for the UK GDPR, which mirrors the EU text here, is blunt: in the vast majority of cases, using AI on personal data will trigger the legal requirement for a DPIA11. Where an AI makes a decision with legal or similarly significant effects on a person without meaningful human involvement, Article 22 adds a further right. A leak of personal data through an assistant can also be a reportable breach, with clocks that Human Oversight and AI Incidents covers later in this module.

The United States has no comprehensive federal privacy statute. About 20 states have their own laws, and California’s is the reference point. Its updated regulations require a documented risk assessment for processing that poses significant privacy risk, a duty that applies to new high-risk processing from 1 January 2026. From 1 January 2027, a business that uses automated decision-making technology to make a significant decision about a California resident, in lending, housing, education, employment or healthcare, must give notice before use, offer an opt-out unless an exception applies, and answer access requests12.

Story: the AI connector that crossed company lines

On 1 May 2025 Asana launched a server built on the Model Context Protocol, an open standard for connecting AI assistants to business software. Through it, a customer’s AI assistant could read that customer’s tasks, projects and comments in Asana and answer questions about them13.

On 4 June the company found a logic flaw in the server. Under certain conditions, information from one customer’s Asana domain could be returned to MCP users in other organizations. Depending on how the connector was used, that could include task details, project metadata, team details, comments and uploaded files. A spokesperson told BleepingComputer that roughly 1,000 customers were affected14. Asana said the cause was a bug, not a hack, and there was no indication that anyone had exploited it or actually seen another organization’s data13.

The feature was a month old and in demand. Its owners could keep it running while engineers patched it, fix it quietly, or take it down and tell customers. Pause here and choose.

Asana took the server offline from 5 June, fixed the code, contacted affected organizations directly and reset every connection, so each customer had to reconnect before using it again. The service came back on 17 June13. Administrators were advised to review MCP access logs and the answers their assistants had generated, looking for anything that seemed to come from another organization14.

Three lessons carry beyond one company. The model did nothing wrong: it answered faithfully from what the connector handed it, and the failure lay in the access check, exactly where the permission tree above puts it. The exposure did not end with the fix, because anything that crossed would sit in another company’s assistant answers and history until someone found it. And for every customer, the first question was simply whether any employee had connected an assistant at all. Only an organization that knows its data paths can answer that in an afternoon.

What this means for leaders

Privacy in AI is not a property of the vendor or of the approval stamp. It is a property of the path, and the path is designed by your organization. Four habits make the difference. Ask for the data path before the demo. Make minimum data, defined purpose, minimum access and controlled retention the default, and require a reason to expand any of them. Insist that retrieval inherit people’s existing access, and that this is tested with real cases. And make the approved route easy enough that shadow AI loses its appeal.

Check yourself

  1. An approved enterprise AI platform means confidential data is safe.
  2. Encryption solves privacy.
  3. An AI assistant can leak information while answering faithfully.
  4. If the provider does not train on our data, there is no privacy risk.
  5. The United States has a comprehensive federal privacy law that covers AI.
  6. Regulators expect most AI uses of personal data to need a data protection impact assessment.

Reflection: trace one path

What comes next

This chapter looked at information reaching the wrong place through design gaps and everyday behavior. The next chapter, Security and AI Attacks, asks the harder question: what happens when someone deliberately tries to manipulate an AI system, or uses it as a new route into the organization?

Laws referenced

General Data Protection Regulation · EU

Regulation (EU) 2016/679

Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.

  • 2018-05-25 — Applies

Last verified 2026-10-08 · official text

California privacy rules on automated decisions (CCPA regulations) · US - California

California Consumer Privacy Act; CPPA regulations on ADMT, risk assessments and cybersecurity audits (approved by OAL Sept 2025)

The most concrete US privacy rule on AI. Businesses that use automated decision-making technology to make a significant decision about a California resident (finance or lending, housing, education, employment or pay, healthcare) must give notice before use, offer an opt-out unless an exception applies, and answer access requests. Processing that poses significant privacy risk needs a documented risk assessment. There is no comprehensive federal privacy statute; about 20 states have their own laws, and California's is the reference point.

  • 2026-01-01 — Updated CCPA regulations take effect; risk-assessment duty applies to new high-risk processing
  • 2027-01-01 — ADMT duties for significant decisions: pre-use notice, opt-out (with exceptions) and access (some firm alerts cite enforcement from 1 Apr 2027)
  • 2028-04-01 — Attestation of 2026-2027 risk assessments due to the CPPA; cybersecurity audits phase in 2028-2030 by revenue

Last verified 2026-10-06

References

  1. Cyberhaven Labs. 2025 AI Adoption and Risk Report. Cyberhaven. 2025.
  2. Cyberhaven Labs. 2026 AI Adoption and Risk Report. Cyberhaven. 2026.
  3. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
  4. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
  5. European Parliament and Council of the European Union. Regulation (EU) 2016/679 (General Data Protection Regulation). Official Journal of the European Union. 2016.
  6. Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
  7. OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
  8. Matthew Finnegan. Microsoft 365 Copilot rollouts slowed by data security, ROI concerns. Computerworld. 2024.
  9. European Data Protection Board. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models. EDPB. 2024.
  10. Baker Donelson. Ten Takeaways from IBM's 2026 Cost of a Data Breach Report. Baker, Donelson, Bearman, Caldwell & Berkowitz. 2026.
  11. Information Commissioner's Office (UK). Guidance on AI and data protection: What are the accountability and governance implications of AI?. ICO. 2023.
  12. White & Case. CPPA finalizes rules on ADMT, risk assessments, and cybersecurity audits requirements under the CCPA. White & Case LLP. 2025.
  13. Jessica Lyons. Asana MCP server back online after plugging a data-leak hole. The Register. 2025.
  14. Bill Toulas. Asana warns MCP AI feature exposed customer data to other orgs. BleepingComputer. 2025.

Further reading

Sources last verified 2026-10-08.