Security and AI Attacks
AI systems take instructions and content through the same channel, language, so anyone who can put words in front of them can try to steer them. Prompt injection has no structural fix today. The leaders who stay safe assume some manipulation will succeed and limit what a fooled system can access, change, send or trigger.
After this chapter you can
- Explain why natural-language input creates a new security boundary around AI systems.
- Distinguish direct injection, indirect injection, jailbreaks and hallucination.
- Recognize attacks beyond the prompt, including poisoning, evasion and model extraction.
- Apply the dangerous-combination test (untrusted input, sensitive data, power to act) to an AI system.
- Ask for layered controls outside the model and a working emergency stop rather than a perfect filter.
Picture an ordinary business email arriving in an employee’s inbox. It reads like routine correspondence, addressed to the person who receives it. Nobody needs to open it. Some days later the employee asks the company’s AI assistant a question on a related topic, and the assistant, searching the mailbox for useful context, reads that email too. Woven into its wording are instructions. The assistant follows them: it gathers sensitive material from the files and messages it can reach, and tucks it into a link that quietly delivers it to a server the attacker controls.
This is not a thought experiment. Researchers at a security start-up reported exactly this chain to Microsoft in January 2025, in Microsoft 365 Copilot. They called it EchoLeak. Microsoft rated it 9.3 out of 10 for severity, fixed it on its own servers, and found no evidence that anyone had used it maliciously12. The researchers described it as the first real-world “zero-click” prompt injection in a production AI system. To get past the assistant’s injection filter, the email addressed its instructions to the human reader rather than to the AI3.
Notice what did not happen. Nobody stole a password. No malware ran. The model was not broken; it did what it was built to do, which is read content and act helpfully on it. Privacy and Confidential Data looked at information reaching the wrong place by accident. The harder case is someone making it happen on purpose.
Language is now an attack surface
Traditional software keeps a hard line between code and data. A customer’s address is stored and displayed; it is never run as a command. Security teams spent decades building and defending that line. AI systems blur it. Instructions and content arrive through the same channel, language, and the model has no reliable way to tell which words are its owner’s instructions and which are simply material it was asked to read.
That matters because modern AI systems do not only talk. They search document stores, read email, call tools, update records and send messages. So the security question changes. It is no longer only whether the model is safe. It is what the whole system can reach and do if someone steers it.
Securing only the model protects one box of four. Identity, permissions, the content the system reads, its tools, its outputs and its actions all belong inside the security boundary. The security community’s reference list of risks for language-model applications puts prompt injection in first place, and its list reads like a tour of that chain: sensitive information disclosure, poisoning, improper handling of outputs, excessive agency and runaway consumption4. The US government’s generative AI risk profile treats information security as one of a dozen core risks for the technology, including attacks by injection and poisoning5.
Four failures that look alike
Four terms get confused in board discussions, and the confusion leads to the wrong fix.
Direct prompt injection is the simplest: the attacker types hostile instructions straight into the system, hoping to bypass its rules, reveal its hidden configuration or trigger a tool. Indirect prompt injection is often the version that matters more for enterprises, because the attacker never talks to your AI at all. The instruction waits inside content the AI will later read: a web page, a support ticket, a contract, a supplier’s PDF, an email. In 2023, Kai Greshake and colleagues showed this working against real, deployed assistants, including a search engine’s chat feature, and argued that retrieval turns every document an AI can read into a possible source of commands6. EchoLeak was the same idea, two years later, inside a corporate mailbox.
A jailbreak is a prompt designed to get around the model’s own safety training. It is a reason to keep controls outside the model, not a separate security program. Hallucination, finally, is not an attack. Nobody is manipulating anything; the model simply produces something wrong, as Accuracy, Hallucination and Reliability showed. The distinction matters because the fixes differ. Better grounding reduces hallucination. It does nothing against an instruction someone planted on purpose.
Attacks on the data and the model
Prompt injection gets the headlines, but it is one family among several. The US National Institute of Standards and Technology keeps a taxonomy of attacks on machine learning. It sorts them into evasion, poisoning and privacy attacks, with model extraction among the privacy attacks, and adds direct and indirect prompt injection for generative systems7. Three of these have been demonstrated convincingly enough that executives should know them by name. Most of the evidence below comes from research demonstrations; public records of these attacks exploited in the wild are thinner, which is a reason to design ahead rather than a reason to relax.
Poisoning plants tainted material in the data a model learns from, so that it misbehaves later, often long after the plant. A common assumption was that this would need a meaningful share of a huge training set. In October 2025, researchers from Anthropic, the UK AI Security Institute and the Alan Turing Institute found otherwise: as few as 250 malicious documents created a hidden “backdoor” in models ranging from 600 million to 13 billion parameters, whatever their size8. The backdoor was deliberately harmless and the work a controlled experiment, but the lesson about scale stands. The same logic applies to the knowledge bases and feedback data your own systems learn from.
Evasion crafts inputs that fool a model at the moment of decision. In a well-known demonstration, a few stickers made a road-sign classifier read a stop sign as a speed limit sign in most video frames from a moving vehicle9. In business the target is more often a fraud screen, a content filter or a document check.
Model extraction uses ordinary queries to recover what a model has learned: in 2024 a research team did so for under 20 US dollars against two production language models10. A model that cost millions to build can leak through its own front door.
The other three families are quieter. Data leakage happens when the system reveals what it can reach, which is why Privacy and Confidential Data insisted that retrieval respect the rights people already have. Unsafe output is generated code, queries or emails passed into other systems without checks, which turns the AI into one link in an attack chain. Runaway cost is a flood of long requests or an agent stuck in a loop, which turns a security incident into a financial one4. Weaknesses in the models, libraries and services you buy are a supply chain question, and Model and Third-Party Risk takes them up.
The dangerous combination
Not every AI system is equally exposed. In June 2025 the developer Simon Willison gave a useful test a memorable name, the lethal trifecta. Any system that combines access to private data, exposure to untrusted content and the ability to communicate externally can be tricked into sending that data to an attacker11. EchoLeak had all three: a mailbox full of private material, an inbox open to anyone on the internet, and a way to reach an outside server.
In October 2025 Meta’s security team turned the idea into a design rule, the Agents Rule of Two. In any one session, an agent should have at most two of three properties: it processes untrustworthy input, it has access to sensitive systems or private data, or it can change state or communicate externally. If a task truly needs all three, the agent should not operate autonomously and needs supervision, such as human approval12.
The power of the rule is that it does not depend on catching the hostile sentence. An assistant that reads outside email and holds private data, but cannot send anything anywhere, can be fooled without being able to deliver the prize. Security teams call everything a manipulated system could access, change, send or trigger its blast radius. The Rule of Two is a way to keep it small by design. Agentic AI and Autonomous Actions, later in this module, takes up the ladder of autonomy itself and how to widen it on evidence.
Why no filter will save you
Your security teams will recognize the pattern. It looks like SQL injection, the attack in which someone types a database command into a web form and the system runs it. The comparison holds in three ways: untrusted input becomes a command, the defense starts by treating all outside input as hostile, and nobody mistakes it for a user being creative.
Here is where the analogy breaks, and it breaks in the place that matters most. SQL injection has a structural fix. Parameterized queries send commands and data down separate channels, so the database never confuses the two. Language models have no equivalent. Everything they read arrives as one stream of words. In December 2025 the UK’s National Cyber Security Centre warned that treating the two attacks as alike is dangerous, because prompt injection may never be mitigated the way SQL injection was, and advised designing systems to limit the impact instead13. The same month, a leading model provider wrote that prompt injection is “unlikely to ever be fully ‘solved’”14.
The evidence on defenses points the same way. In October 2025, researchers from three major AI developers tested twelve published defenses against jailbreaks and prompt injection. Most had reported near-zero attack success when they were introduced, measured against fixed sets of known attacks. Against attackers who adapted to each defense, success rates rose above 90 percent for most, and human red-teamers got through every one15.
None of this means filters are useless. They raise the cost of an attack and stop the lazy ones. It means a filter cannot be the reason you believe a system is safe. So with SQL injection we eliminate the risk; with prompt injection we contain it. We cannot eliminate it.
Contain it: controls outside the model
If some manipulation will succeed, the design question becomes how to make sure one success does not become an incident. The answer is layers, each enforced outside the model, so that a fooled model meets a wall it cannot talk its way past.
The foundation is identity and least privilege. Each AI system runs under its own identity, not a shared, powerful service account, and gets only the data, tools and permissions its task needs, not the full rights of the employee it helps. The boundary is enforced by the systems around the model. An AI that says it is acting for an employee has not been authorized by saying so.
Validation comes next: checking what goes in, what comes out and every tool call against an allowlist, so that generated code is tested before it runs and an outgoing message is checked before it leaves. Human approval belongs on high-impact or irreversible actions, such as payments, deletions, legal messages and account changes. It is a real control only if the approver can see what they are approving. Lisanne Bainbridge’s classic paper on the ironies of automation warned in 1983 that people asked to watch over automated systems are poorly placed to catch the rare failure16. Monitoring and red teaming mean logging what the system actually did and paying people to break it before attackers do. At the top sits the emergency stop: the ability to halt an agent, disable its tools, revoke its credentials and block its traffic within minutes, owned by someone who can act without a meeting.
One more point belongs here, because it is a common mistake in boardrooms. A model provider secures its model service and infrastructure. Your prompts, data, permissions, tools and business logic remain yours. Security is shared, and widely used frameworks, such as NIST’s AI Risk Management Framework, treat security and resilience as properties of the whole system and its operators17.
Story: the web form that steered a sales agent
Salesforce’s Agentforce puts AI agents to work inside a company’s customer-relationship system. One ordinary feature of that system is Web-to-Lead: a public form on a company’s website where a prospect leaves a name, a company and a free-text description, which becomes a lead record. Anyone on the internet can fill it in.
In July 2025 Noma’s researchers showed what that description field, which accepts up to 42,000 characters, could carry. They submitted a lead whose description held instructions written to look like part of a normal request. When an employee later asked the agent an everyday question about the new lead, the agent followed the planted instructions as well: it gathered other CRM records and packed them into the address of an image, so that loading the image would send the data to an outside server19.
The page would normally refuse to load an image from an unknown site. Here it did not, because the server sat on a domain still listed among Salesforce’s trusted addresses, although its registration had lapsed. The researchers bought it for about 5 US dollars19. The researchers themselves rated the chain critical, 9.4 out of 10 on the CVSS severity scale (the 9.3 for EchoLeak was Microsoft’s own rating), and reported it on 28 July 2025 [@noma-forcedleak-2025; @thehackernews-forcedleak-2025]. Measured against the Rule of Two, the agent had all three properties in one session: untrusted input from a form anyone could fill in, sensitive customer data, and a way to send that data out.
The fix sat outside the model. From 8 September 2025 Salesforce enforced its Trusted URLs allowlist for Agentforce and Einstein AI, so that agent output could no longer be sent to untrusted addresses, and it re-secured the expired domain [@noma-forcedleak-2025; @thehackernews-forcedleak-2025]. In its words, the services behind Agentforce “will enforce the Trusted URL allowlist to ensure no malicious links are called or generated through potential prompt injection”20. The research was published on 25 September21.
Two lessons outlast the patch. A planted lead may still fool the model, but fooling it no longer delivers the prize, because a control the model cannot talk its way past has cut the outbound leg of the dangerous combination. And that control needed upkeep: an allowlist entry for a domain nobody was renewing had quietly become the attacker’s way out. A wall outside the model protects only as long as someone owns it. Customers had homework too; the researchers advised auditing existing lead data for suspicious submissions19.
What this means for leaders
Four lessons follow. Treat every AI system that reads outside content as exposed, because any document, email or web page can carry instructions. Ask for blast radius, not cleverness. Put the controls outside the model, where a fooled model cannot talk its way past them. And distrust any claim that prompt injection is solved: ask how a defense performed against adaptive attackers and human red-teamers, not against a fixed list of known attacks.
Check yourself
- An attacker must interact with your AI system to inject instructions into it.
- A secure model means a secure AI application.
- An assistant that only reads data cannot cause a security incident.
- Prompt injection can be engineered away, the way parameterized queries fixed SQL injection.
- A few hundred poisoned documents can backdoor a language model, whatever its size.
- An agent that reads untrusted email and can send messages, but holds no sensitive data, fits the Rule of Two.
Reflection: find the dangerous combination
What comes next
This chapter looked at how AI systems can be manipulated or compromised by someone acting on purpose. The next chapter, Intellectual Property and Copyright, turns to a different exposure: what happens to ownership, licensing and confidentiality when AI creates, transforms and reuses content.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
NIS2 Directive · EU
Directive (EU) 2022/2555
Cybersecurity risk management for essential and important entities. Significant incidents: early warning within 24 hours, incident notification within 72 hours, final report within one month. Management bodies are accountable.
- 2024-10-18 — Applies through national law
Last verified 2026-10-06 · official text
SEC cybersecurity incident disclosure · US - listed companies
Form 8-K Item 1.05 (SEC rule adopted July 2023)
Listed companies disclose a material cybersecurity incident within four business days of determining that it is material. AI-related breaches are included.
Last verified 2026-10-06
References
- Microsoft Security Response Center. CVE-2025-32711: M365 Copilot Information Disclosure Vulnerability. Microsoft Security Update Guide. 2025.
- The Hacker News. Zero-Click AI Vulnerability Exposes Microsoft 365 Copilot Data Without User Interaction. The Hacker News. 2025.
- Pavan Reddy and Aditya Sanjay Gujral. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System. AAAI Fall Symposium Series 2025 (arXiv:2509.10540). 2025.
- OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
- Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
- National Institute of Standards and Technology. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2 E2025. NIST. 2025.
- Alexandra Souly et al. (Anthropic, UK AI Security Institute, Alan Turing Institute). A small number of samples can poison LLMs of any size. Anthropic (paper arXiv:2510.07192). 2025.
- Kevin Eykholt et al. Robust Physical-World Attacks on Deep Learning Visual Classification. CVPR 2018 (arXiv:1707.08945). 2018.
- Nicholas Carlini et al. Stealing Part of a Production Language Model. ICML 2024 (Best Paper; arXiv:2403.06634). 2024.
- Simon Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. simonwillison.net. 2025.
- Meta AI. Agents Rule of Two: A Practical Approach to AI Agent Security. Meta AI blog. 2025.
- Dave Chismon (UK National Cyber Security Centre). Prompt injection is not SQL injection (it may be worse). NCSC blog. 2025.
- OpenAI. Continuously hardening ChatGPT Atlas against prompt injection. OpenAI. 2025.
- Milad Nasr, Nicholas Carlini et al. (OpenAI, Anthropic, Google DeepMind). The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections. arXiv:2510.09023. 2025.
- Lisanne Bainbridge. Ironies of Automation. Automatica 19(6). 1983.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 15: Accuracy, robustness and cybersecurity. Official Journal of the European Union. 2024.
- Noma Security (Noma Labs). ForcedLeak: AI Agent risks exposed in Salesforce Agentforce. Noma Security. 2025.
- Ravie Lakshmanan. Salesforce Patches Critical ForcedLeak Bug Exposing CRM Data via AI Prompt Injection. The Hacker News. 2025.
- Alessandro Mascellino. Critical Vulnerability in Salesforce AgentForce Exposed. Infosecurity Magazine. 2025.
Further reading
- OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
- National Institute of Standards and Technology. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2 E2025. NIST. 2025.
- Dave Chismon (UK National Cyber Security Centre). Prompt injection is not SQL injection (it may be worse). NCSC blog. 2025.
- Simon Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. simonwillison.net. 2025.
Sources last verified 2026-10-08.