Multimodal and Agentic AI
Systems that see, hear and act are no longer two product categories. They are merging into one system that perceives the work and acts through the same screens, voices and connections people use. The leadership decisions move with them: which door the system uses into your software, what it is allowed to watch, and where a person still signs.
After this chapter you can
- Explain how multimodal perception and agentic action are merging into single systems that work through screens, voice, sensors and connectors.
- Compare the two doors into enterprise software - connectors and the screen - on reach, reliability, cost and auditability.
- Describe how open connection standards are changing integration cost and why connectors belong in the AI inventory.
- Distinguish prompted from ambient AI and identify the legal questions that watching and listening raise.
- Test an agent claim by its action log rather than its label, and apply the lessons of a documented stroke-network case.
In June 2025 the research firm Gartner published two forecasts about AI agents in the same press release. By 2028, it expected a third of enterprise software applications to include agentic AI, up from less than 1 percent in 2024. And by the end of 2027, it expected more than 40 percent of agentic AI projects to be canceled, because of escalating costs, unclear business value or inadequate risk controls1.
The same release put a number on the noise. Many vendors, Gartner said, were rebranding existing assistants, chatbots and robotic process automation as agents, a practice it called “agent washing”. Of the thousands of vendors claiming agentic products, it estimated that only about 130 were real1.
These are an analyst firm’s forecasts from mid-2025, not measurements, and both target dates are still ahead, so they should be read as informed opinion. Both can be true at once, and that is the useful thing about them. The capability is real and spreading fast. So is the gap between what is sold as an agent and what can safely be given work. Closing that gap is a leadership task, not a technical one, and it starts with being clear about what has actually changed.
Seeing and doing become one system
Multimodal AI showed that a model can read images, audio, video and documents as well as text. From AI Assistants to AI Agents showed what turns a model into an agent: a goal, a plan, tools and a loop that checks results. Neither idea needs repeating here. What is new is that the two are fusing.
Until recently, perception and action were separate products. One system described a photograph; another, scripted in advance, moved data between applications. In 2026 a single system can look at a screen and click on it, listen to a caller and answer, read a scan and send the alert. Perception is no longer only how the system understands a case. It is how the system finds its way around the software and the people it works with.
That merger changes the executive questions. Whether a system can do a task is increasingly settled, as Where AI Is Going Next showed, and the reliability, cost and autonomy it does it with are the dials to watch. Three new questions sit on top of those dials. Through which door does the system enter your software? What is it allowed to see and hear, and for how long is that kept? And at which point does a named person still sign? The rest of this chapter takes them in turn.
The screen becomes a door
The clearest sign of the merger is the computer-using agent. In October 2024 one AI lab released a model that operates a computer the way a person does: it looks at a screenshot, decides what to do, moves the cursor, clicks and types. The lab itself called the capability experimental and “at times cumbersome and error-prone”2. In January 2025 a second lab released a similar agent that reads screenshots and works through the buttons, menus and text fields people see3.
The attraction for an enterprise is obvious. Much of the work that never got automated sits in old applications with no programming interface, in supplier portals, or in screens stitched together by people copying from one window into another. An agent that can work the screen reaches all of that without an integration project.
The cost is equally plain. Screen-working agents still fail often on ordinary computer tasks, roughly one in three on the benchmark described in Where AI Is Going Next4. They are slow, because every step needs a fresh look at the screen. And they read whatever the screen shows, including text planted to mislead them. The same system card that describes the method lists prompt injection from web content as an added risk, and it relies on asking the user to confirm before critical actions3. How such attacks work, and why limiting what a system can reach is the main defense, is the subject of Security and AI Attacks5.
Two doors into your systems
Every agent that does work in your organization enters your software through one of two doors. The back door is a connector: a programming interface or tool that the system’s builders have deliberately exposed, with defined actions and permissions. The front door is the screen: the agent uses the same interface a person would.
The choice is not technical housekeeping. It decides how far an agent can reach, and therefore how much harm a mistake can do. A connector exposes a short list of actions, such as “create a draft order” or “look up a part”, and each one can be permitted, limited and logged. A screen exposes everything the logged-in account can see and touch.
That suggests a simple rule for leaders. Use connectors for anything consequential or high volume. Allow screen work for the gaps, where no connector exists, the stakes are low and someone is watching, and treat each such use as a signal that a connector is worth building. The controls themselves, from least privilege to approval and reversibility, belong to Agentic AI and Autonomous Actions; the door decides how many of them you will need.
The plumbing is being standardized
The back door is getting easier to build, because the connections are being standardized. In November 2024 one lab published the Model Context Protocol, an open standard for connecting AI systems to data sources and tools in place of one-off integrations6. In April 2025 another introduced Agent2Agent, an open protocol for agents from different builders to discover each other’s capabilities and coordinate tasks7.
Both standards moved quickly to neutral homes. In June 2025 Agent2Agent became a Linux Foundation project, founded with seven large technology companies and supported by more than 1007. In December 2025 the Model Context Protocol moved to a new foundation under the Linux Foundation, co-founded by rival AI labs and backed by the largest cloud providers, and reported about 10,000 active servers one year after launch8.
For an executive, the meaning is practical. The cost of connecting an agent to a new system is falling, so connectors will multiply, many of them added by teams or suppliers without a central decision. Each one is a new door. Treat connectors as you treat AI systems in AI Inventory and Risk Classification: listed, owned and reviewed. A connector someone installed from a public directory is part of your supply chain, and it can carry instructions as well as data.
From prompted to ambient
The third shift is in when the system works. Most AI so far has been prompted: a person asks, the system answers, and nothing happens between requests. Multimodal agents increasingly work ambiently. They listen to a meeting or a call, watch a camera feed or a queue of scans, and act when something they perceive matches a trigger.
Ambient systems are often where the value is, because nobody has to remember to ask, and the delay between an event and a response shrinks. But the governance questions change character. A prompted system sees what someone chose to give it. An ambient one sees everyone in range: colleagues on the call, customers in the shop, patients in the room. The questions become who is being recorded, whether they know, what is kept, and who may look at it later.
Law already reaches some of these choices.
Read the label, test the behavior
Gartner’s estimate that only about 130 agentic vendors are real is a forecast-firm judgment, not a census1. The pattern behind it is easy to check for yourself. From AI Assistants to AI Agents gave the test for whether a system plans, uses tools and acts. With multimodal agents, add three demands to any vendor or internal team.
First, a log of actions taken on real cases, not a demonstration. Second, for each action, the door it used: a connector with defined permissions, or the screen. Third, the points at which a person approved, corrected or stopped the system, and how often. A product that cannot produce these is, for your purposes, an assistant, whatever its name. That is not a reason to reject it. It is a reason to price, govern and measure it as an assistant.
What a satisfactory answer looks like is already documented. Since May 2025 the software platform GitHub has offered a coding agent that takes a task, changes code and proposes the change for review; it is named here as evidence of a design, not as a recommendation11. The agent works through the platform’s own interfaces, not a screen, and can push only to a single branch created for it. It cannot approve or merge its own pull request, and by default the automated checks on its code do not run until a person with write access approves them. Every commit it makes links to the session log of what it did, and administrators receive audit-log events12. The door, the action log and the human signature are part of the product. That is the bar to hold other agents to.
Story: the stroke alert that reached three hospitals
The clearest real example of perception joined to action is older than the current wave of agents, and smaller than most agent pitches. That is what makes it instructive.
A comprehensive stroke center in upstate New York treats patients with large-vessel strokes, which need a clot removed as fast as possible. It receives them from 14 hospitals within about 75 miles, only three of which belong to its own network13.
Before. A patient arrived at a local hospital and had a brain scan. Software processed the perfusion scan and e-mailed results to a list of providers, for every scan whether or not it showed an occlusion. A radiologist reviewed the images and called the care team; the team then arranged the transfer. The alerts were many, and they were easy to ignore.
After. In 2021 the network added software that reads the CT angiogram, flags a suspected large-vessel occlusion, and pushes the alert and the images to the phones of both the local doctors and the thrombectomy team at the center, with secure messaging between them. The system perceives, and it takes one small action: it tells the right people, with the evidence attached. Every treatment decision stays with physicians13.
The results, across 262 patients who had a clot removed, were real and specific. The time from arrival at the center to the start of the procedure fell from about 95 to about 80 minutes. Transfers from the three network hospitals fell from 157 to 120 minutes. But transfer time from all referring hospitals barely moved, from 162 to 159 minutes: the eleven hospitals outside the network were not on the system13.
Two further details matter as much as the gains. The first is what the software caught, which takes two numbers to state. For the occlusions it was designed to find, in the terminal internal carotid artery and the first segment of the middle cerebral artery, it detected 64 of 73, or 88 percent. Across all 111 patients treated after it was introduced, it flagged 69, or 62 percent, because many patients had occlusions outside its design scope, mostly in the carotid artery in the neck, in smaller branches of the middle cerebral artery or at the back of the brain. Those patients still reached treatment through the existing radiologist pathway, which the network kept. The second detail is that the study measured workflow times, not patient outcomes13.
Read as a trajectory, the case contains the whole chapter. The value came from joining perception to a modest, well-chosen action, not from autonomy. The benefit stopped exactly where the connections stopped. And the human pathway was kept, because perception was good but not complete.
What this means for leaders
The capability will keep improving along the dials described in Where AI Is Going Next. The decisions that stay with you are about design. Choose workflows where the evidence arrives as images, sound or screens and the delay is in getting it to the right person or system, because that is where joining perception to action pays first. Prefer connectors to screens for anything that matters, and register every connector. Decide what an ambient system may watch and keep before it is switched on, not after the first complaint. And keep the human pathway running in parallel until the record shows the system catches what people catch.
Check yourself
- Multimodal and agentic AI are still separate product categories.
- An agent that operates the screen is safer than one that uses a connector, because it does only what a person could do.
- Within about a year, the main agent connection standards moved to neutral, multi-company governance.
- Gartner forecast both rapid growth in agentic software and the cancellation of over 40 percent of agentic projects.
- Ambient systems raise mainly technical questions, not legal ones.
- In the stroke network, the alert software made transfer faster from every referring hospital.
Reflection: the evidence that waits
What comes next
Everything in this chapter still happens inside software, even when the input is a camera and the output is a phone alert. The next step takes the same pattern of perceiving and acting into machines that move. The next chapter, AI + Robotics, looks at what changes when an action is physical and cannot be undone with a click.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
EU Digital Omnibus on AI · EU
Regulation (EU) 2026/1744
First amendment to the AI Act. Defers high-risk obligations (Annex III to 2 Dec 2027, Annex I to 2 Aug 2028), adds two prohibited categories, softens the Art. 4 AI-literacy duty to "take measures to support", and simplifies some compliance duties. Art. 50 transparency duties still apply from 2 Aug 2026, with one transition (new Art. 111(4)): providers of generative AI systems placed on the market before 2 Aug 2026 must meet the Art. 50(2) marking duty by 2 Dec 2026.
- 2026-07-24 — Published in the Official Journal
- 2026-07-27 — Entered into force
- 2026-12-02 — Grace period ends for safeguards against two new prohibited uses (non-consensual intimate imagery, child sexual abuse material)
- 2026-12-02 — Art. 50(2) marking duty applies to generative AI systems placed on the market before 2 Aug 2026 (Art. 111(4))
Last verified 2026-10-10 · official text
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
References
- Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner Newsroom. 2025.
- Anthropic. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. Anthropic. 2024.
- OpenAI. Operator System Card. OpenAI. 2025.
- Stanford Institute for Human-Centered AI (HAI). AI Index Report 2026. Stanford University. 2026.
- Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
- Anthropic. Introducing the Model Context Protocol. Anthropic. 2024.
- Google. Google Cloud donates A2A to Linux Foundation. Google Developers Blog. 2025.
- Model Context Protocol project. MCP joins the Agentic AI Foundation. Model Context Protocol blog. 2025.
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024.
- European Union. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689. Official Journal of the European Union. 2026.
- GitHub. GitHub Copilot coding agent in public preview. GitHub Changelog. 2025.
- GitHub. Risks and mitigations for GitHub Copilot cloud agent. GitHub Docs. 2026.
- Nicholas C. Field, Pouya Entezami, Alan S. Boulos, John Dalfino, Alexandra R. Paul. Artificial intelligence improves transfer times and ischemic stroke workflow metrics. Interventional Neuroradiology (doi:10.1177/15910199231209080). 2023.
Further reading
- Gartner. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner Newsroom. 2025.
- Nicholas C. Field, Pouya Entezami, Alan S. Boulos, John Dalfino, Alexandra R. Paul. Artificial intelligence improves transfer times and ischemic stroke workflow metrics. Interventional Neuroradiology (doi:10.1177/15910199231209080). 2023.
- Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
Sources last verified 2026-10-08.