AI Academy · Book
Executives & Directors · Module 02 · Chapter 012

Multimodal AI

Much of what an organization knows arrives as screenshots, photos, recordings and video, and AI can now work with many of them. That makes each extra type of input a design choice, not an upgrade: it brings information, cost, new errors and personal data. Add a modality when it holds information that changes the decision.

≈ 14 min read

After this chapter you can

  • Define a modality and explain how a system can be multimodal in its inputs, its outputs or both.
  • Explain in plain terms how a model takes in images and sound, and why that drives cost.
  • Compare a pipeline of specialist models with one unified multimodal model.
  • Recognize the failure modes, costs and personal data that images, audio and video add.
  • Decide whether an extra modality earns its place in a use case.

Here is a support ticket, invented for illustration but typical of what many business-software companies receive.

An assistant that can read only text learns almost nothing from it. Everything useful sits in the attachments. The screenshot shows the error code. The screen recording shows the three clicks that trigger it. The voice note says which report fails, and that it started after Tuesday’s update. A support engineer would open all three before deciding anything.

A great deal of what organizations know arrives like this: as photos, scans, recordings, diagrams and video, not as tidy text. The models that read the ticket’s text, described in Large Language Models, now increasingly read the attachments too. The question for a leader is not whether that is impressive. It is when the attachment is worth reading, and what reading it costs.

A modality is a choice, not an upgrade

A modality is a type, or channel, of information: text, images, audio, video, and other signals such as sensor readings, telemetry and location. A unimodal system works with one of them; a photo goes in and a label comes out. A multimodal system can process, combine or generate more than one. The research field treats combining modalities as its own problem, not as a sum of single-modality skills: the system has to represent each input, line them up with each other and fuse them into one judgment1.

There are two ways to get the choice wrong, and organizations make both.

A spectrum from text only to every modality, with the right stance in the middle - add the modality that holds information that changes the decision.Text onlyMisses what the attachments holdEvery modalityMore cost, risk and noiseAdd what mattersThe input that changes the decision
Figure 2.12.1 The right stance sits between the extremes. Each modality has to earn its place.

The first mistake is to keep everything as text and lose what the attachments hold. The second is to make everything multimodal because it sounds more advanced, and then pay for it in cost, error and exposure. The rest of this chapter builds the case for the middle position: add a modality when it holds information that changes the decision.

In, out, or both

Multimodality runs in two directions, and product announcements often blur them. What a system can read and what it can produce are separate capabilities.

A two-by-two grid of input and output - text to text, understanding media, generating media, or both.MediaTextInputTextOutput · MediaUnderstandScreenshot plus question, answeredBothHears, sees, replies aloudText to textThe classic assistantGenerateA brief becomes an image
Figure 2.12.2 “Multimodal” says nothing about direction. Ask what goes in and what comes out.

Text in and text out is where most people first met generative AI. Media in and text out is understanding: a chart, a scanned contract or a call recording goes in, and an answer comes out. Text in and media out is generation: a written brief becomes an image, a voice or a short clip. Both at once is a system that hears a question, looks at a shared screen and answers aloud.

A model that is strong in one quadrant may be weak or absent in another. Reading images does not imply creating them, and transcribing speech does not imply replying in a natural voice. How models create new images and sound is the subject of the next chapter, How Generative AI Works. This chapter concentrates on understanding, because that is where many of today’s business uses, and many of the risks, sit.

How a model takes in a picture or a sound

You do not need the engineering, but one idea removes much of the mystery. As Tokens, Context and Embeddings showed, a language model reads text as tokens and works with them as numbers. Multimodal models extend the same trick to other inputs.

Four steps - cut images and audio into pieces, turn each piece into numbers, place them in the same space as words, and reason over them together.SliceImage patches,slices of sound,sampled framesEncodeEach piecebecomes a listof numbersAlignPictures andwords that meanthe same sit closeReasonOne modelweighs all thepieces together
Figure 2.12.3 Images and sound are cut into pieces and placed in the same space as words. That is what lets one model reason across them.

An influential 2021 paper showed that an image can be cut into a grid of small patches and fed to a transformer as if each patch were a word; its title was An Image is Worth 16x16 Words2. In the same year a second team trained a model on 400 million images paired with their captions from the internet, so that a picture and a sentence describing it land close together in one shared space. Without seeing any of the 1.28 million labeled training images of the standard ImageNet benchmark, it matched the accuracy of a classic model trained on all of them3. That shared space is why you can now search a photo library by describing what you want, or ask a question about a chart.

Two consequences follow for leaders. First, a modality is not “seen” or “heard”; it is converted into pieces and numbers, and anything lost in that conversion is lost to the model. Second, every piece counts toward the model’s context and its bill, which is why images, audio and video are priced the way they are.

One model or a team of specialists

There are two broad ways to build a multimodal application. The first is a pipeline of specialists: a speech model turns a call into text, a vision model reads the image, a language model writes the answer, and a coordinating layer joins the results. The second is one unified model that takes every input directly.

Voice assistants show the trade-off. When one widely used assistant replaced its three-model voice chain with a single unified model, its maker reported that average response time fell from 2.8 or 5.4 seconds to 320 milliseconds, and that the chain had lost tone, speakers and background noise [@openai-2024-hello-gpt-4o; @openai-2024-gpt-4o-system-card]. Every hand-off in a pipeline adds time and drops information.

That does not make the unified model the default answer. Pipelines have real strengths. Each part can be tested, tuned or replaced on its own; a narrow speech or vision model is often faster and cheaper at high volume; and the text in the middle is a record a person can inspect. Where the words carry the information, as in most meeting notes, a transcript followed by a language model may be all you need. Where tone, timing, layout or the relationship between inputs matters, translating everything into text first throws away the evidence.

Pipelines are cheap at volume, easy to change part by part and leave a readable text record; one unified model reasons across inputs directly.DesignReasons acrossinputsCheap at volumeSwap one partInspectable middlePipeline of specialistsOne unified model
Figure 2.12.4 Neither design always wins. Choose by what the task needs to keep, and what you need to check.

What each modality costs

Images, audio and video are not priced like text, because they become far more pieces. One major provider’s published counting rules, as of October 2026, make the scale concrete4.

One provider counts 258 tokens per image tile, 1,920 tokens per minute of audio and 15,780 tokens per minute of video.258Tokens per image tileLarger images are cut intoseveral tiles1,920Tokens per minuteof audio32 tokens a second15,780Tokens per minuteof video263 tokens a secondSource: One provider's published token rules · October 2026
Figure 2.12.5 An hour of video counts as about 950,000 tokens, roughly eight times the same hour’s audio.

At those rates a 45-minute recorded call is about 86,400 tokens as audio and about 710,100 as video, about 8.2 times as many. A token count is not a price, though. On some of the same provider’s models audio costs more per token than video, which narrows the gap in money to about two and a half times5; other providers count and charge differently, and prices fall over time. Treat the ratio as an order of magnitude, and price your own volume on the current rate card. The leadership question is not “can it watch the recording?” but “is the picture worth several times the cost of the sound?”

New ways to be wrong

Every added modality brings its own failures, and none of them is the failure of a person. Models invent detail in text, as Large Language Models showed6; seeing and hearing more does not cure that. It adds new forms of it.

Images. In 2024 researchers gave four leading vision-language models seven tasks a child can do: do two circles overlap, how many times do two lines cross, which letter is circled. The models averaged 58 percent; the best reached 78 percent. When the shapes were spaced further apart, accuracy rose to nearly 100 percent7. The lesson is not that the models are useless. It is that they fail on fine spatial detail, exactly the kind found in engineering drawings, charts and crowded screenshots, and that they answer confidently either way.

Speech. A FAccT 2024 study ran more than 13,000 recorded clips through a widely used speech-recognition model. About 1 percent of transcriptions contained whole phrases or sentences that nobody had said, and 38 percent of those inventions were harmful, such as violent language or false associations. They were more frequent for speakers with long pauses, as in aphasia8. One percent sounds small until the transcript becomes a medical note, a complaint record or the input to the next model in a pipeline.

Video. Video is not just many images. It adds sequence, timing and sound, and the model must work out what happened and in what order. On a 2025 benchmark of 900 videos, the best commercial model answered 82 percent of questions about clips under two minutes correctly, but 67 percent for videos of 30 to 60 minutes9.

The best commercial model answered 81.7 percent of questions on short videos, 74.3 percent on medium and 67.4 percent on long videos.Short videos (under2 minutes)81.7%Medium videos (4 to15 minutes)74.3%Long videos (30 to60 minutes)67.4%Source: Fu et al., Video-MME (CVPR), best commercial model · 2025
Figure 2.12.6 The longer the video, the more often the model gets the answer wrong.

Two practical rules follow. Test on your own real inputs, blurred, cropped, noisy and long, never only on the clean examples in a demo. And decide before launch what happens when the model is unsure: it hands over to a person.

Privacy rides along

Text that a person types usually contains what they chose to write. A screenshot contains whatever was on the screen: other customers’ names, salaries, contract values. A photo of a site shows faces and number plates. A call recording carries a voice, which can identify a person, and a video call carries a face and often a home. Every modality you add widens what the system ingests, stores and sends to its provider, which is why masking, retention and where the processing happens belong in the design, not in a review after launch.

Does the modality earn its place?

All of this reduces to one question, asked before any money is spent: does another input hold information that the text does not?

A decision tree - if no other input holds missing information, stay with text; if one does, test it on real inputs, pilot it with a human check if it passes, and do not scale it if it fails.Does another inputhold informationthe text lacks?NoStay with textCheaper and easier to testYesTest on real inputsPassesPilot with a person checkingFailsDo not scale yet
Figure 2.12.7 An extra modality must hold missing information and pass a test on real, messy inputs before it scales.

If no other input holds missing information, stay with text: it is cheaper, faster and easier to test. If one does, find the slice of cases where it matters, test on real examples from your own operation, and pilot with a person checking the output. Alongside the tree sit two questions that never go away: what in this input is sensitive, and what does each interaction cost at real volume? The decisive input is not always where a demonstration suggests, as the story shows.

Story: the demo that looked like live video

In December 2023 Google launched its Gemini model with a video called “Hands-on with Gemini”. In it, a person draws, gestures and plays games at a table while the model seems to watch and comment aloud as things happen. The description under the video noted that, for the purposes of the demo, latency had been reduced and the model’s outputs shortened for brevity10. A post on Google’s developer blog, published the same day, explained the prompting behind the video: the team gave the model still frames from the footage, together with written prompts11. Google told Bloomberg that the video had been made by using still image frames from the footage and prompting via text, as TechCrunch reported the next day12.

Google's December 2023 Gemini video looked like a model watching live video and talking; a footnote said latency was reduced and outputs shortened; the developer post showed still frames with written prompts; Google said the prompts and outputs were real.The videoLooks like live videoThe model seems to watchand talkThe footnoteLatency reducedOutputs shortenedfor brevityThe developer postStill framesplus textChosen inputs andwritten promptsThe responsePrompts andoutputs realMade to inspire developers
Figure 2.12.8 The outputs were real. The input was not the one the video suggested.

The difference lay in where the information lived. In the video, the model recognizes a game of rock, paper, scissors from silent hand gestures. In the written version, it was shown still images of the gestures and asked what the person was doing, with the hint that it was a game12. Part of what made the answer right was in the text, not in the pictures. And a handful of chosen frames is not video. It carries none of the sequence, timing and sound that, as the cost and failure sections showed, are where video’s accuracy falls and its token count climbs. Gemini’s co-lead responded that all the user prompts and outputs in the video were real, shortened for brevity, and that the video illustrated what multimodal experiences built with the model could look like and was made to inspire developers10. Nothing in the record says the model could not do what the video showed. The record says the video did not show how it was done.

For a leader the case is a template for reading any multimodal demonstration. Ask what actually went in: continuous video or chosen frames, sound or a transcript, and whether a written prompt carried part of the answer. Ask who chose the inputs, and how long each answer took. Then run the same task on your own real, messy inputs, live, at the speed and volume your operation needs. The question that settled this post-mortem is the question of the whole chapter: where does the information live?

What this means for leaders

Multimodal AI lets systems use a far richer picture of the business than text alone, and the most promising uses are those where a text-only process visibly loses information today: scanned documents with stamps and handwriting, photos from the field, screens that show what a customer actually did. These are capabilities, not proven results: whether one pays off in your process has to be shown there, in a pilot measured on your own inputs. The discipline is to treat each modality as a line item with a benefit, a cost, a failure profile and a privacy footprint, and to approve it on those terms. Approving “multimodal” as a label, because the system “sees and hears like a person”, is how organizations end up paying for video to transcribe speech.

Check yourself

  1. A model that can read images can also create them.
  2. Video is just a longer series of images.
  3. Multimodal models cannot invent details, because they can see or hear the evidence.
  4. An hour of video can count as several times the tokens of the same hour’s audio.
  5. A pipeline of specialist models is always worse than one multimodal model.
  6. Adding a modality can make a system costlier and riskier without improving its answers.

Reflection: where does your information live?

What comes next

This chapter kept one distinction throughout: understanding many kinds of information is different from generating them. Models that read a screenshot or a recording are also, increasingly, models that write the reply, draw the image or speak the answer. The next chapter, How Generative AI Works, explains how they create something new, and why the same request can produce a different answer every time.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

General Data Protection Regulation · EU

Regulation (EU) 2016/679

Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.

  • 2018-05-25 — Applies

Last verified 2026-10-08 · official text

References

  1. Tadas Baltrusaitis, Chaitanya Ahuja and Louis-Philippe Morency. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2) (arXiv:1705.09406). 2019.
  2. Alexey Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR 2021 (arXiv:2010.11929). 2021.
  3. Alec Radford et al. Learning Transferable Visual Models From Natural Language Supervision. ICML 2021 (arXiv:2103.00020). 2021.
  4. Google. Understand and count tokens (Gemini API documentation). Google AI for Developers. 2026.
  5. Google. Gemini Developer API pricing. Google AI for Developers. 2026.
  6. Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.
  7. Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri and Anh Totti Nguyen. Vision language models are blind. Asian Conference on Computer Vision (ACCV) 2024 (arXiv:2407.06581). 2024.
  8. Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann and Mona Sloane. Careless Whisper: Speech-to-Text Hallucination Harms. ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). 2024.
  9. Chaoyou Fu et al. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. CVPR 2025 (arXiv:2405.21075). 2025.
  10. The Decoder. Gemini co-lead Oriol Vinyals addresses criticism of Google Deepmind's staged multimodal demo. The Decoder. 2023.
  11. Alexander Chen. How it's Made: Interacting with Gemini through multimodal prompting. Google Developers Blog. 2023.
  12. Devin Coldewey. Google's best Gemini demo was faked. TechCrunch. 2023.

Further reading

Sources last verified 2026-10-10.