AI Academy · Book
Executives & Directors · Module 02 · Chapter 013

How Generative AI Works

Generative AI does not fetch a stored answer. It composes a new one, a token at a time for text or by refining random noise for images and video, and every run draws from probabilities. That is why the output varies, why it looks convincing whether or not it is true, and why the real cost moves from making content to choosing and checking it.

≈ 15 min read

After this chapter you can

  • Distinguish composing new output from retrieving stored content, and generative from predictive AI.
  • Explain how text is built one token at a time and how diffusion refines noise into images and video.
  • Explain why outputs vary from run to run and why generation settings should match the task.
  • Recognize why fluent or photorealistic output passes for accurate, and design facts to come from systems of record.
  • Recognize that generation shifts cost from producing content to choosing and checking it.

On the morning of 22 May 2023, a photograph of black smoke rising beside the Pentagon spread across Twitter. A verified account posing as a Bloomberg news feed shared it, and so did accounts with hundreds of thousands of followers1. Just after 10 a.m. in New York, the S&P 500 slipped about 0.3 percent to its low for the day. Within minutes the image was exposed as a fake and the index recovered2.

No camera had taken that picture, and no archive had stored it. It was widely judged to be the work of an image generator: a model that had learned what buildings, smoke and news photos tend to look like, and composed a new picture that matched a description. Nothing in the image was retrieved. All of it was made.

The same is true of every paragraph a chat assistant writes. Leaders who understand how that composing works can answer three questions their organizations face every week: where does the output come from, why does the same request give a different answer each time, and why does wrong output look exactly as convincing as right output?

Composed, not retrieved

A generative model produces new content in response to an instruction and whatever context the application supplies. What it holds is not a library of finished answers but patterns, learned in training, about how text, images, sound and code tend to fit together. At the moment of a request it uses those patterns to build an output through a long run of small choices, each one a matter of probability.

An instruction with context goes into a model of learned patterns, which makes many probabilistic choices to compose a new output - composed rather than retrieved, fluent rather than verified.Your requestInstruction pluscontextThe modelPatterns learnedin trainingSmall choicesEach one aweighted drawNew outputComposed forthis requestComposed, not retrieved. Fluent, not verified.
Figure 2.13.1 Generative systems broadly follow this chain. What differs is how the small choices are made.

This is what separates generative AI from the predictive systems described in The Major Types of AI. A predictive model returns a judgment about its input: a risk score, a category, a forecast. Its output is small, and the world eventually shows whether it was right. A generative model returns something new: a letter, a summary, an image, a block of code. There are many acceptable versions and no single correct one. The practical consequence is about checking. A score can be measured against outcomes. A composition has to be reviewed.

Two pairs of words are worth keeping apart from here on. Composed is not retrieved: the output was never stored anywhere, even when parts of it closely resemble something that was. And fluent is not verified: the model’s skill is producing output that fits the patterns, not checking it against the world.

Two ways to build an output

Modern generative systems compose in one of two main ways, and the difference explains much of their behavior.

Text is built from left to right. As Large Language Models showed, a language model predicts the next small piece of text, a token, adds it, and predicts again, so that a paragraph is hundreds of choices made in sequence3. Each choice is shaped by everything before it, including the choices the model has already made. An unlucky early choice therefore shapes every token after it.

Many image and video generators work differently. They use a method called diffusion, set out in its modern form by Jonathan Ho, Ajay Jain and Pieter Abbeel in 20204. In training, the model is shown images that have been progressively buried under random noise, and it learns to remove a little noise at a time. To generate, the process runs in reverse. It starts from pure static and, over many steps, removes noise in a direction that matches the prompt: first the rough layout, then shapes and light, then fine detail. The whole picture is worked on at once rather than piece by piece.

A table comparing text generation, which adds one token at a time from the instruction, with diffusion for images and video, which refines pure noise over many steps; text varies with each token drawn, images with the starting noise.Step by stepTextImages and videoStarts fromYour instruction and contextPure random noiseEach stepAdds one tokenRemoves a little noise everywhereNumber of stepsOne per token, often hundredsDozens to a thousand refinementsMain sourceof variationWhich token is drawnWhich noise it starts from
Figure 2.13.2 Text grows one piece at a time; diffusion refines a whole picture. Both draw on chance, in different places.

The original method used a thousand refinement steps for every image, which made it slow and expensive4. Later techniques produced images of similar quality 10 to 50 times faster5, which is a large part of why image generation became cheap enough for everyday use. Video applies the same idea across space and time. One leading developer describes its video model as a diffusion model that takes noisy patches of video and learns to predict the clean ones6. Newer systems increasingly blend the two approaches, so treat the table as the mental model rather than a rule about any product.

Every step is a weighted draw

Return to text and zoom into a single choice. Suppose an order assistant has written: “Your order is scheduled to arrive on”. The model does not know the arrival date. It assigns a probability to every possible next token.

Completing the phrase your order is scheduled to arrive on, an illustrative model rates Tuesday most likely at 38 percent and Wednesday close behind at 29 percent; the likeliest token is not necessarily the true date.…arrive on Tuesday38%…arrive on Wednesday29%…arrive on Monday14%…arrive on Thursday9%Something else10%ILLUSTRATIVE NUMBERS
Figure 2.13.3 Illustrative. The model rates every candidate; the likeliest is not necessarily the true one.

Two things follow from an illustrative chart like this one. First, the favorite is only the most likely continuation given the patterns. It is not the arrival date, which lives in the order system. Second, the system then has to choose, and in most applications it does not simply take the top bar. It makes a weighted draw.

There is a good reason for that. Ari Holtzman and colleagues showed that always choosing the most likely text “leads to text that is bland and strangely repetitive”; models that pick only the favorite tend to loop. Their remedy, now widely used, is to draw from the likely core of the distribution and ignore the long tail of improbable options7. The result reads far better. The price is that the same request, on the same model, with the same settings, can come back in different words or with different content.

Even with sampling turned down to its most predictable setting, outputs can still differ between runs, because of how servers batch requests together. AI vs Automation showed one test in which 1,000 identical requests gave 80 different answers, until engineers changed the serving software8. Variation is a property of how these systems are built and run. It can be reduced. It should not be a surprise.

Match the settings to the task

The most common control is a setting called temperature. Lower values concentrate the draws on the likeliest options, so output is more conservative and repeatable. Higher values spread them out, so output is more varied and sometimes more surprising. Image generators have an equivalent: the random starting noise, often called a seed. Keep the seed and the settings fixed and the image can usually be reproduced; change the seed and you get a new image.

A spectrum from predictable settings for extraction and fixed formats to varied settings for brainstorming and concepts, with the marker saying set it per task; neither end makes output true.PredictableExtraction, classification, fixed formatsVariedBrainstorming, concepts, alternativesSet per taskNeither end makes output true
Figure 2.13.4 Choose the setting for the job, then remember that no setting turns a likely answer into a checked one.

The business rule is simple: the setting should match the task. Pulling fields from an invoice wants predictable output in a strict format. Generating fifty names for a new service wants variety. Teams get into trouble when one configuration is used for everything, or when a low temperature is treated as a guarantee. It is not. A lower setting makes the model say its likeliest answer more consistently. If the likeliest answer is wrong, it will now be wrong consistently. Where the words must be identical every time, such as a legal disclosure or a safety instruction, the reliable design is a fixed template that the model does not rewrite.

From noise to picture, and why it looks real

Diffusion explains a property of generated images and video that executives need to understand: they are refined toward what looks typical, not toward what exists. Each step makes the picture more consistent with the patterns the model learned and with the prompt. Nothing in the process consults a map, a floor plan or a photo archive. That is why a generator can produce a convincing photograph of an event that never happened, as it did beside the Pentagon, and why it can render a product, a building or a face in confident detail that matches nothing real.

Diffusion starts from random static, forms a rough layout, refines light and detail, and ends with a finished image after many steps; new starting noise gives a new image, and plausible is not the same as real.Pure noiseRandom staticRough layoutShapes emergeRefineLight, texture,detailFinished imageAfter manysmall stepsNew noise, new image. Plausible is not real.
Figure 2.13.5 Diffusion removes noise step by step toward the prompt. Each step moves toward what looks typical, not toward what is true.

The same mechanism produces the characteristic failures. Physical rules are learned only as visual habits, so they break in the details. The developer’s own report on that video model noted that it “does not accurately model the physics of many basic interactions, like glass shattering”6. Hands with extra fingers, wheels that do not turn and lettering that is almost words are the same phenomenon: locally plausible, globally wrong.

Because generated media can be mistaken for real records, the law now treats it as something to be disclosed.

Fluent is not verified

For text, the equivalent of the convincing fake photograph is the confident wrong sentence. A model produces the likeliest continuation, so an invented figure, deadline or citation arrives in the same polished prose as a correct one. Researchers call fluent but false or unsupported output hallucination, and a widely cited survey of the field treats it as a general property of language generation, not a defect of one product9. Large Language Models showed why training and benchmarks reward confident guessing. The question here is why people so readily believe the result.

The answer is that fluency is what readers use as a shortcut for credibility, and generation is very good at fluency. In a study published in Science Advances, Giovanni Spitale and colleagues found that people could not tell tweets written by an AI language model from tweets written by people. The model’s accurate tweets were easier to understand than the human ones, and its false tweets were more compelling10.

Fluency sits above the waterline; where facts came from, whether figures were calculated, whether the content is current and whether anyone checked it all stay hidden below.WHAT ANY READER SEESFluent sentences · Confident tone ·Tidy formattingWHAT THE TEXT CANNOT SHOWWhere each fact came fromWhether a figure was calculatedWhether it is currentWhether anyone checked it
Figure 2.13.6 Fluency is visible and costs the model nothing. Correctness is invisible until someone checks.

The remedy is design, not exhortation. In a well-built application the facts come from authoritative sources at the moment of the request, such as the order system, the policy store or a deterministic calculation, and the model writes the words around them. Supplying documents at question time is called retrieval-augmented generation11, and RAG and Enterprise Knowledge covers how it works. Grounding reduces invented facts. It does not remove the need to check what matters.

Story: seventy thousand clips for one commercial

In 1995 Coca-Cola first aired “Holidays Are Coming”, its commercial of illuminated red trucks rolling through snowy towns, and it became one of the brand’s best-known films. In November 2024 the company released two AI-generated reimaginations of it, made with specialist AI studios. The music was performed by real artists12. Marketing Week reported that the 2024 version was produced for about a tenth of the cost of a traditional shoot13. Critics were less impressed by what they saw: stiff, uncanny human faces, and truck wheels that turned in odd directions or not at all14.

A timeline from the 1995 original, through the 2024 AI remake criticized for uncanny faces and wheels, to about a month of generating more than 70,000 clips in 2025, then selection and repair, and the 2025 remake directed and finished by people.1995The originalfilmRed trucks,snowy townsNov 2024First AI remakeUncanny faces,wheels that donot turn2025About a monthof generationMore than70,000 clipsThenSelection andrepairAnimals insteadof facesNov 2025Second AIremakePeople directed,chose and finished
Figure 2.13.7 Generation made the clips cheap. Choosing, repairing and directing them was where the work went.

For the 2025 version, the Wall Street Journal reported, five AI specialists generated and refined more than 70,000 video clips over about a month, and around 100 people were involved in the campaign as a whole14. That is roughly 470 clips per specialist for every day of that month, for films that use a small fraction of them. The new films replaced human faces with animals. Viewers still spotted Santa’s fingers briefly distorting into impossible shapes14. Coca-Cola’s head of generative AI put the division of labor plainly: creative direction “has and always will be human-led”, and AI is the engine for execution and production13.

More than 70,000 clips were generated by five AI specialists in about a month, with around 100 people involved in the campaign.70,000+Clips generatedFor one holiday campaign5AI specialistsAbout a month of work~100People involvedAcross the whole campaignSource: Futurism, citing the Wall Street Journal · Nov 2025
Figure 2.13.8 The clips were cheap and plentiful. The scarce resource was people’s judgment about which ones to keep.

Read the case through the mechanics of this chapter. Every clip was a draw: the same prompt, run again, gave a different truck, a different snowfall, a different polar bear, so the team generated tens of thousands and kept the few that worked. Every clip was plausible rather than correct: the wheels and fingers failed because the model had learned how trucks and hands usually look, not how they work. And the cost moved. Coca-Cola said its 2024 remake cost about a tenth of a traditional shoot; the 2025 clip counts show where the effort went instead. The work that remained was deciding what the brief should be, which samples were good enough, what had to be repaired by hand, and whether the result was worthy of the brand. Whether the films succeeded is a matter of taste and of reaction, which was divided. What the case settles is where generative AI’s costs go: away from making, toward choosing and checking.

What this means for leaders

The mechanics translate into a short set of design decisions. Start by asking of any proposed use whether it needs a composition or a fact. Drafting, rewriting, summarizing, translating and visual concepts are compositions, and generation is well suited to them, although the gain in any one process has to be measured there. Dates, prices, balances, eligibility and anything a regulator or customer will rely on are facts, and they should come from systems of record and deterministic software, with the model writing the words around them.

Then decide how much variation the task should have, and test for it before launch. Run the same requests many times and look at the spread. Where the answer must be identical, use templates and fixed content, not a lower setting and hope. Where variation is the point, plan for selection: someone has to choose among the samples, and that someone is part of the cost.

Finally, measure cost per useful outcome rather than cost per output. Suppose, with illustrative figures, that 100 drafts cost 1,000 in total and 70 of them are usable. The figure that matters is about 14.29 per usable draft, not 10 per draft, plus the time of the person who sorted them. As the commercial showed, generation makes the first draft nearly free and makes judgment the scarce input.

Check yourself

  1. A generative AI system finds a stored answer and sends it back.
  2. The same prompt with the same settings can produce a different answer.
  3. Turning the temperature down makes an answer more likely to be correct.
  4. Many image and video generators start from random noise and refine it toward the prompt.
  5. A fluent, confident answer is more likely to have been checked.
  6. Generation makes producing content cheap, so the main cost shifts to choosing and checking it.

Reflection: compositions and facts

What comes next

You now have the mechanics: generative AI composes, it draws on chance at every step, and its fluency says nothing about its accuracy. Of everything that shapes the output, the instruction and the context you supply are where you have the most control. The next chapter, Prompting and Context Engineering — Executive Mental Model, explains how to communicate intent to these systems, and why prompting has become much more than writing a good sentence.

Laws referenced

EU AI Act · EU

Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744

Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.

  • 2024-08-01 — Entered into force
  • 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
  • 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
  • 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
  • 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
  • 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
  • 2028-08-02 — High-risk obligations for AI in products regulated under Annex I

Last verified 2026-10-06 · official text

References

  1. Amanda Silberling. Fake Pentagon attack hoax shows perils of Twitter's paid verification. TechCrunch. 2023.
  2. Bloomberg News. Fake AI Photo of Pentagon Blast Goes Viral and Trips Stocks Briefly. Bloomberg. 2023.
  3. Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
  4. Jonathan Ho, Ajay Jain and Pieter Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS 2020 (arXiv:2006.11239). 2020.
  5. Jiaming Song, Chenlin Meng and Stefano Ermon. Denoising Diffusion Implicit Models. ICLR 2021 (arXiv:2010.02502). 2021.
  6. OpenAI. Video generation models as world simulators. OpenAI technical report. 2024.
  7. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes and Yejin Choi. The Curious Case of Neural Text Degeneration. ICLR 2020 (arXiv:1904.09751). 2020.
  8. Horace He and Thinking Machines Lab. Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. 2025.
  9. Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.
  10. Giovanni Spitale, Nikola Biller-Andorno and Federico Germani. AI model GPT-3 (dis)informs us better than humans. Science Advances 9, eadh1850 (arXiv:2301.11924). 2023.
  11. Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
  12. The Coca-Cola Company. Coca-Cola Refreshes Givers of the Season, Embraces AI-Powered Storytelling in Global Holiday Campaign. The Coca-Cola Company. 2024.
  13. Marketing Week. Coca-Cola launches new AI version of 'Holidays Are Coming'. Marketing Week. 2025.
  14. Frank Landymore. Coke's New AI-Generated Ad Required 100 Staff and 70,000 AI-Generated Clips, and It Still Looks Like Garbage. Futurism. 2025.

Further reading

Sources last verified 2026-10-08.