AI Academy · Book
Executives & Directors · Module 02 · Chapter 014

Prompting and Context Engineering — Executive Mental Model

A prompt tells a model what to do; the context decides what it has to work with. Many wrong answers in production come from what the model could not see rather than from how the request was worded. Leaders who ask "what did the model receive?" before approving a better prompt fix the right problem.

≈ 14 min read

After this chapter you can

  • Explain what a prompt is and which parts make it a good brief, including what a role can and cannot do.
  • Distinguish prompting from training and fine-tuning.
  • Describe context engineering as the design of everything a model receives when it runs, within a limited budget.
  • Recognize the limits of prompting - trust boundaries, rules that software must enforce, and evaluation of every change.
  • Diagnose when a longer prompt is the wrong fix for wrong answers.

In 2022 a team of researchers gave one of the largest language models of its day a set of grade-school math word problems. First they showed it eight example problems, each with only the question and the final answer. It solved about 18 percent of the new problems. Then they changed nothing but the prompt. They used the same eight examples, with the reasoning in each written out step by step. It was the same model with the same parameters and no retraining, and it now solved about 57 percent, more than three times as many1.

Newer models are trained to reason this way by default, as Large Language Models explained, so that particular trick matters less in 2026. The lesson behind it has grown more important. What a model produces depends heavily on what it receives. Two teams can license the same model and build applications of very different quality, and the difference is rarely a magic phrase. It is the design of everything that reaches the model each time it is called.

The core idea

Prompting tells the model what to do. Context engineering decides what the model has available to do it with.

The practice has moved through three stages. At first people simply asked: draft me an email. Then came prompt engineering, the craft of writing a clear brief with a role, a task, constraints and examples. By mid-2025 the people building serious applications had started to use a broader term. The researcher Andrej Karpathy described context engineering as “the delicate art and science of filling the context window” with exactly the information the next step needs2. The wording of the request became one part of a larger design problem.

Three stages from asking a model, to engineering the prompt, to engineering everything the model receives; the model is one part of the system.Ask'Draft me an email.'Prompt engineeringRole, task, constraints, examplesContext engineeringEverything the model receives
Figure 2.14.1 The wording is one rung. In production, the serious work is designing everything the model receives.

For an executive the shift matters because it moves the question. “Who writes our prompts?” is a small question. “Who decides what our AI systems can see, and how do we know it is current and correct?” is a management question, and it has owners, budgets and controls attached.

A good prompt is a brief, not a magic phrase

A prompt is the instructions, questions, examples and material given to a model when it is used. Compare two requests. The weak one says: look at this report. The stronger one says: you are helping a chief operating officer prepare for a board meeting; identify the three most important operational risks in this report; use only the report, and if it does not support a claim, say so; give a five-line summary, then the risks as a numbered list.

A good prompt can contain six parts - role, task, context, constraints, examples and output format - chosen to fit the task.RoleFrames the workTaskExactly what to doContextFacts the task needsConstraintsLimits, sources, no guessingExamplesShow the patternOutput formatCheckable by software
Figure 2.14.2 Six parts of a good brief. Use the ones the task needs; not every prompt needs all six.

One part deserves a warning. A role frames the work, but it grants no expertise. Researchers tested 162 personas on 2,410 factual questions across four families of models; adding a persona did not improve accuracy, and choosing the best persona automatically did no better than chance3. The other parts do plainer jobs. Constraints tell the model to say when it does not know, which reduces unsupported answers without eliminating them; examples show a pattern that large models can pick up without retraining4; a defined format gives the next piece of software something to check. None of this is a secret vocabulary. It is the brief you would give a capable new colleague.

Prompting changes the input, not the model

One distinction is blurred in many vendor meetings, and it is worth keeping clean.

Prompting changes the input and leaves the model unchanged; training or fine-tuning changes the parameters and creates a new model to test and govern.PROMPTINGChanges what the modelis given.Same model - quick to change and roll backTRAINING OR FINE-TUNINGChanges themodel's parameters.A new model to test and governvs
Figure 2.14.3 A prompt lends information to one request. Training changes the model itself.

Prompting changes what the model is given. Nothing inside the model moves, so a prompt can be changed, tested and rolled back in hours. Training or fine-tuning changes the model’s parameters. The result is a new model that needs data, compute, testing and governance. Today’s assistants follow instructions well because their developers post-trained them on human feedback long before anyone typed a request, as Large Language Models described5.

So when a team says “we taught it our policy in the prompt”, the policy sat in the input for that conversation. Unless the application stores it and sends it again, it is gone afterwards. There is a third route between the two: retrieval, which fetches the relevant passage from your own sources and places it in the context for each request6. The next chapter, RAG and Enterprise Knowledge, explains why organizations usually retrieve rather than retrain.

Context engineering designs everything the model receives

Context engineering is the design of everything a model receives each time it is called. One practitioner guide defines it as curating the best set of information for the model at the moment of use, including everything that arrives outside the prompt itself7.

The model's context is assembled for each request from instructions, the request, the conversation, knowledge, business data, tool results and memory.Contextassembled foreach requestInstructionsThe requestConversationKnowledgeBusiness dataTool resultsMemory
Figure 2.14.4 The user types one line. The application assembles the rest before the model is called.

Take one question, as an illustration. A site engineer at a construction firm asks an assistant: can we pour the slab on Thursday? A well-built assistant does not pass that sentence straight to the model. It assembles the request with the engineer’s site and role, the current cold-weather clause of the concrete specification, the forecast from a weather service, the pour schedule from the project system and the earlier messages in the conversation. Only then does it call the model. A poorly built one sends the question alone, and the model answers fluently from general patterns about concrete, which is exactly what you do not want.

The difference shows up in production numbers, not only in examples. LinkedIn’s engineers changed what their customer service assistant retrieved: instead of passing past support tickets to the model as loose chunks of text, they kept each ticket’s structure and links to related issues. Retrieval quality rose by 77.6 percent against the earlier approach, and after about six months in use by the customer service team, the median time to resolve an issue had fallen by 28.6 percent8. The model did not change; what it was given did. It is one company’s own report, but it shows the size of the stakes.

The wording of the instructions is now one part of the solution. The rest is design: which documents to fetch, which of the user’s data to include, which tools to call, what history to keep and what to leave out. Tokens, Context and Embeddings showed what the context window is; the design question is what goes into it.

Context is a budget: relevance beats volume

Context is limited, so everything sent competes for space, and every page adds cost and delay. The same guide puts it plainly: context is “a finite resource with diminishing marginal returns”7. Tokens, Context and Embeddings showed the evidence that models use information buried in the middle of a long input less reliably than information at its start or end9.

Narrowing from every project document to the cold-weather clause gives the model less but more relevant information, which usually improves quality and lowers cost.4,000 pagesEvery project document60 pagesConcrete specification2 pagesCold-weather clauseSent with the request
Figure 2.14.5 Less, but more relevant. The narrow context is usually the better and the cheaper answer.

For the engineer’s question, the assistant could send every project document, the whole concrete specification or the two-page cold-weather clause with the request. The third is usually best: less information, but the right information. Irrelevant pages can crowd out the ones that matter, lower quality and raise the bill. Context engineering is also cost engineering. Think of it as packing for a trip: a bigger suitcase does not make a better journey, and you pack for this trip rather than every trip you might take.

Three further design choices follow. Keep instructions, evidence and the task clearly separated, so the model can tell what it is being asked from what it is being shown. Summarize long conversations with care, because a summary can drop the one detail that mattered. And remember that an assistant which seems to remember you is being sent earlier messages or stored notes as context. That memory often holds personal information, so it needs a purpose, a retention limit and an owner.

Not everything in the context is an instruction

Context engineering is also a control problem, in three ways.

Instructions come from organization rules, developer instructions and the authorized user; retrieved documents, emails, web pages and tool output are data, never orders.OrganizationrulesEnforced in software tooTrustedDeveloperinstructionsShape the behaviorTrustedUser requestAuthorized data onlyCheckedRetrievedcontentDocuments, emails, web pages, tool outputNever obeyedTRUSTRISES
Figure 2.14.6 Instructions flow from the top. Everything retrieved is evidence to read, not an order to follow.

First, rules. Suppose the specification in the context says no pour below a set temperature without heated enclosures. The model can use that rule when it answers. But for anything that costs money, safety or trust if broken, the rule must also be enforced in ordinary software, for example in the system that approves the pour. Never make the model the only control.

Second, permissions. Each user should reach only the data they are authorized to see, and that check happens before anything enters the context, not after the model has read it.

Third, trust. Instructions come from the top of the stack: the organization’s rules, the developer’s instructions, then the user’s request. Retrieved documents, emails, web pages and tool results sit at the bottom. They are evidence to read, never orders to obey. A document that contains the words “ignore your instructions” is a prompt-injection attack, which no wording alone reliably stops10; Security and AI Attacks in Module 6 covers the defenses.

Treat every prompt change like a software release

A prompt should never be called good because one example worked. Small, apparently meaningless changes can move results a long way. In one study, changing only the formatting of a few-shot prompt, such as separators, spacing and capitalization, shifted one open model’s accuracy on a task by up to 76 percentage points, and the sensitivity did not disappear with larger models or more examples11.

Every prompt change is managed like code - versioned, evaluated on the same set of real questions, released with an owner and rollback, and traced in use.VersionOne numberedchangeEvaluateSame set ofreal questionsReleaseNamed ownerand rollbackTraceModel, prompt,sources, toolsEvery promptchangeManaged like code
Figure 2.14.7 Prompts are application logic. One changed sentence changes the application for every user at once.

Prompts are application logic, so manage them as code. Make each change a numbered version. Evaluate it by running version two against the same representative set of real questions as version one and comparing the results. Release it with a named owner and a way to roll back. Then trace it in use: which model and prompt version answered, what was retrieved, which tools were called, and the cost and quality signals, without logging sensitive data you do not need. That trace answers the question every executive eventually asks: why did the AI say that?

The discipline also exposes the most common trap. If answers are wrong because knowledge is missing or outdated, retrieval finds the wrong document, the user’s data is absent or a tool is unreachable, a longer prompt fixes none of it. The following case shows what that looks like in public.

Story: the support agent that invented a policy

In April 2025, developers using Cursor, an AI coding editor, found themselves logged out whenever they switched between machines. One emailed support and received a reply from “Sam”. The logouts were expected, Sam explained: the product was “designed to work with one device per subscription as a core security feature”12.

In April 2025 Cursor users were logged out across machines, an AI support agent explained it with a policy that did not exist, users cancelled, and the company apologized and labeled AI replies.Mid-April 2025Unexpected logoutsUsers switching machinesare signed outSupport reply"Sam" explainsa policyOne deviceper subscriptionDays laterPublic backlashPosts spread; someusers cancelResponse"We have nosuch policy"Bug fixed, refund, AIreplies labeled
Figure 2.14.8 The policy never existed. The agent filled a gap in what it knew with something plausible.

There was no such policy. Sam was an AI agent, and its replies were not labeled as AI. The answer spread across Reddit and Hacker News, and some users announced they were cancelling their subscriptions12. Michael Truell, a co-founder, replied that it was “an incorrect response from a front-line AI support bot”. The company had rolled out a change to session security; the logouts turned out to come from a race condition on slow connections, since fixed. The affected user was refunded, and the company said that AI is its first filter for email support and that AI replies would now be clearly labeled13.

Cursor has not published how its support agent was built, so read what follows as the diagnosis the public record supports, not an inside account. The question the agent received was about behavior caused by a recent engineering change. Nothing in what it could see appears to have explained that change, so it had no facts about the cause. A language model with no relevant facts does not fall silent; it composes the most plausible answer, and a security policy is a plausible reason for a logout. There was no visible path for “I don’t know, let me pass you to a person”, and no label that would have told the user to treat the reply with care.

Users saw a confident policy answer; below it, the agent lacked facts about the recent change, a hand-off rule, a policy check and an AI label.WHAT USERS SAWA confident answer about a policyWHAT THE CONTEXT LACKEDFacts about the recent changeA rule to say 'I don't know' and hand offA check on stated policiesA label saying it was AI
Figure 2.14.9 The failure sat below the wording, in what the agent could see and what nothing checked.

Read through this chapter’s lens, each gap is a context or control decision, not a wording one. Current facts about known incidents belong in the context of a support agent. A constraint to say “I don’t know” and route to a person belongs in its instructions, and the routing belongs in software. Any statement of company policy can be checked against the published policy before it is sent. And labeling AI replies, which the company did, tells customers how much weight to give them. A cleverer sentence in the prompt would not have supplied any of these.

What this means for leaders

The mental model fits in one question: what could the model actually see? Most of the useful leadership work follows from asking it. Treat context as a designed product with an owner, not a by-product of whoever wrote the first prompt. Insist that critical rules live in software and that retrieved content is never treated as an instruction. Budget for evaluation sets and tracing as part of the application, not as optional extras. And when a team proposes a longer prompt to fix wrong answers, ask for the diagnosis first.

Check yourself

  1. A sufficiently good prompt can make a weak model perform any task.
  2. Putting our policy in the prompt trains the model on it.
  3. Telling the model it is a senior expert reliably makes its answers more accurate.
  4. More context always improves the answer.
  5. Changing only the formatting of a prompt can change accuracy dramatically.
  6. A document retrieved into the context should be treated as evidence, not as instructions.

Reflection: what does your assistant see?

What comes next

You now have the mental model: a prompt is a brief, the context is a designed product, and the model can only be as good as what it receives. For an enterprise, the hardest part of that design is the knowledge itself: thousands of documents, constantly changing, each visible only to some people. The next chapter, RAG and Enterprise Knowledge — Executive Mental Model, explains how retrieval connects models to what an organization knows, and why that is usually the better path than retraining.

References

  1. Jason Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022 (arXiv:2201.11903). 2022.
  2. Andrej Karpathy. +1 for "context engineering" over "prompt engineering". X (post). 2025.
  3. Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee and David Jurgens. When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models. Findings of the Association for Computational Linguistics, EMNLP 2024. 2024.
  4. Tom B. Brown et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 2020.
  5. Long Ouyang et al. Training language models to follow instructions with human feedback. NeurIPS 2022 (arXiv:2203.02155). 2022.
  6. Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
  7. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan and Jeremy Hadfield. Effective context engineering for AI agents. Anthropic Engineering. 2025.
  8. Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang and Zheng Li. Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering. Proceedings of SIGIR 2024 (ACM), pp. 2905-2909. 2024.
  9. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12. 2024.
  10. Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
  11. Melanie Sclar, Yejin Choi, Yulia Tsvetkov and Alane Suhr. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. ICLR 2024 (arXiv:2310.11324). 2024.
  12. Benj Edwards. Company apologizes after AI support agent invents policy that causes user uproar. Ars Technica. 2025.
  13. Thomas Claburn. Cursor AI's own support bot hallucinated its usage policy. The Register. 2025.

Further reading

Sources last verified 2026-10-08.