Tokens, Context and Embeddings
Tokens are the pieces of text a model reads and the meter you pay by. Context is everything the model can see in one request, and a bigger window does not mean better use of it. Embeddings are learned numbers that place content by meaning, so you can find the few passages worth sending. Leaders who keep the three apart are better placed to design assistants that are cheaper, faster and more accurate.
After this chapter you can
- Explain what a token is and why token counts, which vary by language, drive cost and speed.
- Describe what fills a context window and why a bigger window does not guarantee better use of it.
- Distinguish tokens from embeddings and explain how embeddings compare meaning.
- Name the three limits of meaning search: exact identifiers, domain fit and the sensitivity of vectors.
In 2020, one of the best-known language models could take in 2,048 tokens at once, about a few pages of text1. Four years later, a technical report described models that could recall and reason over millions of tokens in a single request, with near-perfect scores on tests that hide one fact in an enormous input2. On that evidence, the obvious design for an enterprise assistant is to give it everything: every manual, every contract, every email thread.
Then, in 2025, researchers changed the test in one small way. They wrote questions that shared almost no words with the hidden fact, so a model had to recognize the meaning rather than spot a matching phrase. At 32,000 tokens, a modest fraction of those advertised windows, 11 of the 13 models they tested fell below half of their own short-input scores. The best model fell from 99.3 percent to 69.7 percent3.
Both findings are true. Models can accept far more text than ever, and filling that space can make their answers worse. Explaining why, and deciding what to do about it, takes three ideas that are often confused with each other: tokens, context and embeddings.
Three words, three jobs
The three terms travel together in most conversations about AI systems, and they name three different things.
A token is a unit of text. An embedding is a list of numbers that represents meaning. One is counted, the other is compared. Context sits between them: it is the space into which the chosen text is placed, measured in tokens and filled, in a well-built system, with the help of embeddings. The rule that ties them together is simple to state and hard to follow: give the model the right information, not all information.
Tokens: what the model reads and what you pay for
A language model does not see a page the way a person does. Before the model reads anything, a component called a tokenizer cuts the text into pieces from a fixed vocabulary. A token can be a whole common word, part of a longer word, a punctuation mark or a space. A word such as unfiltered may be split into several pieces; a code like a lot number may be split almost character by character. Each token is then mapped to a number, and it is these numbers the model works on. It writes its answer the same way, one token at a time, as Large Language Models showed.
Tokens matter to leaders for a plain reason: they are the meter. AI services commonly bill by the tokens sent and the tokens returned, and the time a request takes rises with both. A longer prompt costs more to read, and a longer answer takes longer to write, because the model produces it token by token.
Token counts also vary more than most budgets assume. Tokenizers are built mostly from English-heavy text, so the same message can need many more tokens in another language. Petrov and colleagues found that the same text translated into different languages can differ in token length by up to 15 times4. For a company that works in many languages, that difference carries through to cost, response time and how much fits in the window. The most reliable forecast is one made on your own documents, in your own languages.
One source of confusion is worth clearing up here. Inside a language model, each token number is immediately turned into a learned vector, and the original Transformer paper calls these vectors embeddings5. Those internal vectors are part of the model’s machinery. When an architecture diagram shows “embeddings” next to a search index, it usually means something else: one vector per question or passage, produced by a separate embedding model so that content can be compared. That second kind is the subject of later sections.
Context: everything the model sees in one request
For each request, a model receives a particular set of text, and that set is its context. The user sees only the tip: the question they typed. Underneath, the same request usually carries much more.
The context window is the limit on how many tokens a model can take into account at once. Everything in the request competes for that space. When a conversation or a document outgrows it, something must be dropped or summarized, usually by the application rather than the model.
Context is not training. What a model learned is held in its parameters and stopped changing when training ended; context is supplied at the moment of use and is gone when the answer is written. Putting a document into a request lends it to the model for one answer and changes nothing inside the model. An assistant that seems to remember you is storing notes and sending the relevant ones back as context, which Prompting and Context Engineering covers along with how to design what goes in.
Why a bigger window is not a better answer
If the window is large enough to hold everything, why not send everything? Because a model does not use every part of its context equally well.
The best-known evidence is a 2024 study titled Lost in the Middle. Liu and colleagues placed the document holding an answer at different positions among many others. Models did best when the answer sat at the start or the end of the input and worst when it sat in the middle. One model’s accuracy dropped by more than 20 percent; in the worst case, with the answer buried among 20 or 30 documents, it scored lower than it did with no documents at all6.
Later work points the same way. The NoLiMa results in the opening of this chapter show the effect sharpening when question and answer use different words, which is the normal case in business, where a customer’s phrasing rarely matches the policy’s3. A 2025 technical report tested 18 models and found that performance grew increasingly unreliable as inputs lengthened, even on simple tasks such as repeating text back7.
The two headline studies measure different things. Lost in the Middle compares one model’s accuracy as the answer moves to different positions in the same long input. NoLiMa compares each model’s long-input score with its own short-input score. Both point the same way, but their percentages are not interchangeable.
So the case against “send everything” has four parts. Every extra token is billed. Every extra token adds processing time. The passage that matters can be buried and used less reliably. And everything in the window is visible to the model and potentially to the user, whether or not that user should see it. Too little context has its own failure: the model fills the gap with something plausible. The target is in between, relevant context in a sensible order.
A long window still has real uses. Reading one long contract end to end, or comparing two full versions of a specification, are tasks where the whole document is the relevant context. The window is headroom for those cases, not a design for all of them.
Embeddings: meaning as coordinates
If the goal is to send only the passages that matter, something has to find them. That is the job of embeddings.
An embedding is a long list of numbers, often hundreds or thousands of them, that an embedding model has learned to produce for a piece of content. The numbers act as coordinates. Content with similar meaning lands close together, and unrelated content lands far apart, even when the words overlap. The idea is older than today’s language models. In 2013, Mikolov and colleagues showed that learned word vectors captured relationships as directions: take the vector for king, subtract man, add woman, and you land very close to queen8.
A useful picture is a map. On a road map, two places are close because of where they are, not because their names sound alike. An embedding is an address on a map of meaning. “Can we print barrel-aged on the label?” and “Is matured in oak casks allowed?” share almost no words, yet they are neighbors. “Return the empty barrels” shares a word with the first and lands somewhere else entirely.
The map picture breaks in one useful place. A road map has two dimensions you can read; a map of meaning has hundreds, and no single number in an embedding means anything a person could name. You cannot inspect a vector and see “about oak barrels”. You can only compare it with others.
Because comparison is the point, the same technique serves many purposes beyond feeding a language model: search, grouping similar complaints, recommending products and spotting duplicate records. It also reaches beyond text to images, sound and products, which Multimodal AI takes up. Using embeddings to fetch your own documents at question time is the core of retrieval-augmented generation, which RAG and Enterprise Knowledge covers in full, including why retrieval must apply each user’s permissions before it looks for similar content.
Where meaning search falls short
Embeddings are powerful and easy to oversell. Two limits belong in every design review.
First, some things must match exactly. A lot number, a certificate reference, an error code or a clause number is not “similar” to its neighbor; it is right or wrong. Meaning search can rank a near-miss above the exact identifier. Keyword search remains strong: across 18 benchmark datasets, the long-established keyword method BM25 proved a robust baseline, and embedding-based retrievers often did worse on data unlike what they were trained on9. Strong systems therefore combine meaning search, exact lookup and filters. No embedding model is best at everything, either: a benchmark of 33 models found that no single one dominated across tasks, so the test that matters most is a set of your own real questions with known right answers10.
Second, embeddings are not anonymous. It is tempting to treat a list of numbers as harmless, but researchers have recovered 92 percent of short, 32-token texts exactly from their embeddings, and recovered full names from a dataset of clinical notes11. OWASP’s 2025 list of the top risks for applications built on language models includes weaknesses in vectors and embeddings as a category of its own12. Store and protect vectors with the same access rules as the documents they came from.
Story: a label desk stops sending the archive
Suppose a mid-sized wine importer with a label-compliance desk. Before each new wine ships, its compliance officers check the label and the paperwork against the labeling rules of the target market and against the firm’s own record of past decisions, a task that turns on precise wording. The firm builds an AI assistant to suggest whether a label passes and explain why, with an officer signing off every label.
The first design is generous. Every request carries the producer’s whole file, including technical sheets and certificates in French, Italian, Spanish and Portuguese, the full labeling rules for the market and the desk’s entire archive of past label decisions. The window is large enough for most of it, and the team trims the rest from the end.
Within weeks, three symptoms appear. The bill runs several times over plan, because every question pays for the full archive, and documents in some languages cost far more tokens than the English ones. Officers wait so long for answers that many stop asking. And the suggestions miss things an officer would catch. A draft label says “barrel-aged”; the producer’s sheet says the wine was “matured in oak casks”. The two phrases share almost no words. The past decision that settled exactly this kind of claim sits halfway down a long archive, and the assistant does not use it. Worse, an account manager working for one producer can see prices and notes written about a competing producer’s wines, because everything goes to everyone.
The team does not change the model. It changes what the model sees. Notes are filtered by producer account first. Labeling rules and past decisions are indexed as embeddings, so “barrel-aged” finds the rule on oak-aging claims and the earlier decision on a similar wine, while an exact lookup handles lot numbers and certificate references character for character. Each request now carries about a dozen passages instead of an archive. Every week the desk reruns a set of past labels whose correct decision is known and checks that the right passages come back.
The result, in this illustration, is the pattern the research predicts: fewer tokens per request, so a lower bill and faster answers, and the decisive precedent near the top of a short context instead of lost in the middle of a long one. Nobody on the first team was careless. They filled the window because they could.
What this means for leaders
The practical lessons are few and durable. Treat tokens as a unit cost, and ask for cost per answer, measured on your own documents and languages, before you approve a budget. Treat a larger context window as headroom, not as a reason to stop choosing what goes in. Expect good systems to find content both by meaning and by exact match, and expect the team to prove retrieval quality on a set of your own questions. And treat embeddings as sensitive data, governed like the documents they came from.
Check yourself
- Tokens and embeddings are two names for the same thing.
- The same text can need very different numbers of tokens in different languages.
- If the context window is big enough, sending everything gives the best answers.
- Two questions with almost no words in common can have close embeddings.
- Embedding-based search makes keyword search obsolete.
- Embeddings are just numbers, so they need no protection.
Reflection: what does one request carry?
What comes next
Everything in this chapter has been text: text cut into tokens, text placed in context, text turned into coordinates of meaning. Much of what organizations handle is not text at all. It is a photo of a damaged pallet, a scanned delivery note, a recorded call. The next chapter, Multimodal AI, shows how the same ideas stretch across images, audio and other kinds of input, and where they break.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
References
- Tom B. Brown et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 2020.
- Gemini Team, Google. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv technical report 2403.05530. 2024.
- Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, Seunghyun Yoon and Hinrich Schütze. NoLiMa: Long-Context Evaluation Beyond Literal Matching. Proceedings of the 42nd International Conference on Machine Learning (ICML 2025). 2025.
- Aleksandar Petrov, Emanuele La Malfa, Philip Torr and Adel Bibi. Language Model Tokenizers Introduce Unfairness Between Languages. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). 2023.
- Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
- Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12. 2024.
- Kelly Hong, Anton Troynikov and Jeff Huber. Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma technical report. 2025.
- Tomas Mikolov, Wen-tau Yih and Geoffrey Zweig. Linguistic Regularities in Continuous Space Word Representations. Proceedings of NAACL-HLT 2013. 2013.
- Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava and Iryna Gurevych. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS 2021 Datasets and Benchmarks Track. 2021.
- Niklas Muennighoff, Nouamane Tazi, Loïc Magne and Nils Reimers. MTEB: Massive Text Embedding Benchmark. Proceedings of EACL 2023, pages 2014-2037. 2023.
- John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov and Alexander M. Rush. Text Embeddings Reveal (Almost) As Much As Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023). 2023.
- OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
Further reading
- Ashish Vaswani et al. Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017). 2017.
- Niklas Muennighoff, Nouamane Tazi, Loïc Magne and Nils Reimers. MTEB: Massive Text Embedding Benchmark. Proceedings of EACL 2023, pages 2014-2037. 2023.
- Kelly Hong, Anton Troynikov and Jeff Huber. Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma technical report. 2025.
Sources last verified 2026-10-08.