Insert token: how LLMs work, one token at a time

Samuel Batista

Samuel Batista / September 23, 2026

8 min read • ––– views

How does a model get from "The Eiffel Tower is in" to "Paris, France"? Press start to follow the words and word pieces, called tokens, through the animation. The switch at the top lets you compare three model designs, explained below.

Text. It starts as plain text.

  • Box heats up: heavy math
  • Dots along a pipe: data moving to or from memory
  • Dots smash into one: information compressed
Token IDs are real (GPT-4’s tokenizer); word scores, attention strengths and expert choices are illustrative. Real chips run every layer on the same compute units.

The model first processes the text you've given it, a stage called prefill. It can do much of that work in parallel because the prompt's tokens are already known. Once it starts generating, each new token depends on the ones just chosen, so it proceeds one at a time; that's decode. Processing a long prompt can mean a wait before the first token appears. During generation, moving the model's learned numbers, its weights, from memory into the chip's working storage often takes more time than the arithmetic itself.

Questions you might have

How do tokens change each other?

In "Maria lost her keys, so she," understanding "she" depends on connecting it to "Maria." Attention lets a model make connections like that by bringing information from earlier tokens into the one it's processing.

Each token is represented by a vector, a list of numbers. An attention layer transforms that vector into three lists: a query used to find relevant information, a key used for matching, and a value containing information to pass along. It compares the current token's query with the keys of the available tokens, then combines their values according to how well they match. This weighted combination is how information from several positions can contribute to one token's representation.

In a model generating text this way, attention can use the current token and earlier ones, but not future ones. The animation shows "in" drawing information from "Eiffel Tower," giving the model context for predicting a location. Those particular connections are illustrative, rather than measurements from a running model.

What is the KV cache for?

The KV cache stores the keys and values that earlier tokens produced at each layer. Once the model has processed "The Eiffel Tower is in," for example, it can reuse those numbers when it processes "Paris." It doesn't need to run the whole prompt through the network again.

That reuse works because earlier tokens can't attend to later ones: adding "Paris" doesn't change the keys and values already computed for the prompt. Full attention still reads the stored entries for each new token, but storing them avoids recalculating them. The growing stack beside each layer in the animation represents this cache.

For a sense of the cost, Llama 3.1 70B needs about 33 GB of cache for 100,000 tokens when stored at 16-bit precision, in addition to roughly 140 GB for its weights at that precision. Longer conversations use more storage and require more cache data to be read at each step, which can slow generation down.

What is the state in linear attention?

The linear-attention design shown here combines information from successive tokens into a fixed grid of numbers, called the state. Each new token updates that grid and retrieves information from it. Think of updating a running average: you retain information about what came before without keeping a separate entry for every item. The actual state is much richer than one average, and learned update rules control what it retains.

Its storage and the work needed to update it don't grow with the text. But information can be lost as the state changes, and there's no separate entry to revisit for an earlier token. Some models combine this approach with full attention to balance cost and recall. Kimi Linear, for example, uses three linear-attention layers for every full-attention layer. The fixed memory applies to its linear layers, not the whole model.

What does the feed-forward block do?

After attention has brought in information from the surrounding text, the feed-forward block, also called the MLP, transforms each token's representation independently. Its weights were learned during training, so the same input patterns produce similar changes each time the model runs.

The block first expands the token's list of numbers into a larger one. You can think of that as testing for many learned patterns at once. An activation function controls how strongly the resulting values contribute, and another transformation combines them back into the original number of dimensions. The output is then added to the token's representation. These are numerical operations learned from examples, rather than explicit if-then rules written by a programmer.

Researchers studying these blocks found patterns associated with wording in earlier layers and meaning in later ones. Other work traced factual recall to feed-forward blocks in the models studied. That helps explain how learned associations can enter the computation, though the animation's Eiffel Tower example doesn't trace the exact route a real model uses to recall Paris.

In a dense model, every token uses the same feed-forward block at a given layer. A mixture-of-experts model offers several such networks and selects a few for each token. That lets it have more learned weights without using all of them on every step, although the unused weights still need storage.

Where does the thinking happen?

There isn't a separate thinking block. In the model shown here, each layer uses attention and a feed-forward block to update the token's representation. The early layers' results become inputs to later ones, so information can be combined and transformed through many steps before the model predicts anything.

At the end of the prompt, the representation for "in" has been shaped by "The Eiffel Tower is in" as a whole. The model uses it to predict the next token. Once "Paris" is chosen, that token goes through the network too, using the cached information to predict what follows. How AI models think goes further into what can happen within those repeated computations.

What's special about the LM head?

The language model head converts the final representation into a score for each token in the vocabulary. At the start of the model, an embedding table turns a token ID into a list of numbers. The LM head goes in the other direction, from numbers to candidate tokens, though it isn't necessarily the same table or a mathematical inverse.

To continue the prompt, we only need the prediction from its final position. In the animation, that's "in." Earlier positions would predict tokens already present in the prompt, so their predictions aren't needed to choose the next one.

The scores are converted into probabilities, and the generation software selects a token. It can take the highest-scoring one or sample from the distribution. Temperature adjusts that distribution: lower values concentrate probability on the stronger candidates, while higher values spread it out. Sampling is one reason the same prompt can produce different answers.

Where does reinforcement learning from human feedback (RLHF) fit in?

The animation shows a trained model running, so its weights are already fixed. Pretraining is where it first learns from large amounts of text: it predicts the next token, compares that prediction with the actual one, and adjusts its weights to make the observed text more likely. Those adjustments affect attention and feed-forward blocks alike.

Post-training then develops behaviors such as following instructions and producing useful answers. RLHF is one approach: people compare model responses, those preferences train a reward model, and the language model is optimized to produce responses that receive higher rewards. Other methods use example answers or automatically checked results. The model's basic structure can stay the same through these stages; what changes is the learned weights and therefore the responses it tends to produce.

How does the model know when to stop thinking?

For a model that writes out intermediate reasoning, those steps are generated through the same token-by-token process as the answer. DeepSeek R1, for example, separates its reasoning from its answer with <think> and </think> markers. Generating the closing marker signals the transition to the answer. Other models and apps can use different formats or impose their own limits.

Training shapes when the model makes that transition, but the software running it can also intervene. In the s1 experiments, researchers shortened reasoning by inserting an ending marker, or extended it by preventing the model from stopping and appending "Wait." In those experiments, extra reasoning sometimes led the model to correct an earlier mistake.

Each reasoning token requires another pass through the model and, with a standard KV cache, another stored entry at every layer. That takes time and memory. It also leaves intermediate work available for later tokens to use, including the tokens in the final answer.