How AI models think

Samuel Batista

Samuel Batista / August 01, 2026

14 min read • ––– views

A large language model, at the level of hardware, is nothing but numbers. Billions of them, multiplied and added in enormous grids. And yet systems built this way write working code, catch subtle errors in arguments, and occasionally surprise the people reading them. This post is about the gap between those two facts: how a mountain of arithmetic turns into behavior that looks like thinking. Along the way it settles an argument I keep running into, the claim that these systems are glorified autocomplete. That claim turns out to be accurate about one thing and wrong about the thing that matters.

1. Words become geometry

Before a model can do anything with your words, it has to turn them into something it can compute with. That first step is called tokenization: the text gets chopped into tokens, small pieces that are usually whole words or word-fragments. Common words like "think" survive intact; a rarer word like "tokenization" gets split into a few chunks. A model never reads letters or words, it reads tokens, and everything else in this post happens to tokens.

The second step is where it gets interesting: every token is assigned a long list of numbers, a vector. Early models used 768 numbers per token; today's biggest models use many thousands. These numbers are not arbitrary labels. They are coordinates in a space, and training arranges that space so that geometry encodes meaning (where a word sits tells you what it means).

Related concepts end up near each other. Relationships become directions: traveling from "man" to "woman" moves you roughly the same way as traveling from "king" to "queen." Nobody programmed that in. It emerged because geometry is the cheapest way to squeeze an internet's worth of text into a machine of fixed size. That's the first answer to how numbers can hold meaning: they hold it the way a map holds a country, by position and direction, not by lookup.

2. One pass through the machine

A common guess is that a model secretly writes text to itself and re-reads it in a hidden loop. The truth is stranger: while a model is thinking, no words exist anywhere inside it. When a model produces a single token, the computation runs through the full depth of the network exactly once: in a large modern model, around ninety layers. What flows through those layers is one long list of numbers holding many overlapping pieces of information at once. If you've watched an AI app think out loud before answering, that's something else, an exception I'll get to.

That vector is called the residual stream, and it's the closest thing the model has to a mind. Picture a shared whiteboard. Each layer reads the whole board, computes something, and writes notes back. Take the prompt in the diagram below. The word "models" enters ambiguous: on its own it could mean AI models, fashion models, or business models, and its starting numbers carry all those meanings at once. Early layers consult the surrounding words and sharpen it to AI models; middle layers add abstract annotations like a question, not a statement, the asker wants an explanation. The final layers do a different job: they stop adding understanding and start converting it into a decision, narrowing everything on the board down to a score for every possible next token. Early layers ask "what am I reading?"; final layers ask "what do I say next?"

So how does a layer actually compute its notes? Two mechanisms alternate the whole way down. Attention is how words talk to each other: each word queries the rest of the text for what's relevant and pulls it in, which is how "models" found "think" sitting next to it and settled on the AI meaning. Feed-forward blocks are where knowledge lives: they take what attention just gathered and enrich it with facts absorbed during training, the way hearing "Paris" summons France, the Eiffel Tower, and croissants without any effort on your part. One mechanism moves information around; the other knows things. A layer is just these two steps, and the whole network is that pair repeated roughly ninety times.

So there are two loops. The inner loop is a single trip through all the layers, called a forward pass, and it happens entirely in numbers. The outer loop appends the chosen word to the text and runs the whole trip again. Only the outer loop involves words at all:

Prompt (tokens)howdomodelsthink?residualstreamAI models, not fashion modelsa question, not a statementthe next word takes shapeLayer 1Layer 2Layer 91Layer 92· · ·reads the whiteboard,writes notes backstill just numbers:no word exists yetone word outappended,pass repeats
The inner loop (solid): one trip through all the layers, entirely in numbers. The outer loop (dashed): the chosen word is appended to the text and the whole trip runs again.

3. The token bottleneck

That last step destroys an enormous amount of information. By the final layer, the model's internal state is a huge list of numbers, the accumulated result of the entire pass. The output is a single word picked from a menu of a couple hundred thousand options. This is lossy compression, the same trade you already know from JPEG artifacts and blurry low-resolution video: squeeze something rich into something small, and detail is lost in the process. Here the squeeze is extreme: an entire computation forced through a keyhole, once per word.

This one fact explains two behaviors that otherwise look puzzling.

First, it explains why asking a model to show its work, the technique known as chain of thought, helps so much. The number of layers is fixed, so each pass can only do so much thinking, and whatever the pass figured out internally is destroyed at the keyhole. But anything the model writes down survives. A written step gets read back in on every later pass, re-activating the internal patterns that produced it, so the next pass picks up where the last one left off instead of starting over. Thinking out loud chains many short passes into one long computation: the page becomes the model's working memory. The extended thinking modes in modern AI products are exactly this, packaged up: the model writes out reasoning before its answer, and the app hides or summarizes it. That's the one real case of output being fed back in, and it's a feature bolted on top, not the core mechanism.

Second, it explains why you can't fully trust a model's explanation of its own reasoning. When a model tells you how it got an answer, it isn't reading back a log of what happened inside, because there is no log. It's generating more text through the same keyhole. Anthropic's interpretability researchers caught this happening. Asked to add 36 and 59, Claude got there by two internal routes running at once: a rough estimate of the total, plus a separate precise calculation of the last digit. Asked how it did it, Claude described carrying the one, the way you were taught in school. That's not lying, because lying means knowing the truth and hiding it. The model can't see what it actually did, so it answers with the most plausible story available: the textbook method it read about in training. Some would say the model hallucinated that explanation, but call it whatever you like: hallucinated an answer, came up with an answer, autocompleted an answer, they all describe the same mechanism. I find this stuff fascinating: the way these models fail is nothing like the way we do.

4. Why "glorified autocomplete" misses the point

The case for calling models glorified autocomplete is simple: they were trained to do exactly one thing, predict the next token. True. The mistake is assuming a simple goal means a simple machine. Training works like this: take a simple rule (guess the next word, get corrected, adjust), run it on enormous amounts of hardware over enormous amounts of text, and let it grind for months. The rule stays simple the whole time. The machine it leaves behind does not: to keep getting better at the guessing game, the model is forced to grow real internal machinery, rough approximations of the mental models you and I use to make sense of the world. It has to hold those mental models in mind while it hunts for each next word. That, I'd argue, is what makes these systems seem like they can genuinely think and anticipate.

Why does predicting the next word teach a machine so much? Because predicting arbitrary human text well is arbitrarily hard. To predict the next move in a written-out chess game, you need to know the board. To predict the resolution of a murder mystery, you have to track who knew what and when. To predict the last line of a proof, you have to do the math. Cheap statistical shortcuts get you partway; after that, the only way to get better is to actually understand the thing the text is about: the game, the mystery, the math.

Autocomplete describes the interface, not the mechanism. It's like calling a chess grandmaster someone who picks legal moves: true about the output, silent about everything behind it. The same goes here: autocomplete tells you what comes out of the model, one word at a time, and says nothing about the mental models doing the choosing.

5. The evidence for mental models

This would be just a hopeful story if we couldn't look inside. We can, partially, and what's been found is direct. The cleanest demonstration is Othello-GPT: a small model trained on nothing but sequences of legal moves in the board game Othello. No board shown, no rules stated. Researchers looked inside and found a live picture of the current board. Better: when they surgically edited that internal picture, the model's next moves changed to match the altered board. The model wasn't looking up memorized answers. It had built its own mental model of the game, what researchers call a world model, from nothing but text.

Other findings point in the same direction. Models form concepts independent of any language: the same internal pattern lights up whether an idea appears in English, Chinese, or code. The concept lives inside the model once; each language is just a different doorway to it. Models also develop small internal routines, like "find where this pattern appeared earlier and continue it," that show up abruptly during training, like a switch flipping. The model discovers an algorithm. One result is strong evidence that models aren't just autocompleting one word after the next: when Claude writes a rhyming couplet, it picks the rhyme word for the end of the second line before writing the beginning of that line, and editing that internal plan changes the whole line. The model isn't only choosing the next word; it's holding the shape of the whole line in mind as it writes each one, planning ahead and improvising toward its own target.

The newest result ties these threads together. Anthropic researchers recently found that Claude contains something like a global workspace, a term borrowed from neuroscience, where it describes how the brain pulls whatever you're consciously focused on into one place. Claude's version is a small set of internal patterns sitting at the center of its reasoning. It holds only a few dozen concepts at a time and accounts for less than a tenth of the model's activity, but what's in it is what the model is thinking about right now. If that sounds like a privileged corner of the whiteboard, that's exactly what it is. Ask Claude to silently think of a sport, and researchers can read the choice in the workspace before a single word comes out; swap the internal pattern for soccer with the one for rugby, and Claude reports rugby. Give it mental math, and the intermediate values surface there silently, in the right order. And the workspace specializes: most routine processing bypasses it entirely, and removing it leaves fluent text intact while multi-step reasoning falls apart.

6. Where the intuition comes from

Put the pieces together. A forward pass does an enormous amount of work in parallel: no search, no backtracking, no deliberation inside it. Thousands of learned patterns light up together, combine, and settle on the likely next words. It takes the same fixed amount of time every time, and none of it can be narrated. That's what intuition is in humans too. A strong chess player who says a move looks right before calculating a single line is describing the same kind of process. You can't fully explain why a sentence sounds off to you either: the computation happened everywhere at once, not as a chain of steps you could narrate.

This also explains why bigger models feel smarter. Models cram in far more patterns than they have room for, a phenomenon called superposition: patterns packed together like overlapping radio stations, interfering with each other. A bigger model relieves the pressure. Patterns get cleaner, more numerous, more finely separated. A bigger model doesn't just know more facts; it can tell apart shades of meaning a smaller model is forced to blur together. Modern designs push further. Many models are now built as a team of specialists, routing each word to a few experts out of hundreds, so the same machinery isn't forced to handle both poetry and protein folding. And more layers buy more thinking steps before the answer has to be squeezed through the keyhole.

So is it just scale? My answer: scale, spent well. Capacity matters because of where it goes: into cleaner patterns, finer distinctions, and dedicated specialists. Size is the budget, not the magic.

7. What this does not settle

For me, the evidence rules out the "fancy lookup table" story: planning, mental models, and self-taught routines are not things a lookup table does. What the evidence does not settle is whether models understand in the deeper sense you might care about. Finding the board inside Othello-GPT proves the information is there and actually being used; whether the model understands what a board is, the way you do, is a harder question, and nobody knows how to answer it yet. The workspace is the same: it looks a lot like what neuroscientists describe when they talk about conscious processing, but the researchers are explicit that nothing in it shows the model experiences anything. The skeptics still have real points too: these systems need vastly more examples than humans to learn anything, their model of the world is built from descriptions of the world rather than contact with it, and nothing persists between conversations unless it's engineered in from outside.

One more caveat: an AI model is a poor witness about its own case. It has no reliable introspective access to its computation. Ask one how it reached an answer and it can't actually check; it writes a plausible answer to your question, the same way it writes everything else. The workspace research opens the first real window here: Claude can accurately report what's currently sitting in its workspace, and the report can be checked against the internal patterns themselves. But the window is narrow: you can see what the model is thinking about, not how it thinks. So the argument can't be settled by asking the model, and it can't be settled by pointing at the training objective either. It has to be settled by looking inside, and we've only just started looking.

The short version

Meaning is stored as geometry. Thinking happens in a single silent pass through many layers, and only the conclusion gets squeezed into a token. Predicting text well is so hard that training forces genuine mental models to form inside the network, and we've caught them in the act: representing game boards, discovering algorithms, planning rhymes before writing toward them, and gathering their thoughts in a small workspace we can now read. What looks like intuition is the forward pass itself: massive parallel pattern-matching, arriving at answers it can't narrate. Autocomplete is just the keyhole we watch it through.

If you want the engineering side of this story, how attention evolved from GPT-2 to today's models and how modern designs manage a fixed-size memory that has to decide what to keep, overwrite, and let fade, see From GPT2 to Kimi3, Explained.


Written with Claude (Fable 5). Practicing what I preach: I supplied the questions and the skepticism, and the model supplied the explanations.