Follow the white rabbit: DeepSeek V4's architecture in 3D

Samuel Batista

Samuel Batista / September 24, 2026

6 min read • ––– views

In Insert token, we watched a model write a sentence. This time we're looking inside DeepSeek V4-Pro, which can work with about a million tokens, the words and word pieces that make up its input. Take the red pill to follow one token through its 61 layers, then watch it generate the next ones.

As a conversation grows, looking back through it becomes more expensive. V4 reduces that work by combining the information stored for older tokens into fewer entries, while keeping separate entries for the most recent 128. Different layers use different amounts of compression: some combine 128 tokens per entry, while others combine four and select the entries most relevant to the current token. It's a little like keeping notes on a long book at different levels of detail, except V4's notes are arrays of numbers, not written summaries.

Questions you might have

Why is a million tokens so hard?

With full attention, processing a new token involves comparing it with every earlier token at each layer. A longer conversation means more comparisons, and the model repeats that work for each token it generates.

The model also needs somewhere to store the earlier tokens' keys and values, the numbers attention uses to find and retrieve information. That's the KV cache. The Llama example in Insert token needs about 33 GB of cache for 100,000 tokens at 16-bit precision. Extending the same setup to a million tokens would take about 330 GB for one conversation, before accounting for the model itself.

Long documents aren't the only reason to need that space. A reasoning model's intermediate work takes up tokens too, as do the tool results and conversation history accumulated during a long task. DeepSeek designed V4 with these uses in mind: its report puts the compute per generated token at 27% of V3.2's, and the KV cache at 10%, with a million-token context.

How does the compression work?

Heavily Compressed Attention (HCA) combines the numerical representations of 128 tokens into one entry. It uses weighted averages, allowing some tokens to contribute more than others. Training teaches the model how to assign those weights, including different weights for different parts of the representation. The result takes much less space, but can't preserve every detail of the original tokens.

Compressed Sparse Attention (CSA) produces an entry for every four tokens, so it retains more detail. Each entry also draws on the previous group of four, letting information cross the boundary between groups. Both designs are described in the report's attention section.

Alongside either kind of compression, a sliding window keeps separate representations of the most recent 128 tokens. That preserves finer detail nearby and makes tokens in an unfinished group available immediately. In "Follow the white," for example, "white" can attend to "Follow the" even though their group hasn't yet become a compressed entry.

What is the lightning indexer?

The lightning indexer chooses which compressed entries a CSA layer will read. At V4's maximum context length, compression still leaves about 262,000 entries. The indexer estimates their relevance to the current token, then passes the top 1,024 to attention, alongside the recent-token window.

This first pass is cheaper because it uses smaller representations and lower-precision numbers. Think of checking search results before opening the pages: you use a quick comparison to narrow down where to spend more effort. V4 adapts this approach from DeepSeek V3.2's sparse attention.

The indexer still checks every compressed entry, so its work grows with the conversation. And selection can miss something useful: an older entry it skips won't be read by that layer's attention on that step.

Why four lanes instead of one?

A token normally passes from layer to layer as one list of numbers, called the residual stream. Each layer adds its result to that list. V4 carries four lists instead, which gives it more room to preserve different information as the token moves through the network.

This approach is called hyper-connections. A layer reads a weighted combination of the four lists, does its usual computation, and distributes the result back across them. The lists also mix with one another between layers. If that mixing repeatedly amplifies their values, the increases can compound and destabilize training.

V4's manifold-constrained hyper-connections (mHC) restrict that mixing to weighted averages. The mixing weights form a four-by-four grid, with nonnegative entries and rows and columns that each sum to approximately one. V4 approaches that constraint by repeatedly adjusting the row and column totals. That's the grid settling in the animation: the model can redistribute information without that mixing step repeatedly magnifying it.

Why have 384 experts if each token uses only six?

Each expert is a smaller neural network with its own learned weights. Having many experts gives the model more capacity, but running all of them for every token would be expensive. V4 selects six of the 384 routed experts in each layer and also runs one shared expert. Across the model, a token uses about 49 billion of its 1.6 trillion parameters.

DeepSeek's reason for using many small experts is that they can specialize and be combined in different ways. The shared expert handles patterns useful across tokens, reducing the need for routed experts to learn the same things.

Training also has to keep the work reasonably balanced. If the router keeps choosing the same experts, others get little use. DeepSeek adjusts a selection bias for each expert, making underused experts easier to select and overloaded ones harder. That changes which experts run, without directly changing how much their outputs contribute.

The inactive experts still need storage. V4 uses 4-bit numbers for their weights to reduce that cost, but selecting fewer experts mainly saves computation; it doesn't make the rest of the model disappear.

What's the catch?

Compression can lose information, and selection can skip information that's still stored. An HCA entry can't preserve every detail of 128 separate representations. CSA keeps more detail, but an older entry has to pass the indexer's selection before that layer can use it. Neither approach guarantees that the model will retrieve a particular detail from a long conversation.

One test of that ability is MRCR. It puts similar requests throughout a long conversation, then asks for a particular earlier response, such as the second poem about tapirs. DeepSeek reports that V4-Pro's retrieval stays steady up to about 128,000 tokens and declines beyond that. The test shows the limitation, but doesn't establish how much of it comes from compression.