Circuit City: Kimi K3's architecture in 3D

Samuel Batista

Samuel Batista / September 26, 2026

6 min read • ––– views

After DeepSeek V4, here's a look inside Moonshot AI's Kimi K3. The city below follows a token, a word or piece of a word, through the model's 93 layers of processing as it works toward choosing what comes next.

What I find most striking is how K3 remembers what you've written. Most of its layers keep updating a fixed amount of memory as they read, so a longer conversation doesn't need more storage in those layers. The tradeoff is that new information can displace older details. Other layers keep a separate set of numbers representing each token, giving the model somewhere to look back for details the fixed memory may have lost. Those are the notebooks and memory blocks in the animation: two ways of storing information about the conversation, with different costs. K3 uses both.

Questions you might have

Where are the weights and the context kept?

In the animation, the buildings hold the weights: the 2.78 trillion numbers K3 learned in training. The memory blocks and notebook pages hold the context: what your conversation has left behind so far. The two sit side by side in every layer, but they behave very differently.

The weights are fixed once training ends. Every conversation runs through the same ones, and talking to K3 doesn't change them. The context belongs to a single conversation and changes with every token, since each new token adds a memory block to every MLA layer and updates every KDA notebook.

On real hardware, both are kept in the memory of the chips running the model. The weights are far too large for one chip, so they're spread across many, loaded once, and shared by every conversation those chips are serving. Each conversation's context is kept alongside them while it's being processed, which is why a long conversation costs more memory to serve than a short one.

Why does K3 need two kinds of memory?

Imagine reading a long document while updating a single page of notes. You can keep the main ideas, but you can't expect every detail to survive. Kimi Delta Attention (KDA) has a similar constraint: it combines information from successive tokens into a fixed grid of numbers. That's the "notebook" in the animation, and it can lose information as it's updated.

Multi-head latent attention (MLA) keeps a compressed numerical representation of each token separately. Each violet memory block in the animation represents one of those stored records. When processing a new token, an MLA layer can compare it with all the stored representations and draw information from the relevant ones. Earlier entries stay available as new ones arrive, though that doesn't guarantee perfect recall.

Keeping those entries takes more memory as the conversation grows, and checking them takes more work. K3 balances the costs with 69 KDA layers and 24 MLA layers: every fourth layer uses MLA, along with the final layer.

What does the delta in Kimi Delta Attention mean?

"Delta" means difference. Each token supplies a key, a set of numbers used to look up related information, and a value, the information to store. The delta rule compares what the memory returns for that key with the new value, then uses the difference to adjust the memory. That lets it correct an existing association instead of simply adding more information on top.

KDA also controls how much of the previous memory to retain. Different parts can fade at different rates, depending on the incoming token, so some information can persist while other information is replaced more quickly.

How does K3 keep track of word order?

The order in which you update a memory matters. Reading "dog bites man" produces a different sequence of updates from reading "man bites dog." KDA carries information forward in that order, and it also mixes information from each token with the three immediately before it. You can see both operations in the implementation.

K3's MLA layers don't add their own position encoding, the extra numbers many models use to mark where a token sits. They receive representations that have already passed through KDA layers, so they can still use information about word order. Turning off that extra encoding doesn't make the model blind to sequence.

What are the banks?

The banks show how K3 reuses work from earlier layers. In a typical transformer, each layer adds its result to the numbers being passed forward. A later layer receives the combined result, rather than separate results it can choose between.

K3's Attention Residuals keep some of those results separate. It groups layers into blocks of 12 and saves each completed block's combined output as a "bank." A later layer weights those saved outputs and the current block's running result, taking more from some and less from others. This lets it use an earlier stage's work directly, even after more layers have processed the token.

The animation shows this selection once per row; the model does it before both attention and the following network computation.

What are the experts, and why do they get fewer numbers?

An expert is a smaller neural network within the model. In each of K3's 92 expert layers, a routing mechanism chooses 16 of 896 experts to process the token. This gives the model many sets of learned patterns to draw on without running all of them for every token.

Moving information to those experts also costs time when they're spread across chips. K3 reduces each token's representation from 7,168 numbers to 3,584 before sending it to the selected experts, then expands their combined result back to the original size. That halves the amount of data in each representation being sent. Two additional shared experts process every token at the full width.

What's the catch?

Updating the same memory over and over makes it harder to return to an earlier point in the text. With separate entries for each token, you can discard the entries after that point. With KDA, those later tokens have already changed the shared memory, so you need a saved earlier state or a way to recompute it.

That matters when a server reuses the beginning of a prompt, or when a small model drafts several tokens for K3 to check. If K3 rejects a draft token, its memory must also return to the point before that token. The storage savings come with extra work in the software serving the model; Moonshot calls out prompt reuse as one of those challenges.