Lightning Machine: GLM-5.3's architecture in 3D

Samuel Batista

Samuel Batista / October 11, 2026

11 min read • ––– views

In this example, a user asks GLM-5.3 what they named a dice roller much earlier in a long conversation. Following First Light, this animation traces how another model builds an answer.

The model works with tokens: words, word pieces, punctuation and special markers, each represented by a list of numbers. The towers show GLM-5.3's 78 processing layers. Each layer stores a record of the numbers it receives. A small network called an indexer selects up to 2,048 records to read. Attention combines information from those records and adds it to the current token's numbers. A neural network then calculates another update before the token moves to the next layer.

The conversation, answer, scores and selections are illustrative, not a recording of GLM-5.3 generating an answer. Model dimensions and token IDs come from GLM-5.3's released files. Before you start, choose how to watch: with a narrator, with captions, or both. You can switch later under Presentation, the film-strip icon in the control bar. The narrator's voice comes from Kokoro-82M, an open-weight text-to-speech model by hexgrad. Study mode turns the music off and pauses after each caption until you continue. Quiet mode shows captions with all sound off.

Questions you might have

What are the towers, the rods and the orbs?

The towers hold the weights: numbers learned during training that stay fixed throughout a conversation. At each tower's foot, 64 slats represent its attention heads: separate calculations that read the same records in different ways. Above them, rings represent smaller neural networks called experts. The crown is a shared expert that runs for every token, while only a few rings run for each token. Towers that choose records with their own indexer also have a spire.

The rods are stored records of tokens. Each layer makes its own records from the numbers it receives, so these aren't 78 identical copies. Together, they form the key-value cache, or KV cache. Later tokens can read these records instead of running the earlier tokens through the model again.

Each layer makes its records from the numbers arriving at that layer, using its own weights. Those numbers change between layers, so one layer's records can't substitute for another's. Building the KV cache from scratch means running every token through every layer, which is what makes prefill expensive for a long conversation. Keeping the cache between turns means only new tokens need processing; without it, the whole conversation must be prefilled again. Here, the earlier 20,000 tokens are already cached, so prefill processes only the question's 15 tokens.

Each orb carries 6,144 numbers that change as a token passes through the layers. We call them its working memory. The orb holds the token's numbers as they're being updated; the rods keep information for later tokens to read. The label under each orb starts as its token's text, then switches to hex digits when the first result is added. The digits stand in for the orb's numbers, with illustrative values rather than the model's real ones, and change only when another result is added. At the output head, the orb's numbers score every possible next token, and the label locks, letter by letter, into the one that's picked.

Why read at most 2,048 tokens?

Reading every record takes more work as a conversation grows. GLM-5.3 uses DeepSeek Sparse Attention to reduce that work: a small network called an indexer gives each record a quick score, then attention reads up to 2,048 of the highest-scoring records. The indexer still has more records to check as the conversation grows. But attention never has to read more than 2,048 records for a token.

Records that aren't picked stay in that layer's cache, ready for a later token to read. If there are 2,048 or fewer, attention can read them all. And even when several new tokens are processed together, each can read only its own record and earlier ones, never the text that comes after it. The implementation applies both rules.

Why do neighbouring layers share one list?

Neighbouring layers often choose similar tokens to read. Rather than have every layer choose again, GLM-5.3 lets some reuse the previous layer's list. For example, layers 3, 4 and 5 read the token positions chosen by layer 2. Each still reads its own records using its own weights; only the list is shared.

The configuration puts indexers at layers 0, 1 and 2, then 6, 10, and so on: 21 indexers for 78 layers. The IndexCache paper explains why this saves work.

What's in a record?

A record holds information about a token when it arrives at a layer, before attention and the neural network update it. In the animation, it's a rod in front of that layer's tower.

Attention needs two things from a record: a key to help decide how much it should contribute, and a value containing information to pass on. Each head turns the current token's numbers into a query: the numbers it compares with each key to score the match. Those scores become shares that add up to 100%, with higher scores giving larger shares. The head multiplies each value by its share, then adds the results together.

Storing a full key and value for every head would take a lot of memory. Instead, Multi-head Latent Attention stores a compressed record of 512 numbers, plus 64 for the token's position in the text. Each head uses these to produce the keys and values it needs.

That's 576 numbers per token per layer. Full keys and values for all 64 heads would need about 57 times as many. This comparison covers the stored records, not the entire model's memory use. Layers with an indexer also store a separate 128-number key per token for its quick scoring step.

Each rod's top is coloured by one number from its record, using the same colours as the orb and input table. Because layers make different records, each group of rods has a different colour pattern. Where a layer has an indexer, a pale square on the rod represents its separate indexer key.

What does the dimmer inside a unit represent?

A unit is one small calculation within the neural network. It multiplies its inputs by learned weights and adds the products together. It does this with two sets of weights, giving two sums. A calculation called SiLU turns the first sum into a multiplier. Multiplying the second sum by it gives the unit's output.

The dimmer represents that multiplier. Unlike a household dimmer, it can go below zero or above one, so it can reverse the sign or increase the size of the other sum. The whole calculation is called SwiGLU.

How big are the numbers really?

The panel's numbers are examples chosen to be easy to read, not values measured while running GLM-5.3. The real weights can be much smaller. In the released model weights, for example, the first unit in layer 0 has input weights typically around 0.01 in size, with none farther from zero than 0.07. A unit adds thousands of products of these small numbers.

Before attention and before the neural network, the model also adjusts the size of the numbers it receives. It measures their overall size, divides them by that measure, then multiplies each by a learned scale. This is RMS normalization. This controls the overall size of the inputs without making all the numbers equal or forcing each one into a fixed range.

What is the tree around the face?

Each point represents one of the vocabulary's 154,880 tokens, each identified by a number called its ID. The animation places lower IDs near the trunk and higher IDs farther out, with special tokens grouped at the root. The tokenizer, which splits text into tokens, can join shorter pieces into longer ones. These branches don't show those joins or map related meanings. They're just a way to arrange the vocabulary on screen; an ID doesn't tell us how often a token will appear in this conversation.

The face represents the model's next-token predictor, called the output head. It uses the last token's 6,144 numbers to give every possible next token a score. Those scores become probabilities: chances of being picked. The server, the software running the model, uses them to choose a token. It doesn't always have to choose the most likely one. Feeding the chosen token back into the model lets it predict what follows. The choices and odds shown here are illustrative.

Are the experts assigned subjects?

No one assigns an expert a subject such as "math" or "writing." Training shapes both what the experts do and how the model chooses between them.

In each of the first three layers, all the units in one neural network work on every token. This is called a dense network. Every later layer instead chooses from 256 smaller networks, or experts. The router scores them for the current token and chooses eight, with an adjustment that helps spread the work between experts. The router also sets how much each chosen expert contributes to their combined result. The shared expert's result is then added too.

The shared expert runs on every token. The idea is for it to learn patterns many tokens need, rather than having each of the other experts learn those patterns separately. That's the aim described in the DeepSeekMoE paper, not proof of what any particular GLM expert has learned.

Counting the weights from the configuration gives about 743 billion weights in the main model, with about 40 billion used for each token. Most of the rest sit in experts that token doesn't use. These counts don't include the extra prediction layer.

How does the extra layer guess ahead?

GLM-5.3 has one extra prediction layer, with about 10 billion weights. It has its own records, indexer, attention and experts, but uses the main model's input table to look up tokens' starting numbers and its output head to score possible next tokens. It combines the last token's working memory with the newly chosen token's starting numbers to guess what comes after the new token. That guess is called a draft; the main model still has to check it.

The animation first follows this after The is picked, then shows the same process throughout the reasoning. Later, the main model picks Snake and the extra layer proposes Eyes. Both go through the main model together. Processing Snake checks whether Eyes was a good guess, while processing Eyes predicts the full stop. In this example the guess is accepted, so that pass adds two new tokens: Eyes and ..

Processing the pair together can save time because some weights can be loaded once and used for both tokens. But the model still has to calculate results for both, and they may select different experts. The saving is greatest when waiting for weights to load takes more time than the calculations.

A guess can also fail. After the full stop, the extra layer proposes a line break, but the main model picks <|user|>, the token that hands the turn back to the user. The server throws away the draft and its stored records, then stops. Drafting and checking this way is called speculative decoding. The animation illustrates both outcomes, not a measured speedup. Its exact-match checks are a simple example; servers can also check guesses by comparing probabilities. How much time it saves depends on how often guesses succeed and how much work drafting and checking take.

Why does it think before it answers?

GLM-5.3's chat format starts a reply with a thinking section. The model generates this text using the same layers it uses for the answer. As it processes that text, it stores records that later tokens can read. This gives it more text to work from before answering, but isn't a transcript of every calculation inside the model. Software using the model can request more or less thinking with reasoning_effort.

How does it know when to stop?

Training teaches it when to produce special ending tokens; the server knows to stop when one appears. GLM-5.3's generation settings list three: <|user|> hands a chat back to you, <|observation|> hands control to a tool, and <|endoftext|> ends a document. Picking the stop token ends this run without another trip through the layers.

What changed from GLM-5.2?

The arrangement of layers shown here is the same. Z.ai says it used the same starting model as GLM-5.2. Further training, called post-training, improved its coding and ability to handle tasks with many steps. That training changed the weights, not the model's architecture.

Made with Claude Opus 5.5, GLM-5.3, GPT-6 Astra, Kokoro-82M, three.js and Web Audio API.