+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++
+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++
+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++
+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++
+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++
Designing Kolibri: Architecture Trade-Offs from First Principles
TL;DR
This blogpost investigates the impact of the main architectural choices in autoregressive
language models on training and deployment cost, focusing on the allocation of parameters
and FLOPs, how FLOPs scale with context length and sequence-mixer1 state2 size. We derive these quantities
and analyse how they vary across recent open-weight architectures. It also includes an interactive
tool to configure your own model and compare it with recent open-weight releases.
Introduction
A language model’s architecture decides how its capacity and computation are distributed.
Many of these choices affect quality, training cost and serving cost at the same time. The
effect of a choice on quality has to be measured empirically, but the costs follow directly
from the configuration:
Total parameters set the weight-memory footprint, and with it the minimum hardware
needed to deploy the model.
FLOPs per token, multiplied by the training-token horizon,3 set training compute. They are also a useful proxy for prefill cost, which is typically compute-bound.
Sequence-mixer state2,
e.g. the key-value (KV) cache for an attention-based sequence mixer, sets how many
sequences, and how much context, fit in a decode batch, and how many bytes each decode
step has to read.
Which of these becomes the serving bottleneck depends on the workload. Small-batch,
short-context decoding is limited by reading the weights. Prefill and large-batch decoding
of short queries are limited by compute. Long-context decoding is limited by reading the
sequence-mixer state.
In this blogpost we derive each quantity for MoE decoders and analyse recent open-weight
models:
Parameter allocation: how active parameters
are allocated between the FFN (a Mixture of Experts (MoE) layer in all models we analyse here), the sequence-mixer (e.g. attention) and the Language
Modelling (LM) head.
FLOP allocation: how FLOPs per token are allocated
between components and scale with sequence length.
Sequence-mixer state: how much state
each sequence holds and each decode step reads, depending on sequence length and sequence-mixer
type.
We visualise these quantities in an interactive explorer below that lets you change the
configuration and see the effect on all three axes.
Architecture Configurations
The explorer below is a simplified subset of an internal tool we use when designing model
architectures. It helps us quantify the trade-offs between different architectural choices
before committing to a configuration. Inspired by Sebastian Raschka’s LLM Architecture Gallery, we compare recent open-weight architectures. We initialise the explorer with two Aleph
Alpha models, Kolibri Origin and Kolibri, the latter of which was recently released under
the Apache 2.0 licence.4 These provide two concrete
reference points for the design choices discussed throughout the post. You can modify either configuration
using the sliders or add your own with +.
Parameter Allocation
In MoE architectures, total and active parameters can differ by an order of magnitude. Total
parameters mainly affect the memory footprint, hence the minimum deployment requirements of
a model. Active parameters typically account for most of the FLOPs during training and
prefill. At small batch sizes, they also largely determine how many weight bytes each decode
step reads from memory.5
If you play with the sliders, you’ll notice that every slider except the full-attention/SWA
split and the window size affects total or active parameters. We now derive each component’s
contribution to the two totals.
FFN: Dense Layers, Routed and Shared Experts
Routed expert count and expert width determine the capacity of the FFN. Top-k and
shared experts determine how much of that capacity is active for one token.
A SwiGLU expert uses three linear projections:
gate, up and down. The gate and up projections map from the hidden dimension d to
the intermediate dimension de, while down projection maps back from de to d. This gives
3×d×de parameters per expert. Multiplying this number by the number of experts
and the number of MoE layers gives the routed total parameters. For the active routed parameters
we substitute E with k, since only the top-k experts will be used.
Prouted=Lmoe×3×d×de×{Ektotalactive
In the same way we can derive the shared experts
parameters. Note that multiple shared experts are equivalent to a single shared expert of larger
width
ds, since all shared experts are always active and their outputs are combined.
So we only consider a single shared expert here.
Pshared=Lmoe×3×d×ds
Many MoE decoders keep their first layers dense6. They can have their own width dff. In the sliders above dff is tied
to the active expert width, so dff=k×de, which is a common choice among
the models analysed here (e.g. MiMo-V2.5, MiniMax M3, Kolibri Origin).
Pdense=Ldense×3×d×dff
The router maps the full residual stream d to E logits, so its size grows
with the expert count and hidden size. However, it is a very small part of total parameters,
and we count it with the FFN.
Prouter=Lmoe×E×d
Sequence Mixer: Attention Projections
The Query, Key, Value and Output projections of one layer depend on the hidden width, the
query width and the KV width. Their total then scales with the number of attention layers.
In a GQA layer we project the hidden
state of width
d into a query of width dq and two KV tensors of width dkv, then project the attention result from dq back to d.
With the group size g=hq/hkv, MHA
(g=1) and MQA (hkv=1, so
g=hq) are the two edge cases of GQA.
Norms
Each decoder layer commonly carries two (only pre-norm) to four (sandwich norm) RMSNorms, with scale vectors of
width d, plus one final norm before the LM head. The factor four below assumes
sandwich norms.
Pnorms=(4L+1)×d
Embedding and LM Head
The product of vocabulary size V and hidden width d sets both the LM head
and the embedding. The latter only affects the total parameters, since it determines only a
light lookup operation instead of a matrix multiplication.
An untied input embedding and LM head each contain d×V weights. Tying them stores
one shared matrix instead of two, but does not remove the dense vocabulary projection from training
compute. You can verify how only the total parameters change using the embedding-tying toggle.
Sorting the models by total parameters, active parameters or the active-to-total ratio reveals no obvious relationship with the sequence-mixer share. The choice of sequence-mixer
share comes primarily down to empirically measured quality. Additionally, increasing the sequence-mixer
compute while keeping the state size constant increases arithmetic intensity and makes the primarily
memory bound kernels more efficient.
The ratio between sequence mixer and FFN parameters is mainly controlled by the relative
widths of the attention projections and the intermediate active width of the FFN. Using the
parameter counts derived above,
The residual-stream width d cancels out. Changing d alone therefore leaves
this split unchanged, while changing the attention or expert width directly shifts the allocation.
Finally, the LM head’s share shrinks as models grow. Most weight matrices grow quadratically
with the width, such as d×dq with dq∝d, while
the LM head (d×V) grows only linearly, because the vocabulary size V usually
stays fixed. You can see this clearly with the LM head stacked first and the models sorted by active parameters.
FLOP Allocation
Training compute per token has two main components: parameter FLOPs and attention FLOPs.
Parameter FLOPs come from multiplying activations with weight matrices and are independent
of the context length. Attention FLOPs come from the QK⊤ and AV products,
which involve no learned weight matrix but depend on the attention width and on the context length.
The equations in this section reuse the parameter-allocation symbols defined above.
The parameter-FLOPs term depends only on the active weights: the 6 comes from three matrix
multiplications of equal size, each costing two FLOPs per weight per token, one multiply and
one add. The forward pass (fwd) computes Y=XW.
Backpropagation then needs two more products of the same shape. The gradient with respect to
the activations (bwdA) propagates the error signal back into
the preceding layer, while the gradient with respect to the weights (bwdW) measures how the loss depends on each weight and how it should be updated to reduce the
loss.
Embedding lookup, norms and activations are not considered by the previous equation since,
at this scale, their contribution is negligible.7
Attention FLOPs involve no weight matrices: QK⊤ computes one dot product per token and
AV
applies one value vector per token. This results in 4×dq FLOPs per query-key pair,
and the same
×3 multiplier accounts for fwd, bwdA
and bwdW. We count the keys a query attends under the causal
mask, n/2 on average in a full-attention layer. FLOPs grow with the sequence length
because the query attends every retained key, which is all previous tokens in a full-attention
layer and at most the window min(n,w) in a SWA layer. Note that GQA leaves dq and the attention FLOPs unchanged: it only narrows
the KV width, which reduces the sequence-mixer state size, discussed in the next section.
Beyond Full and Sliding-Window Attention
The closed-form derivations below focus on standard and sliding-window attention, but
several architectures in our comparison use other sequence mixers. We only give the
intuition of these alternatives here and refer interested readers to the original papers for
their detailed formulations.
At a high level, these alternatives follow two main approaches. Sparse attention, such as Native Sparse Attention (NSA), reduces the number of tokens accessed from a longer context. Recurrent or linear sequence
mixers, such as Mamba-2 and Gated DeltaNet, instead compress information from the prefix into a fixed-size state. Recent models often
combine these mechanisms with attention in hybrid architectures. In our comparison, DeepSeek
V4, GLM-5.3, MiniMax M3 and Qwen3.8-Flash-Next use sparse attention, while
Qwen3.8-Flash-Next (Gated DeltaNet), Kimi K3 (Kimi Delta Attention) and Nemotron 3 (Mamba-2)
combine attention with recurrent layers.
FFN
The FFN transforms each token independently and, as we have seen before, is where most of
the capacity is spent. At fixed top-k, inactive routed expert matrices add
capacity that costs memory footprint but no FLOPs.
Each MoE layer adds one small router matrix on top.
Frouter/ token =6×Lmoe×E×d
Sequence-Mixer Projections
The sequence mixer models interaction between tokens. Its Query (Q), Key (K), Value (V)
and Output (O) linears depend on hidden width, attention width, KV width and depth, but
are independent of sequence length per token.
Fproj/ token =6×2×Lattn×d×(dq+dkv)
Pure Attention
It includes the attention score computation QK⊤ and the value aggregation AV, which depend on query width and retained keys. A full-attention layer retains every
previous token, so a query attends n/2 keys on average.
Fattn/ token =12×Lfull×dq×n/2
A sliding-window layer retains at most w of them.
Fattn/ token =12×Lswa×dq×min(n,w)
LM Head
The output projection scales with hidden width times vocabulary size.
Fhead/ token =6×d×V
FLOP Allocation vs. Context Length
Adjusting the sequence-length slider in Figure 3, you can see how the FLOP composition changes significantly with context length. At 16K context, pure attention accounts for roughly 3–51% of training FLOPs across the selected models.
As context approaches 1M tokens, this contribution can become dominant, reaching 98% for .
Parameter-bound FLOPs grow linearly with the number of tokens, while full-attention FLOPs
grow quadratically with sequence length. As a result, the FLOP mix shifts increasingly
toward pure attention as context grows, especially in architectures with many full-attention
layers.
Across models, the total fraction of FLOPs spent in the sequence mixer at 16K sequence length is more constrained, ranging from roughly 33% to 65%. As pure-attention
FLOPs are reduced through e.g. sparse attention, projection linears account for a larger fraction of total FLOPs. follows this pattern by reducing the number of full-attention layers and replacing them with
sliding window attention, substantially lowering the pure-attention share from 51% to 27% at 16K
sequence length compared to .
FLOP Scaling vs. Context Length
As Figure 4 shows, differences in compute cost become increasingly
pronounced as context length grows. has the highest FLOPs across the full range, combining a large active model size with 24 full-attention
layers out of 93. sits at the opposite end at long context, with only a small fraction of its sequence mixer scaling
quadratically with context length.
and need
similar FLOPs at 16K context, but at 1M tokens MiMo-V2.5 needs about 5× more: its full-attention
layers attend to every previous token, while DeepSeek V4.1-Flash attends to a fixed number of
selected, compressed tokens and scores the full prefix only with a light indexer.
follows the same
trend. Compared with , it uses sliding-window attention and roughly 5× fewer full-attention layers, so the
compute advantage widens as context length increases.
reaches the lowest FLOPs/token at long contexts in this comparison. Additionally, together with
it is the flattest curve over context length, with only 13% and 33% increase respectively from
4K to 1M.
Multiplying FLOPs by the token horizon D gives us the total training FLOPs, Ctraining≈D×Ftrain/ token 8 and dividing by your achieved training
FLOP/s gives an estimate of your training duration. Consider that the real training cost is far
more complex than just FLOPs. Some of the other major factors, not addressed here, are the overlap
of computation and communication, the quantisation of your weights, kernel efficiency and kernel
launch overhead, activation recomputation and expert imbalance.
FLOPs are a useful proxy for prefill cost, but they do not directly translate into faster
decoding, where memory movement becomes more important. This motivates the next section on
sequence-mixer state size.
Sequence-Mixer State Size
At every decode step, the new query must interact with the representations of earlier
tokens. Those key and value representations were already computed during the previous steps
and never change, so we can cache them instead of re-projecting, trading FLOPs for HBM
capacity and more I/O. The attention computation itself is unaffected, since at each step we
simply append the K and V vectors of the current token.
What limits decode speed depends on the regime: A decode step multiplies every weight matrix
with only one token per sequence, so at short context and small batch it is dominated by
streaming the model weights from HBM to the compute units (memory bound); for MoE, the
weights read are those of every expert selected by at least one sequence, which approach all
experts as the batch grows. At large batch and short context, weight reuse raises arithmetic
intensity and compute can become the limit (compute bound). Once contexts grow long, the
bottleneck is the I/O cost of reading the state for every generated token, which, unlike the
weights, is not shared between sequences (memory bound). These are only general regimes: the
limiting factor depends on the arithmetic intensity of the operation and on the hardware
specifications, in particular the bandwidth (byte/s) and the throughput (FLOP/s), both of
which also depend on the numerical precision. For a deeper explanation of these topics, read
our post Inference performance from first principles.
Mstate=c×B×b×dkv×Lkv×nret
Cached Representation
Standard attention stores separate key and value vectors (c=2). Multi-head
latent attention (MLA, introduced in DeepSeek-V2) instead caches a single vector per token (c=1): a low-rank content latent
plus the positional component required by its decoupled RoPE path.
Batch
In a naive implementation every sequence of the batch B has its own independent state.
Prefix caching instead reuses a shared
prefix (e.g. the system prompt) across sequences, so one copy of that state serves many requests.
Bytes per Element
Values in the cache are commonly stored in 1 byte (FP8) or 2 bytes (BF16). Methods such as KIVI and KVQuant go down to 2–4 bits, but
are not yet common in production serving. Recurrent states are often kept in FP32 instead.
KV Width
The number of KV heads times the head dimension sets dkv. GQA shares one KV head across a group of query heads, while MQA uses a single KV head for all queries.
Attention Layers
By default every attention layer keeps its own state. Hybrids replace most full-attention
layers with sliding-window or recurrent layers, and research directions such as
cross-layer attention (CLA) let
groups of layers share one layer’s K and V. DeepSeek V4.1-Flash, for example, computes its
compressed cache in four layers and lets the layers in between reuse it.
Retained Tokens
Full attention retains the entire prefix. A common and robust variation is the sliding window, which keeps only the newest w tokens. Selective attention, such as NSA, reads only a sparse subset of the prefix at each step. This cuts the bytes read per
step, but not the bytes stored, since any block may be selected later.
Recurrent Layers
Mamba-2 and Gated DeltaNet layers fold the whole prefix into a fixed state per head, dk×dv elements for
GDN and dh×dstate for Mamba-2, plus the last few inputs of a short convolution.
A Qwen3.8-Flash-Next GDN layer keeps 48 heads of 128×128, 0.79M elements, as
many as one of its attention layers caches (K, V and indexer key) for about 680 tokens,
but the same at any context length.
Sequence-Mixer State Across Architectures
Figure 5 shows the total sequence-mixer state and the amount
of state read at each decode step, which capture two different constraints.9 Total state size determines how many sequences and how much context can fit in a decode batch,
while the state read per token is more directly tied to decode speed in the long-context, memory-bound
regime.
This distinction is particularly important for sparse attention. Many sparse mechanisms
reduce the number of tokens accessed at each step, lowering memory traffic and FLOPs,
without reducing the total state that must be stored. At 1M tokens, holds 72B elements per sequence but reads 11B per decode step. Sliding-window and recurrent
layers reduce both, while techniques such as KV compression or cross-layer sharing can also shrink
the stored state itself.
This produces a substantially different ranking from FLOPs. , for example, has lower long-context FLOPs than , while the latter achieves a smaller sequence-mixer state through KV compression and
reuse. and
DeepSeek V4.1-Flash similarly occupy comparable parameter scales but very different points in
terms of state size: at 1M tokens, 72B elements per sequence for MiniMax M3 against 1.7B for DeepSeek
V4.1-Flash.
also improves substantially
over by replacing most full-attention layers with sliding-window attention, reducing the state that
grows with context length: at 1M tokens it holds 10.8B elements per sequence, 5× less than Kolibri
Origin’s 53.7B.
Conclusion
In this blogpost, we present our Model Explorer, used for designing the Kolibri
architecture. We analyse how parameters and FLOPs are allocated across the main model
components for popular open models and how the sequence-mixer state grows with sequence
length under different architectures. Additionally, we derive formulas for FLOPs and
parameters for our Kolibri models. We are excited to share this with the community, so
others can explore architecture trade-offs and reason about architecture from first
principles.
Acknowledgements
We are grateful to Ahmed Hammam, Fabien Benureau, Jonas Knupp, Simon Thel, Thomas Burns and
Yasser Jadidi for their thoughtful review. We thank Alexander Wortmeier and Noé Beckerle
Vallejo for their support bringing it to our website and Helena Treeck and Svenja Fahlisch
for keeping everything on track.