+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++

+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++

+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++

+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++

+++ Meet us at SCCON in Berlin, 13–15 October +++ Experience Kolibri, our new sovereign model, live in action for the public sector +++ Hall 26A, Booth 126 +++

Designing Kolibri: Architecture Trade-Offs from First Principles

TL;DR

This blogpost investigates the impact of the main architectural choices in autoregressive

language models on training and deployment cost, focusing on the allocation of parameters

and FLOPs, how FLOPs scale with context length and sequence-mixer1 state2 size. We derive these quantities

and analyse how they vary across recent open-weight architectures. It also includes an interactive

tool to configure your own model and compare it with recent open-weight releases.

Introduction

A language model’s architecture decides how its capacity and computation are distributed.

Many of these choices affect quality, training cost and serving cost at the same time. The

effect of a choice on quality has to be measured empirically, but the costs follow directly

from the configuration:

Total parameters set the weight-memory footprint, and with it the minimum hardware

needed to deploy the model.

FLOPs per token, multiplied by the training-token horizon,3 set training compute. They are also a useful proxy for prefill cost, which is typically compute-bound.

Sequence-mixer state2,

e.g. the key-value (KV) cache for an attention-based sequence mixer, sets how many

sequences, and how much context, fit in a decode batch, and how many bytes each decode

step has to read.

Which of these becomes the serving bottleneck depends on the workload. Small-batch,

short-context decoding is limited by reading the weights. Prefill and large-batch decoding

of short queries are limited by compute. Long-context decoding is limited by reading the

sequence-mixer state.

In this blogpost we derive each quantity for MoE decoders and analyse recent open-weight

models:

Parameter allocation: how active parameters

are allocated between the FFN (a Mixture of Experts (MoE) layer in all models we analyse here), the sequence-mixer (e.g. attention) and the Language

Modelling (LM) head.

FLOP allocation: how FLOPs per token are allocated

between components and scale with sequence length.

Sequence-mixer state: how much state

each sequence holds and each decode step reads, depending on sequence length and sequence-mixer

type.

We visualise these quantities in an interactive explorer below that lets you change the

configuration and see the effect on all three axes.

Architecture Configurations

The explorer below is a simplified subset of an internal tool we use when designing model

architectures. It helps us quantify the trade-offs between different architectural choices

before committing to a configuration. Inspired by Sebastian Raschka’s LLM Architecture Gallery, we compare recent open-weight architectures. We initialise the explorer with two Aleph

Alpha models, Kolibri Origin and Kolibri, the latter of which was recently released under

the Apache 2.0 licence.4 These provide two concrete

reference points for the design choices discussed throughout the post. You can modify either configuration

using the sliders or add your own with +.

Parameter Allocation

In MoE architectures, total and active parameters can differ by an order of magnitude. Total

parameters mainly affect the memory footprint, hence the minimum deployment requirements of

a model. Active parameters typically account for most of the FLOPs during training and

prefill. At small batch sizes, they also largely determine how many weight bytes each decode

step reads from memory.5

If you play with the sliders, you’ll notice that every slider except the full-attention/SWA

split and the window size affects total or active parameters. We now derive each component’s

contribution to the two totals.

FFN: Dense Layers, Routed and Shared Experts

Routed expert count and expert width determine the capacity of the FFN. Top-k and

shared experts determine how much of that capacity is active for one token.

A SwiGLU expert uses three linear projections:

gate, up and down. The gate and up projections map from the hidden dimension d to

the intermediate dimension de, while down projection maps back from de to d. This gives

3×d×de parameters per expert. Multiplying this number by the number of experts

and the number of MoE layers gives the routed total parameters. For the active routed parameters

we substitute E with k, since only the top-k experts will be used.

Prouted=Lmoe×3×d×de×{Ektotalactive

In the same way we can derive the shared experts

parameters. Note that multiple shared experts are equivalent to a single shared expert of larger

width

ds, since all shared experts are always active and their outputs are combined.

So we only consider a single shared expert here.

Pshared=Lmoe×3×d×ds

Many MoE decoders keep their first layers dense6. They can have their own width dff. In the sliders above dff is tied

to the active expert width, so dff=k×de, which is a common choice among

the models analysed here (e.g. MiMo-V2.5, MiniMax M3, Kolibri Origin).

Pdense=Ldense×3×d×dff

The router maps the full residual stream d to E logits, so its size grows

with the expert count and hidden size. However, it is a very small part of total parameters,

and we count it with the FFN.

Prouter=Lmoe×E×d

Sequence Mixer: Attention Projections

The Query, Key, Value and Output projections of one layer depend on the hidden width, the

query width and the KV width. Their total then scales with the number of attention layers.

In a GQA layer we project the hidden

state of width

d into a query of width dq and two KV tensors of width dkv, then project the attention result from dq back to d.

With the group size g=hq/hkv, MHA

(g=1) and MQA (hkv=1, so

g=hq) are the two edge cases of GQA.

Norms

Each decoder layer commonly carries two (only pre-norm) to four (sandwich norm) RMSNorms, with scale vectors of

width d, plus one final norm before the LM head. The factor four below assumes

sandwich norms.

Pnorms=(4L+1)×d

Embedding and LM Head

The product of vocabulary size V and hidden width d sets both the LM head

and the embedding. The latter only affects the total parameters, since it determines only a

light lookup operation instead of a matrix multiplication.

An untied input embedding and LM head each contain d×V weights. Tying them stores

one shared matrix instead of two, but does not remove the dense vocabulary projection from training

compute. You can verify how only the total parameters change using the embedding-tying toggle.

Sorting the models by total parameters, active parameters or the active-to-total ratio reveals no obvious relationship with the sequence-mixer share. The choice of sequence-mixer

share comes primarily down to empirically measured quality. Additionally, increasing the sequence-mixer

compute while keeping the state size constant increases arithmetic intensity and makes the primarily

memory bound kernels more efficient.

The ratio between sequence mixer and FFN parameters is mainly controlled by the relative

widths of the attention projections and the intermediate active width of the FFN. Using the

parameter counts derived above,

The residual-stream width d cancels out. Changing d alone therefore leaves

this split unchanged, while changing the attention or expert width directly shifts the allocation.

Finally, the LM head’s share shrinks as models grow. Most weight matrices grow quadratically

with the width, such as d×dq with dq∝d, while

the LM head (d×V) grows only linearly, because the vocabulary size V usually

stays fixed. You can see this clearly with the LM head stacked first and the models sorted by active parameters.

FLOP Allocation

Training compute per token has two main components: parameter FLOPs and attention FLOPs.

Parameter FLOPs come from multiplying activations with weight matrices and are independent

of the context length. Attention FLOPs come from the QK⊤ and AV products,

which involve no learned weight matrix but depend on the attention width and on the context length.

The equations in this section reuse the parameter-allocation symbols defined above.

The parameter-FLOPs term depends only on the active weights: the 6 comes from three matrix

multiplications of equal size, each costing two FLOPs per weight per token, one multiply and

one add. The forward pass (fwd) computes Y=XW.

Backpropagation then needs two more products of the same shape. The gradient with respect to

the activations (bwdA) propagates the error signal back into

the preceding layer, while the gradient with respect to the weights (bwdW) measures how the loss depends on each weight and how it should be updated to reduce the

loss.

Embedding lookup, norms and activations are not considered by the previous equation since,

at this scale, their contribution is negligible.7

Attention FLOPs involve no weight matrices: QK⊤ computes one dot product per token and

AV

applies one value vector per token. This results in 4×dq FLOPs per query-key pair,

and the same

×3 multiplier accounts for fwd, bwdA

and bwdW. We count the keys a query attends under the causal

mask, n/2 on average in a full-attention layer. FLOPs grow with the sequence length

because the query attends every retained key, which is all previous tokens in a full-attention

layer and at most the window min(n,w) in a SWA layer. Note that GQA leaves dq and the attention FLOPs unchanged: it only narrows

the KV width, which reduces the sequence-mixer state size, discussed in the next section.

Beyond Full and Sliding-Window Attention

The closed-form derivations below focus on standard and sliding-window attention, but

several architectures in our comparison use other sequence mixers. We only give the

intuition of these alternatives here and refer interested readers to the original papers for

their detailed formulations.

At a high level, these alternatives follow two main approaches. Sparse attention, such as Native Sparse Attention (NSA), reduces the number of tokens accessed from a longer context. Recurrent or linear sequence

mixers, such as Mamba-2 and Gated DeltaNet, instead compress information from the prefix into a fixed-size state. Recent models often

combine these mechanisms with attention in hybrid architectures. In our comparison, DeepSeek

V4, GLM-5.3, MiniMax M3 and Qwen3.8-Flash-Next use sparse attention, while

Qwen3.8-Flash-Next (Gated DeltaNet), Kimi K3 (Kimi Delta Attention) and Nemotron 3 (Mamba-2)

combine attention with recurrent layers.

FFN

The FFN transforms each token independently and, as we have seen before, is where most of

the capacity is spent. At fixed top-k, inactive routed expert matrices add

capacity that costs memory footprint but no FLOPs.

Each MoE layer adds one small router matrix on top.

Frouter/ token =6×Lmoe×E×d

Sequence-Mixer Projections

The sequence mixer models interaction between tokens. Its Query (Q), Key (K), Value (V)

and Output (O) linears depend on hidden width, attention width, KV width and depth, but

are independent of sequence length per token.

Fproj/ token =6×2×Lattn×d×(dq+dkv)

Pure Attention

It includes the attention score computation QK⊤ and the value aggregation AV, which depend on query width and retained keys. A full-attention layer retains every

previous token, so a query attends n/2 keys on average.

Fattn/ token =12×Lfull×dq×n/2

A sliding-window layer retains at most w of them.

Fattn/ token =12×Lswa×dq×min(n,w)

LM Head

The output projection scales with hidden width times vocabulary size.

Fhead/ token =6×d×V

FLOP Allocation vs. Context Length

Adjusting the sequence-length slider in Figure 3, you can see how the FLOP composition changes significantly with context length. At 16K context, pure attention accounts for roughly 3–51% of training FLOPs across the selected models.

As context approaches 1M tokens, this contribution can become dominant, reaching 98% for .

Parameter-bound FLOPs grow linearly with the number of tokens, while full-attention FLOPs

grow quadratically with sequence length. As a result, the FLOP mix shifts increasingly

toward pure attention as context grows, especially in architectures with many full-attention

layers.

Across models, the total fraction of FLOPs spent in the sequence mixer at 16K sequence length is more constrained, ranging from roughly 33% to 65%. As pure-attention

FLOPs are reduced through e.g. sparse attention, projection linears account for a larger fraction of total FLOPs. follows this pattern by reducing the number of full-attention layers and replacing them with

sliding window attention, substantially lowering the pure-attention share from 51% to 27% at 16K

sequence length compared to .

FLOP Scaling vs. Context Length

As Figure 4 shows, differences in compute cost become increasingly

pronounced as context length grows. has the highest FLOPs across the full range, combining a large active model size with 24 full-attention

layers out of 93. sits at the opposite end at long context, with only a small fraction of its sequence mixer scaling

quadratically with context length.

and need

similar FLOPs at 16K context, but at 1M tokens MiMo-V2.5 needs about 5× more: its full-attention

layers attend to every previous token, while DeepSeek V4.1-Flash attends to a fixed number of

selected, compressed tokens and scores the full prefix only with a light indexer.

follows the same

trend. Compared with , it uses sliding-window attention and roughly 5× fewer full-attention layers, so the

compute advantage widens as context length increases.

reaches the lowest FLOPs/token at long contexts in this comparison. Additionally, together with

it is the flattest curve over context length, with only 13% and 33% increase respectively from

4K to 1M.

Multiplying FLOPs by the token horizon D gives us the total training FLOPs, Ctraining≈D×Ftrain/ token 8 and dividing by your achieved training

FLOP/s gives an estimate of your training duration. Consider that the real training cost is far

more complex than just FLOPs. Some of the other major factors, not addressed here, are the overlap

of computation and communication, the quantisation of your weights, kernel efficiency and kernel

launch overhead, activation recomputation and expert imbalance.

FLOPs are a useful proxy for prefill cost, but they do not directly translate into faster

decoding, where memory movement becomes more important. This motivates the next section on

sequence-mixer state size.

Sequence-Mixer State Size

At every decode step, the new query must interact with the representations of earlier

tokens. Those key and value representations were already computed during the previous steps

and never change, so we can cache them instead of re-projecting, trading FLOPs for HBM

capacity and more I/O. The attention computation itself is unaffected, since at each step we

simply append the K and V vectors of the current token.

What limits decode speed depends on the regime: A decode step multiplies every weight matrix

with only one token per sequence, so at short context and small batch it is dominated by

streaming the model weights from HBM to the compute units (memory bound); for MoE, the

weights read are those of every expert selected by at least one sequence, which approach all

experts as the batch grows. At large batch and short context, weight reuse raises arithmetic

intensity and compute can become the limit (compute bound). Once contexts grow long, the

bottleneck is the I/O cost of reading the state for every generated token, which, unlike the

weights, is not shared between sequences (memory bound). These are only general regimes: the

limiting factor depends on the arithmetic intensity of the operation and on the hardware

specifications, in particular the bandwidth (byte/s) and the throughput (FLOP/s), both of

which also depend on the numerical precision. For a deeper explanation of these topics, read

our post Inference performance from first principles.

Mstate=c×B×b×dkv×Lkv×nret

Cached Representation

Standard attention stores separate key and value vectors (c=2). Multi-head

latent attention (MLA, introduced in DeepSeek-V2) instead caches a single vector per token (c=1): a low-rank content latent

plus the positional component required by its decoupled RoPE path.

Batch

In a naive implementation every sequence of the batch B has its own independent state.

Prefix caching instead reuses a shared

prefix (e.g. the system prompt) across sequences, so one copy of that state serves many requests.

Bytes per Element

Values in the cache are commonly stored in 1 byte (FP8) or 2 bytes (BF16). Methods such as KIVI and KVQuant go down to 2–4 bits, but

are not yet common in production serving. Recurrent states are often kept in FP32 instead.

KV Width

The number of KV heads times the head dimension sets dkv. GQA shares one KV head across a group of query heads, while MQA uses a single KV head for all queries.

Attention Layers

By default every attention layer keeps its own state. Hybrids replace most full-attention

layers with sliding-window or recurrent layers, and research directions such as

cross-layer attention (CLA) let

groups of layers share one layer’s K and V. DeepSeek V4.1-Flash, for example, computes its

compressed cache in four layers and lets the layers in between reuse it.

Retained Tokens

Full attention retains the entire prefix. A common and robust variation is the sliding window, which keeps only the newest w tokens. Selective attention, such as NSA, reads only a sparse subset of the prefix at each step. This cuts the bytes read per

step, but not the bytes stored, since any block may be selected later.

Recurrent Layers

Mamba-2 and Gated DeltaNet layers fold the whole prefix into a fixed state per head, dk×dv elements for

GDN and dh×dstate for Mamba-2, plus the last few inputs of a short convolution.

A Qwen3.8-Flash-Next GDN layer keeps 48 heads of 128×128, 0.79M elements, as

many as one of its attention layers caches (K, V and indexer key) for about 680 tokens,

but the same at any context length.

Sequence-Mixer State Across Architectures

Figure 5 shows the total sequence-mixer state and the amount

of state read at each decode step, which capture two different constraints.9 Total state size determines how many sequences and how much context can fit in a decode batch,

while the state read per token is more directly tied to decode speed in the long-context, memory-bound

regime.

This distinction is particularly important for sparse attention. Many sparse mechanisms

reduce the number of tokens accessed at each step, lowering memory traffic and FLOPs,

without reducing the total state that must be stored. At 1M tokens, holds 72B elements per sequence but reads 11B per decode step. Sliding-window and recurrent

layers reduce both, while techniques such as KV compression or cross-layer sharing can also shrink

the stored state itself.

This produces a substantially different ranking from FLOPs. , for example, has lower long-context FLOPs than , while the latter achieves a smaller sequence-mixer state through KV compression and

reuse. and

DeepSeek V4.1-Flash similarly occupy comparable parameter scales but very different points in

terms of state size: at 1M tokens, 72B elements per sequence for MiniMax M3 against 1.7B for DeepSeek

V4.1-Flash.

also improves substantially

over by replacing most full-attention layers with sliding-window attention, reducing the state that

grows with context length: at 1M tokens it holds 10.8B elements per sequence, 5× less than Kolibri

Origin’s 53.7B.

Conclusion

In this blogpost, we present our Model Explorer, used for designing the Kolibri

architecture. We analyse how parameters and FLOPs are allocated across the main model

components for popular open models and how the sequence-mixer state grows with sequence

length under different architectures. Additionally, we derive formulas for FLOPs and

parameters for our Kolibri models. We are excited to share this with the community, so

others can explore architecture trade-offs and reason about architecture from first

principles.

Acknowledgements

We are grateful to Ahmed Hammam, Fabien Benureau, Jonas Knupp, Simon Thel, Thomas Burns and

Yasser Jadidi for their thoughtful review. We thank Alexander Wortmeier and Noé Beckerle

Vallejo for their support bringing it to our website and Helena Treeck and Svenja Fahlisch

for keeping everything on track.