If you work in AI infrastructure, you live with two facts that seem to contradict each other: there is a massive GPU shortage, yet the GPUs we have sit at roughly 30% utilization.

First, a definition. By utilization I mean share of throughput: the tokens a GPU actually serves, divided by the tokens it could serve. This isn't what nvidia-smi reports, or the fraction of time that the GPU is busy. An LLM decoding a single request keeps a GPU "busy" 100% of the time while producing a small fraction of the tokens it could. Providers are paid per token, so throughput is the number that matters.

The contradiction goes away once you look at how GPU fleets are sized. A provider has to buy enough GPUs for the daily peak, and traffic peaks at roughly twice the daily average. At peak, providers aim to run at about 70% of capacity and keep the rest as headroom so latency doesn't spike. So:

35% is the ceiling: the best any single-model deployment can average, before redundancy, rounding up to whole GPUs, or bad forecasts make it worse. The shortage is a shortage of peak capacity; the waste is in the average.

It's Not Just the Long Tail

Take a mid-sized model on OpenRouter: Qwen 3.8 27B. OpenRouter currently routes about 56B tokens a day to it (prompt + completion), split across 16 providers.

OpenRouter only reports tokens, not how many GPUs each provider runs, so we have to estimate. We express everything in B200 equivalents:

- Output volume. The weighted cache-hit rate is 63%, which is consistent with prompt-heavy, agentic traffic. At input-to-output ratios between 3:1 and 10:1, that's 14B to 5.1B output tokens a day, or 162k to 59k output tokens per second.

- Market size. At throughput-optimized batch sizes, one B200 serves roughly 5,000 output tokens per second. So all of this model's OpenRouter traffic would fill 12 to 32 B200s running flat out.

- Each provider's average demand is its token share times the market. Parasail's 18.7% share works out to 2.2 to 6.1 B200s.

- Capacity is sized for peak, with a floor of one GPU. We assume no redundancy, which is likely a more favorable assumption than the reality for most providers.

For Parasail, 2.2 GPUs of average demand needs 7 GPUs of capacity, so it runs at 31%. Ionstream's 1% share is 0.12 to 0.32 GPUs of demand on a single GPU, so it runs at 12–32%.

The seven largest providers carry about 87% of the traffic, and all of them sit at 29–35%, pressed against the ceiling, and their peak demand needs several GPUs. Weighted by tokens, the whole market runs at 28–33% utilization. The tail runs anywhere from 1% to 32%.

One caveat: some of these providers also serve traffic outside OpenRouter. That pushes their real utilization up, so for them this is a lower bound. For most providers though, OpenRouter is the main source of demand for this model.

The obvious fix would seem to be more demand: give a provider an order of magnitude more tokens and utilization should rise. But it doesn't.

Volume Does Not Fix It

Take a frontier model with real demand: Kimi K3. OpenRouter routes it about 210B tokens a day, nearly 4× Qwen 3.8 27B, split across 15 providers.

The difference is size. Kimi K3 has 2.8T parameters, about 1.4 TB of weights even at MXFP4. The smallest deployment that can serve it is a full node of 8 B300s, 2.3 TB of HBM, which leaves ~0.9 TB for KV cache. We price B300s at $7/GPU-hour, typical of 1–3 year reservations (on-demand pricing is far higher). Demand for a model this size comes in whole replicas, not individual GPUs.

The estimate follows the same steps as for Qwen:

- Output volume. The cache-hit rate is 85%. At input-to-output ratios between 3:1 and 10:1, that's 608k to 221k output tokens per second.

- Market size. We estimate one B300 serves ~1,000 output tokens per second (500–1,500). It has the same 8 TB/s of HBM bandwidth as a B200 and ~1.5× the FP4 compute. All of Kimi K3's OpenRouter traffic would fill 28 to 76 replicas running flat out.

- Provision for peak, now counting whole replicas and with a minimum of one. Together's 40.3% share averages 11 to 31 replicas and runs at ~35%.

The top providers sit at 31–35%, and the market as a whole runs at 30–34% weighted by tokens, which is the same ceiling as Qwen. The tail is much worse: DeepInfra's 0.1% share is 220 to 610 output tokens per second, on a replica that can serve roughly 8,000. That's 3–8% utilization. At DeepInfra's own list prices, that traffic earns $34–50 an hour, and the replica costs $56 an hour at reserved rates.

Stealing an Idea From Operating Systems

If one model can't fill a GPU, how about adding a second model? In theory, it's a no-brainer: the same work on half the GPUs. Utilization is additive in throughput share, so a model at 30% and another at 40% become one GPU at ~70%. The 35% ceiling applies to each model, not to the GPU.

A B200's 180 GB of HBM can hold the weights and KV cache of several mid-sized models at once. Small-active MoE models like Gemma 4 26B-A4B are cheap enough per token that it's hard to justify a whole GPU for one. And the long tail of open models keeps growing, with every new model needing somewhere to run.

So why doesn't everyone do this? There are several obstacles:

- Inference engines assume they own the GPU. By default, vLLM and SGLang reserve about 90% of GPU memory at startup for weights and KV cache. A second engine on the same device either fails to start or has almost nothing to work with. Getting past this takes a runtime designed from the start to host several models at once.

- Naive sharing is destructive. When two engines launch work on the same GPU with no coordination, their kernels contend for the same compute cores and memory bandwidth. In our tests on one B200, Gemma 4 dropped from 1,729 tok/s to 26 tok/s when an 8B embedding model ran alongside it. Models have to cooperatively hold and yield a "lease" on the GPU.

- Long jobs hog the device. An image model like Qwen-Image takes ~2.4 seconds per generation, while an LLM produces a token every ~4 milliseconds. If the image job holds the GPU until it finishes, every LLM stream stalls for the full 2.4 seconds. Yielding has to extremely fine grained.

- Not all GPU time is worth the same. At list price, a generated image earns ~5× more per GPU-second than LLM tokens do. Handing out leases evenly leaves money on the table, so the lease has to weigh each co-tenant by what its GPU time earns.

- Not every pair belongs together. Two models whose daily peaks coincide will overflow the GPU at peak (more on this worst case below). Placement has to pair models whose peaks fit within the GPU's capacity, and weigh what each earns per GPU-second.

- Models have to load and unload quickly. Demand shifts, models launch and fade, and yesterday's pairing stops making sense. Re-pairing only pays off if moving a model is fast relative to how long the demand shift lasts. If a model takes minutes to come up, you have to keep spare GPUs waiting, which is exactly the waste co-location removes. Loading also has to happen on a GPU that's already serving, so the models already running must give up memory without restarting or dropping traffic.

Follow this idea to its conclusion and you arrive at something that feels like cooperative multitasking, the scheduling model of early operating systems like Windows 3.x. Programs shared one processor by voluntarily yielding control at safe points. The difference is the stakes: when the shared processor costs $7 an hour, every model that yields well saves you a GPU.

Two Losing Deployments, One Profitable GPU

Take two mid-sized open LLMs: Qwen 3.8 27B and Gemma 4 26B-A4B. The assumptions:

- Hardware: one B200 at $7/GPU-hour.

- Prices: OpenRouter's market rates. Qwen averages $0.15 / $2.58 per million tokens across providers, weighted by token share (cache discounts included). Gemma's median is $0.10 / $0.335 per 1M.

- Traffic mix: 3:1 to 10:1 input-to-output for Qwen, and 3:1 up to our production mix of 12:1 for Gemma.

- Full load: our measured single-B200 throughput for each model at concurrency 16, and with speculative decoding: 2,070 tok/s for Qwen and 1,729 tok/s for Gemma.

- Demand: each model at 30% utilization, the typical top-provider figure from earlier.

Revenue is the same in both columns because the same users send the same tokens. Only the cost changes. On its own GPU, Qwen roughly breaks even (−$0.20 to +$2.21/hr), and Gemma loses $4.83–5.99/hr. Gemma's tokens are too cheap to pay for a GPU even at full load. Put them on the same GPU and the ~40% of capacity Qwen leaves idle carries Gemma for free. The pair goes from losing up to $6 an hour to earning up to $4.

This isn't hypothetical. Over the past fourteen days, we've been running exactly this pairing in production at inference.muna.ai, serving both models, along with two embedding models and an image generation model, from one B200 at a time. We served 2.5B tokens (prompt + completion), with Qwen streams running at ~135 tok/s @p50. The best part? None of our users could tell that the models were being co-located on a single GPU.

The Catch: Aligned Demand Peaks

The everyday cost is latency, not revenue. LLM batches are elastic. When one model waits for the GPU, its next step just runs a bigger batch, so throughput holds and each token arrives a little later. The slowdown grows roughly as 1/(1−s), where s is the other model's share of the GPU. In our benchmark, Qwen's time per output token rose from 3.93 ms to 4.54 ms (+16%) while sharing the GPU with an image model. The model with more lease weight gets its turn sooner, so the leasing system decides who absorbs the delay. We favor Qwen, the higher-value model: in production, Qwen waits on the lease ~5% of the time, and Gemma waits ~35% (88% @p95).

The worst case is aligned peaks. Revenue is demand served, moment by moment. We can model each model's daily traffic as a smooth wave that peaks at twice its average. If two models each average 30% of the GPU and peak at the same hour, their combined demand swings around 60% and hits 120% at the peak. For about seven hours around the peak, demand exceeds what the GPU can serve. The traffic above that line is ~7% of the day's total, worth $0.30–1.25 an hour, against the $7 an hour saved by dropping a GPU. Even this worst case beats two dedicated GPUs by $5–7 an hour. But the overflow grows quickly: at 30% + 40%, demand is over the line for nine hours, and 14% of traffic overflows.

So the question isn't whether to co-locate, but which models go together. The constraint is on the peak of the combined demand, not the sum of each model's peak:

where d_i(t) is model i's demand at time t, as a share of the GPU's throughput. Two models whose peaks line up can still share a GPU if their combined peak fits. Two models that each look small can overflow if they spike together.

This is a data problem, and an inference provider already has the data. It has demand histories for every model, by hour and by day of week. It knows how much of a GPU each model uses per token served, and how much memory its weights and KV cache take. And it has the catalog of models it could place. Together, these turn placement into a packing problem over time-varying demand. Find the sets of models whose combined curve stays under the line, and pack as many models as fit onto each GPU. Demand shifts as models launch, trend, and fade, so placement should keep re-solving itself. And fast cold starts are what make this possible: when a model loads in seconds, moving it to a different GPU is a scheduling decision, not a migration.

Most GPU-hours in inference are paid for, not used. Sizing for the peak caps any single-model deployment at ~35% utilization. That ceiling doesn't move with volume; it stays fixed across a mid-sized model to a 3T parameter frontier model with large, consistent demand.

At Muna, we are building a runtime that dynamically places many models on one GPU. We have proven that it works for mid-sized models, across LLMs, embeddings, and image generation models. Now, we are working on extending co-location to frontier models that span multi-GPU nodes, where every GPU must yield in lockstep. We have built an entire inference serving stack to power this, starting from a general-purpose Python compiler, to an open-source inference server.

The next 3× in inference won't come from faster kernels. It will come from the GPU-hours we're already paying for.