For the last three years, the miracle of the LLM has been bundling. Need to extract something from a document? Call the LLM. Want to search for information, rank five options, decide whether something looks suspicious, choose the next tool, browse a website, write the response, verify the work, or solve a genuinely difficult reasoning problem? Call the LLM. GPT, Claude, Gemini and Kimi collapsed an extraordinary number of previously separate capabilities into one general-purpose product. That was one of the great product breakthroughs of generative AI: instead of stitching together a dozen narrow systems, developers could put one sufficiently intelligent model in the middle of an application and let it do almost everything.
That convenience has hidden something important. These capabilities have radically different computational requirements, and different models are already better at different parts of the bundle. As the variance in performance, latency, and cost across those capabilities grows, the economically rational architecture changes.
I think we are entering the Great Unbundling of Intelligence.
Why now? Agents turn one human objective into hundreds or thousands of machine decisions, so small differences in the cost and latency of each step compound into the economics of the entire product. Inference has become real cost of goods sold increasingly making cost and ROI a consideration. Smaller and open-weight models are good enough to absorb much more routine work. Search, retrieval, reranking, memory, and computer use are becoming separate infrastructure layers. Applications have better evals and can actually tell when a cheaper system is good enough. And paradoxically, frontier models becoming much better makes unbundling safer: an application can aggressively use cheap intelligence because a frontier model is waiting at the top of the stack whenever the cheaper layer won’t work. The architecture starts flipping from frontier by default and optimize later to cheap by default, frontier on exception. With the massive inference scale of agents,different steps inside the same objective deserve radically different amounts and kinds of intelligence.
That is why Jev arrived at exactly the right moment. TypeSafe pulled one common capability out of the generative bundle: judgment. Give Jev messy state, a bounded question, and defined possible answers, and it returns structured decisions and probabilities instead of open-ended text. The underlying idea is not new. What is new is packaging one cognitive operation directly at the model layer at a cost and speed suited to an agent’s inner loop. If the job is simply “look at this and decide,” the model no longer needs to generate a sentence first. Jev arrived exactly when agents turned it into an economic problem.
A useful way to understand this is the decoder tax. If software only needs urgent / not urgent, continue / stop, allowed / blocked, or which tool, today we often ask a language model to generate the answer and then convert the output back into the decision the application wanted. Agents do this constantly. A surprising amount of their inner loop may not need generation at all.
Now consider what that does to the economics of the LLM bundle. Anthropic reportedly has gross margins above 80%. That is what a powerful bundle looks like. Claude can write, classify, extract, rank, judge, code, plan, verify, and reason through a single product. Today a simple classification or verification task costs frontier-level pricing simply because it happens inside a Claude call. Unbundling asks a much more uncomfortable question: what is the independent market price of every capability hiding inside that call? Difficult reasoning may deserve an enormous premium. But should deciding whether an email is urgent? Should ranking five retrieved documents? Should checking whether a browser action succeeded? If judgment gets competed toward fractions of a cent, ranking moves to rerankers, search to search providers, exact work to code, and ordinary generation to cheaper or open-weight models, applications do not need to attack Anthropic’s margin directly. They can attack the mix of work that raises Anthropic margins.
That creates adverse selection for frontier inference. Today the frontier model receives a blended workload. As the application gets better at routing, it systematically strips away the cheap, predictable, high-volume work. Browser models take repetitive computer control. Code takes deterministic reasoning. Small and open-weight models take ordinary generation. The frontier lab increasingly receives the highest order work: strange, ambiguous, long-horizon problems where cheaper systems are uncertain. Frontier intelligence could therefore become simultaneously more capable and more valuable per invocation, but less frequently invoked. And there is a potential double punch. Not only could fewer operations reach the frontier model; the average operation that does reach it may require more test-time compute because routing has preferentially selected the hardest problems. The threat to frontier inference isn't just cheaper frontier inference. It is the decomposition of the workload underneath it.
Once capabilities become separable, somebody has to stitch them back together. That is why the application layer becomes more important. Instead of asking which model should handle an entire task, the application decides which capabilities the task requires. The application becomes an intelligence compiler: take a human objective, break it into cognitive operations, buy the cheapest sufficient intelligence for each one, recombine the result, and escalate only when necessary.
The next step is one layer deeper than asking which model should answer this prompt? It is asking who should supply each capability inside this task? Once a capability has a stable interface, providers can be benchmarked, swapped, routed, and priced independently. The Great Unbundling does not eliminate the AI margin pool. It forces every capability to prove that it deserves one.
And this changes more than margins. Cheap judgment changes which products are economically possible. At a fraction of a cent and a few hundred milliseconds, you can put a judge after every agent action instead of checking work occasionally. You can put intelligence inside browser and voice loops where multi-second inference is too slow. And you can move from sampling 1% of a system to evaluating 100% of it like every support conversation, retrieval, transaction, claim, contract clause, or customer account. It creates continuous supervision, continuous QA, and entirely new product loops and use cases.
Most decisions software could make today are never made because intelligence is still too expensive to apply continuously. Once judgment, ranking, verification, and other capabilities become cheap enough, software can evaluate every customer, every workflow, every agent action, and every change in state.
Cheap intelligence expands the surface area over which software can afford to be intelligent. The real unlock is a world in which intelligence stops being an event and becomes a background property of technology. What if all products can afford to think about everything for close to free? We spent the last several years asking how many capabilities we could stuff into a single model. We may spend the next decade pulling them apart and recomposing them at the application layer.