We are fast approaching a point where the primary users of software libraries are agents, not human engineers. If agents are the users, then a library should be judged by how little code agents need to write correct programs with it, not by how it reads to a human. That is why I am excited to release LibraryDesignBench, a two-phase benchmark that scores an agent-written library solely by how much it helps the future agents that use it.

Asking agents to write libraries for other agents

LibraryDesignBench asks agents to write effective libraries, and we intentionally give them design flexibility. We never prescribe signatures their library must expose. Nor do we ever guide them to specific abstractions or patterns they should use. We intentionally give them underspecified, vague, and ambiguous library specifications because frontier agents need to envision how future agents will use their library, what they will need, and what design will work best. LibraryDesignBench gives agents this vast freedom because the evaluation needs to serve as a test bed for understanding what patterns agents actually prefer, not what we think they will prefer.

Evaluating A Library Based on How Agents Use It

A library that implements something correctly does not immediately provide any value – it needs to be correct and make future code simpler. Thus, the only way to measure a library’s quality is to observe how much future agents benefit from its design decisions. LibraryDesignBench does this in two phases:

- Design Phase: The agent implements a full library from an intentionally non-prescriptive specification.

- Evaluation Phase: We evaluate the library by observing multiple different implementer agents attempt to solve problems using it.

The only aspect we check in the design phase is if the library is installable in their internet-restricted environment, as we pre-install all libraries for evaluation since trying to install a library is not part of the signal we care about. We then task three agents (GPT-5.6 Luna on Codex, DeepSeek v4.1 Flash on mini-SWE-agent, and GLM 5.3 Flash on mini-SWE-agent) with solving programming problems with the library in as little code as possible. Each problem agent library is run in its own environment without internet access, with high reasoning, and with the library pre-installed. We use a highly prescriptive prompt to ensure agents attempt to fully exploit the library.

We now take these solutions and score each with:

Pass rate is the percentage of tests passed, and we square it to penalize incorrectness more harshly than simplicity. Simplicity is defined as:

is the set of static metrics we use to compare how far off the solution written with the library is from the golden reference, , written with the real production library. Each ratio is capped at 1, so a solution that beats the reference gets no extra credit. The four metrics are:

- Source Lines of Code: how much code the agent wrote.

- Cyclomatic Complexity: how many branches and loops the agent needed.

- Cognitive Complexity: how hard the code is to follow, with extra weight on nesting.

- Halstead Volume: size in operators and operands, which dense one-liners can’t game.

Section 2 of the paper covers the scoring and standard errors in more detail.

Agents Copy Human Designs But Worse

We evaluate 11 designer setups, covering 9 frontier models with some run in more than one harness, all at high reasoning. Each setup attempts each of the 15 library design tasks three times, which produces 45 libraries. The three implementer agents then use each library to solve every problem in its task. Across all 242 problems, that comes to 242 problems × 3 libraries × 3 implementers = 2,178 trials per setup.

We compare against two baselines that use the same 2,178 trials. In the first, implementers have no library and use a separate prompt. In the second, they have the human-written production library pre-installed. Neither baseline has three designed libraries to vary, so we instead run each problem-implementer pair three times.

Opus 5.5 designs libraries that help downstream agents more than the human-written production libraries do. With them, implementers pass just as many tests while writing simpler code. Fable 5.1 roughly matches production. Correctness barely separates designers, since every setup passes about the same share of tests, so the ranking comes down to how much code implementers still have to write. At the other end, DeepSeek V4 Pro’s library actively hurts: implementers do worse with it than with no library at all. The harness also matters. Fable’s libraries are more useful when designed in mini-SWE-agent than in Claude Code, and Astra’s are slightly more useful in mini-SWE-agent than in Codex.

Beyond harness peculiarities, the biggest trend we observed is that on 11/15 tasks, agents copied a design pattern from the corresponding production library. Surprisingly, these patterns are the exact ones the implementer agents used when given the production library. Our failure analysis indicates that the implementer agents write more complex code than needed because interfaces are either too rigid to use or require verbose code. In the former case, agents reimplement functionality the libraries provide, adding bugs in the process.

Can Agents Use Libraries Without Instruction?

Our main evaluation prompt is rather heavy-handed because we need agents to fully exploit the library to measure its upper bound. But we also want to understand how models will behave given minimal additional instructions. Thus, we also release LibraryUseBench, our evaluation phase in which agents must use the production library to write as little code as possible. All agents use mini-SWE-agent and use High thinking. Opus 5.5, unsurprisingly, is the best model evaluated, but Sonnet 5.5 is close behind at just over half the cost per problem. The gap between them is small, but it holds up when we compare the two problem by problem. GPT-6 Luna is by far the cheapest model, but it passes only 71% of tests, and GPT-5.6 Luna rounds out the lower end.

Do Agent-First Designs Help?

Designers tend to copy the production library, so we tried steering them away from it. We append a guidance prompt to the spec. It says the library will be used only by coding agents scored on how little code they write, and it asks the designer to:

- write the programs downstream agents would want to write first, then design the API to match them,

- have fresh subagents solve tasks with the library, and fix the library wherever they write extra code or guess a name wrong,

- ship those programs as runnable examples.

Guided libraries copy fewer names from the production library and fold multi-step patterns into single calls. In clirs, for example, builder chains become one task-level call. Downstream solutions get shorter, and the score improves for every implementer, but it still lands just below the production library at roughly twice the design cost. Moving away from human designs helps, but we still haven’t found what the design agents prefer.

Conclusion

The full leaderboards and tasks are at ldbench.com. You can view the full repo here and the task repo here. You can read the full technical report here. If you have a task you would like to see, open an issue on the repo or get involved in the Discord.

This would not be possible without my amazing collaborators: Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, and Ludwig Schmidt. I also want to thank John Yang, Parth Asawa, Xavier Garcia, Ryan Carelli, Arun Kumar, Floriad Brand, and Nick Roberts for their helpful feedback and discussions. LDB is supported by DARPA, NSF, Prime Intellect, and Snorkel AI through the Open Benchmarks Grant.