LibraryDesignBench is a two-phase benchmark that evaluates how well agent-designed software libraries help other agents write code. The benchmark has agents design libraries from vague specifications, then measures quality by observing how effectively different agents use those libraries to solve problems, scoring based on correctness and code simplicity metrics.