Skip to content

Build1 publisher2 min readPublished Updated

Snorkel AI's LibraryDesignBench grades agent-written libraries by the code other agents write with them

Snorkel AI's LibraryDesignBench puts the best agent-designed library run at 48.9, ahead of 46.6 for human-written production libraries. Because the score also pays for shorter downstream code, the lead measures how easily other agents could use each interface.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Snorkel AI's LibraryDesignBench grades agent-written libraries by the code other agents write with them
Generated illustration

What happened

  • Gabriel Orlanski, a Snorkel AI research fellow and University of Wisconsin-Madison Ph.D. student, submitted the LibraryDesignBench paper on September 29.
  • A designer agent works from an open-ended specification and a few example uses, with no required interface, method signatures or tests to follow.
  • The test set covers 15 library-design tasks and 242 problems across Rust, Python, TypeScript and Haskell, modelled on libraries such as clap and pandas.
  • In Haskell, an agent-written library scored below having no library at all in 70% of designer-task pairs.
  • The authors report that agent-focused design guidance, and letting designers test their work with subagents, both raised downstream scores and produced simpler programs.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Checking an agent-written library against its spec will miss the failure the paper found most often: the capabilities were present and the calling agents still worked around them.
  • exposure On this data, a Haskell team that accepts agent-designed abstractions is more likely than not to leave its calling agents worse off than with no library.
  • capability Both improvements act on the designer, so a team can try them without retraining or swapping the agents that consume the library.

Each downstream solution earns credit twice. It scores for passing hidden tests, and again for simplicity measured against a reference solution written with the established production library [5]. The published figure is a 0-to-100 composite, not a pass percentage [5]. So a designer can raise its score by shortening the implementers' code while their pass rate stays flat [5].

The top run's pass rate is 86.6% [8]. Having any library at all is worth 12.2 points over the no-library reference of 34.4 [2][8]. The best agent design adds 2.3 on top of the human library [1]. The source does not give the human reference's pass rate, so those 2.3 points cannot be split between correctness and brevity. Runtimewire, which reported the results, wrote that they do not mean an AI-generated library is generally better than a human-built one [13].

I would put the consumer result first. Agent designers reproduced the human library's abstractions in 11 of the 15 tasks, about 73% [6][3]. The downstream agents still underused what they were given. They reimplemented capabilities the library already had [7]. The paper's analysis attributes much of that added code to rigid or difficult-to-use interfaces, and not to missing features [7]. One listed run scores below the no-library reference [10].

A leaderboard position predicts a team's own results only if that team's setup looks like the benchmark's. Here the consumers are three fixed implementer agents [3]. The problem set averages about 16 problems per task [4]. The harness alone moves scores: the same model scores differently when paired with different coding-agent tools [11].

I think this is the right way to grade code whose main caller is another agent. The library's own source never enters the score. Only the code other agents write against it does [14][5]. Plenty of libraries look tidy in review and are awkward to call. In a codebase where agents write internal helpers for other agents to call, I would measure the same property.

What to watch

  • A published pass rate for the human-library reference, to separate correctness from brevity in the 2.3-point lead.
  • Runs with implementer agents other than the three fixed ones, to test whether the ranking survives a change of consumer.
  • Effect sizes for agent-focused guidance and subagent testing; the reporting so far gives only their direction.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories