Skip to content

Build1 publisher3 min readPublished

The harness, not the model: 250 scores were the acceptance spec for a solo MusicXML editor

A developer let Claude Code write most of a collaborative sheet music editor and kept the definition of "correct" for himself. The transferable artifact is the round-trip corpus.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying The harness, not the model: 250 scores were the acceptance spec for a solo MusicXML editor
Photo: dev.to

What happened

  • ScoreTail is a browser-based score editor built mostly alone by its author with a lot of AI-assisted development.
  • The stated goal was a "Google Docs for sheet music": open a browser tab, start writing notation, share a URL, and have someone else edit the same score in real time; the author says nothing like it existed.
  • MusicXML is described as a huge spec covering pitch, duration, voices, staves, backup/forward, divisions, ties, slurs, ornaments, tuplets, lyrics, chord symbols, dynamics and repeat structures.
  • The author states that implementing all of MusicXML solo, the traditional way, is a multi-year project.
  • Late last year the author decided to let AI write most of the implementation (Claude Code about 90 percent of the code, Antigravity about 10 percent) while he focused on defining what "correct" means through architecture decisions and test cases rather than typing every line.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer publishing as tan-z-tan has released ScoreTail, a browser-based collaborative sheet music editor built largely solo, with Claude Code producing roughly 90 percent of the implementation and Antigravity about 10 percent [1][6]. The part of that worth copying is not the tool choice: by the author's own account, what made the project viable was a corpus of about 250 real-world scores wired into a round-trip test [7][8].

The scope explains why. MusicXML covers pitch, duration, voices, staves, backup and forward, divisions, ties, slurs, ornaments, tuplets, lyrics, chord symbols, dynamics and repeat structures, and the author's estimate for implementing all of it alone the traditional way is a multi-year project [4][5]. The bet made late last year was to let the model do the typing while the human defined correctness through architecture decisions and test cases [6].

The harness is cheap for a specific reason: nobody has to write down the right answer. Each score is imported, exported, and imported again, and any structural drift between the first and second import is treated as a bug [9]. The pass condition references only the input file, so none of the roughly 250 scores needs a hand-authored expected output [12]. That is what let the corpus be assembled from things that already exist: Bach, Beethoven, Chopin, Brahms, Debussy, plus the LilyPond regression test suite [8]. Import, export, re-import means three crossings of the format boundary per file, on the order of 750 per full run [11]. The loop is then mechanical: the model writes code, the round-trip suite catches regressions, the model fixes them, repeat, with humans reserved for the judgment calls [10].

It matters where the bugs actually live. The author names backup and forward as the hardest part of MusicXML: the format is a flat, sequential XML stream, but a measure is not sequential, because it carries simultaneous voices and, in a piano part, two staves sharing one measure, so voice A's notes are written and then a backup element rewinds the internal time cursor before voice B starts from the same beat [13]. Any reader or writer therefore has to keep its own timeline position independent of document order, and the author says the combinations of backup with multi-voice and multi-staff are where most real-world parser bugs sit [14]. A generated unit test would not have found those; a Chopin file does.

The limits of the approach are visible in the rest of the design. Yjs gives structural convergence across clients but has no model of a musically valid tree, which the author treats as two separate problems [20], and the fallback is an auto-undo that rolls back the last operation when an export produces an invalid document [21]. The editing rule is correctness over flexibility: if beat 4 of a 4/4 measure holds a quarter note, changing it to a half note has to overflow into the next measure, auto-insert a tie, or be rejected outright [15], because the target user is not assumed to know notation [17]. Those are acceptance decisions too, and they are the ones the round-trip suite cannot make.

Two things to watch. Rendering goes through Verovio, a C++ engraving engine compiled to WebAssembly, and full-score re-rendering on every edit blocks the main thread on large scores; the source cuts off mid-sentence on the fix [22][23]. And the post reports no pass rate, runtime, or failure count for the 250-file suite [24], which is the number that would tell you whether the harness is an acceptance gate or a progress bar.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories