Skip to content

Science3 publishers3 min readPublished

A jazz model names the player 94% of the time, and says which part of the playing gave them away

Cambridge researchers split melody, harmony, rhythm and dynamics into separate model inputs, then identified 20 pianists from 84 hours of recordings. The code and a web app are public.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Three University of Cambridge researchers used AI in a new paper in Nature Machine Intelligence to better define the musical tics of 20 of the most celebrated jazz pianists of all time.
  • The researchers trained a variety of supervised learning models to identify 20 iconic jazz musicians from a curated dataset of 84 hours of recordings.
  • The best model obtains 94% accuracy across 20 classes.
  • The paper introduces a multi-input architecture that represents four musical domains separately (melody, harmony, rhythm and dynamics), a design that allows accurate identification of individual performers and examination of which musical elements most strongly distinguish between individual artists.
  • The authors release open-source implementations of their models and an accompanying web application for exploring their results.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

Three University of Cambridge researchers have published a paper in Nature Machine Intelligence that uses machine learning to characterise the playing habits of 20 of the most celebrated jazz pianists [1]. Their best model identifies the performer with 94% accuracy across 20 classes [3], which moves "style" out of the register of criticism and into the register of measurement, with the obvious downstream uses in authorship attribution, education, cultural heritage research and historical analysis [12].

The training set is 84 hours of recordings [2], comprising 1,629 performances by players including Chick Corea, Thelonious Monk, Oscar Peterson and Bill Evans [6]. That averages roughly 3.1 minutes per performance [10]. Audio was converted to MIDI, a format that records which notes were played, when, and where on the keyboard [7]. So this is symbolic analysis of note events, not raw waveform timbre.

The architectural choice is the interesting part. Instead of one undifferentiated input, the authors built a multi-input model that represents melody, harmony, rhythm and dynamics separately, which lets them both classify performers and inspect which musical elements do the discriminating [4]. That is a direct response to a problem the paper names: deep networks already identify visual artists, writers, composers and performers with high accuracy, but they are hard to interpret, which limits their value for explaining creative process or teaching, so the current challenge is accuracy and interpretability together [14].

For scale: with 20 balanced classes, guessing yields 5%, so 94% is about 18.8 times chance [9]. A second model trained on more granular features, including the specific notes each pianist used to build melodies and chords, scored 91% [8]. The gap looks small and is not: 6% error against 9% error means the granular model makes half again as many mistakes [11].

The case for bothering is partly that the cues in question are not all audible. Tiny deviations in rhythmic timing or in which beat gets emphasised can escape notice entirely and may never appear in sheet music [15], which is precisely why manual stylistic analysis has been slow, unscalable and confined to a short list of canonical figures such as Beethoven, Cervantes, Dante and Raphael [13]. Jazz improvisation is a reasonable stress test because it fuses composition, improvisation and performance in one artifact [16].

On release: the authors say they are publishing open-source implementations of the models plus a web application for exploring the results [5], and according to Scientific American that app provides graphs of each player's tendencies alongside audio samples illustrating their use of melody, harmony, rhythm and dynamics [17]. Anyone planning to reuse the corpus rather than the code should read the paper's availability statements carefully; what the abstract commits to is implementations and the app, over a dataset it describes as curated [2][5].

Two things to watch. First, generalisation beyond a closed set: the stated practical payoffs are suggesting authors for unattributed works and separating authentic pieces from forgeries [19], and a 20-way classifier that must pick one of 20 known pianists is not yet a system that can say "none of these." Second, whether the per-domain attributions survive expert scrutiny, since the pedagogical claim, that this reveals influence between artists and can train young musicians [18], rests on the explanations being right rather than merely being available.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories