Product1 publisher3 min readPublished
Sony and Warner Music Sue Anthropic Over Alleged Copyright Piracy
Sony Music and Warner have sued Anthropic over pirated music and lyrics, joining Universal. For publishers the live question is which crawlers still get the full text, and what evidence would show a training promise held.
The Product Desk · Product desk

What happened
- Sony Music and Warner sued Anthropic this week over copyright, accusing the company of illicitly pirating music and song lyrics from their catalog to train its AI models.
- Universal brought a similar action in January, so Anthropic is now the target of all three major music publishers.
- Anthropic agreed a year ago to pay $1.5 billion in a settlement over pirated books.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint Blocking the article costs citations: an answer engine then works from metadata alone and tends to cite the competitor who stayed open.
- decision The call moves down from the site to the individual crawler line, and each one needs a second answer alongside allow or deny: what evidence would show the operator kept the copy.
- exposure The conduct alleged here happened outside any publisher's site, so a media company's own logs and crawler rules are not the place the risk shows up, and the plaintiffs with standing are the catalog owners.
- precedent A $1.5 billion settlement already exists as a reference point for rights holders negotiating with Anthropic, and three music plaintiffs are now in front of the same defendant.
The person who has to act on this is whoever owns the site's robots.txt. Pete Pachal, who writes the Media Copilot newsletter, wrote in Fast Company: "I see this more and more in my consulting work: the instinct to protect IP makes publishers reluctant to even do GEO testing." [7]
That file, the Robots Exclusion Protocol, sets which bots may scrape a site [8]. Two kinds matter here. Training bots harvest content into archives used to build models and keep the copy; retrieval bots pull specific information to answer one query in real time and use it once [9]. The search crawlers behind discovery do keep an index, which Pachal called "a card catalog, not a model" [10].
Blocking training bots absent a licensing deal is close to standard among publishers now [11]. Retrieval is the decision that costs something. Block the article and the answer engine has only metadata to work from, and a competitor who allowed the crawler is likely to be the one the engine favours [12]. Pachal also wrote that turning that visibility into good business outcomes is "far from guaranteed" [16].
His recommendation is to block training crawlers, selectively allow retrieval crawlers where AI visibility matters, and then build your own verification [13]. The column ends at that colon, leaving the steps unstated [14].
The music case shows where that recommendation stops short. Sony and Warner allege torrenting of catalogs and wholesale copying of lyrics from third-party websites [3], and the filing says co-founder Benjamin Mann personally conducted or directed the torrenting and discussed it openly in Slack channels [4]. A robots.txt line governs crawlers on your own site [8]. It does not govern a torrent, and your own server logs would not have recorded one.
So verification splits into two jobs of very different difficulty. The one a publisher can actually do is watch what happens to text it allowed for retrieval, since the fear Pachal describes is that the company keeps the article and uses it for training or lets users read the full text [19]. The second job, auditing how a model was built, is out of reach. Anthropic agreed a year ago to pay $1.5 billion to settle claims over pirated books [5], and on this suit told Axios: "we intend to defend ourselves robustly in court" [6]. That settlement came roughly a year before the Sony and Warner filing, with Universal's January action falling in between [17].
The column is written for publishers deciding what to expose, and the question of what a company building products on these models should ask a vendor about training data sits outside its scope.
For Monday, two things to establish for each crawler line in the file. First, whether letting that bot in puts the full text in front of a reader you can count later. Second, what artifact would surface if the operator kept the copy: a logged fetch pattern, a verbatim passage in an answer, a paywalled paragraph appearing in a chat window. When you can only answer the first, the permission runs on trust, and it can still be the right trade where the answer slot is worth more than the exposure.
What to watch
- Whether Anthropic's answer to the Sony and Warner complaint contests the torrenting and Slack allegations against Benjamin Mann.
- Whether any retrieval crawler operator starts publishing checkable evidence of what it did with a fetched article, instead of a policy page.
- Whether the three music publishers' cases consolidate, and whether any settles near the $1.5 billion books figure.