Product1 distinct publisher3 min readPublished
Cara filters AI images and ships Glaze, and neither tool reaches the moment a stranger downloads 12 terabytes of it and posts the archive to Reddit. Three scrapes in August also pushed up the hosting bill.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
Cara's team learned about the dump from its own users, who tagged the account because the scraper was, as Jingna Zhang described it to WIRED, gloating on Reddit and looking for other people to join him in doing something with the dataset [8]. Detection ran through a subreddit. Nothing in a stream of requests for public pages announces that the client on the other end is assembling a training set, which is why the alarm came from outside the system rather than inside it.
The two features Cara advertises sit at either end of that request without touching it. The AI filter governs what gets into the library, and Glaze works on the pixels themselves, masking style so a model that trains on the file has a harder time mimicking it [4]. Neither is in the path when a client asks for a public page and keeps what the server sends, and WIRED reports that preventing scrapes is nearly impossible [5].
The arithmetic on the transfer is worth doing on paper. Twelve million works inside 12 terabytes averages about a megabyte an image [1], and the under-$10 figure the scraper claimed comes to roughly 83 cents per terabyte on his side of the wire [2]. The second dataset covered about 71 percent of the library by count [3]. The third covered about 1 percent [4], and the small one is the one that carried user bios.
Then there is what artists did, which is the number a product team should actually look at. Some deleted their portfolios and abandoned the site, and Zhang says that if it makes them feel better, she supports it [16]. These are people who left platforms like Instagram, where content is explicitly available to Meta as training data [18], for an app whose product was a promise about consent. When the promise failed, they removed the asset. The deletions, not any session-depth chart, are what recorded that shift.
Every remedy in the story arrives after the copy exists. Zhang's legal GoFundMe was past 83 percent of its $120,000 goal by Thursday [15][5], she is a plaintiff in two artist class actions, one against Stability AI, Midjourney and others and one against Google [10], and the man who took the whole library now regrets it and is collaborating with her on an open-source protection tool [11]. Meanwhile the laws have not caught up, which is how a scraper gets to call this technically legal [9].
The promises on a marketing page tend to sort by where each one is actually enforced. Intake you control outright. Serve time you control at a price, which is exactly what Zhang is naming when she says the team has done the right things within limits without making the app horrible to use, and that temporary login gates are not really a solution to an internet-wide problem [17]. Everything downstream of a successful request is a request you are making of strangers. Zhang's own line is that Cara has almost certainly been scraped before and cannot guarantee complete security [19], and that is a far better sentence to have published in advance than to be explaining afterwards.
Ranked by verification strength, evidence, and original report placement.
Cara is an image-sharing social media and portfolio app that photographer Jingna Zhang and a small crew of volunteers have maintained since early 2023.
Cara has attracted about 1.5 million artists, drawn by shared opposition to the unauthorized use of their work to train AI models.
Cara filters out AI images and offers protective features including Glaze, a tool meant to mask the style of images picked up by scrapers in order to disrupt AI mimicry.
WIRED reports that preventing the scrapes themselves is nearly impossible for Cara.
Beginning on August 13, Cara was subjected to three major scrapes, which spiked its server fees and alarmed creators who had migrated there from platforms like Instagram.
A redditor using the name MandarinDawnPoppy994 posted a 12-terabyte archive of 12 million works from Cara, more or less its entire library of publicly available images, on the subreddit r/DefendingAIArt, writing in the since-deleted post that it was a fun project and that the process cost him less than $10.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
The labs got better at watching their agents escape. They did not get better at stopping them.1 distinct publisher
product
A billion downloads, and nobody will say what a download is1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
build
1.5% of Hugging Face repos take 99.2% of downloads, and the ceiling is Chinese1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, two interviews, few receipts
The events are specific and dated — August 13, August 22, 12 million works, 8.5 million links, 123,000 images — but every one of them reaches us through a single WIRED story built on Zhang and the scraper talking. The strongest artifact is Hugging Face's quoted statement, which is a party speaking for itself on the record. The weakest link is the number in the headline: under $10 comes from a Reddit post that no longer exists.
The scrapes shipped; the defenses did not
What is verifiably in the world sits on the wrong side of this fight: two live redistributions, one on Hugging Face and one on Academic Torrents, plus a hosting bill that moved. On Cara's side the concrete deployments amount to login gates its own founder calls a stopgap, and a fundraiser at 83 percent of goal. The protective tool everyone is talking about exists as an agreement between two people.
Redemption arc runs ahead of the artifacts
Two small stretches, pulling the same direction. WIRED's frame lands on the scraper-turned-collaborator, but the tool that would justify that framing does not exist yet, and his contrition, his motives and his deletion of 12 terabytes are all attested only by him. The scale figures also blur: 8.5 million links is a list of addresses, not 8.5 million images taken. The underlying harm, by contrast, is not exaggerated at all — artists are deleting portfolios and the whole public library really did travel.
A fundraiser, two lawsuits, and an anonymous source
Nearly everyone quoted has a stake in how this reads. Zhang is raising $120,000 for legal fees and is a plaintiff in two artist class actions against Stability AI, Midjourney and Google, so a story about lawless scraping serves her case; she also, notably, argues against her own platform's safety pitch. The scraper speaks anonymously after doxing and death threats, with obvious reason to cast the episode as a technical project that got away from him. Hugging Face's statement is a policy position protecting its own hosting posture, issued while it was under takedown pressure.
Solid on what happened, soft on the numbers
That three scrapes hit Cara in August, that datasets went up on Hugging Face and Academic Torrents, and that Hugging Face refused the URL takedown — these hold up well on one outlet's reporting. Confidence drops on the quantities and the aftermath: no server-cost figure, no verification that the archive was deleted, and a headline price supplied by the person who paid it.