Security1 distinct publisher3 min readUpdated
A free 15.5 GB dump holds no passwords and no payment data, but every row carries Google Ad Manager cohort labels that chess.com's public API does not expose.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
A 15.5 GB file containing more than 7.3 million chess.com user records was posted to two data-leak forums this week at no cost and with no ransom demand attached [1]. According to a technical analysis by Ransomnews, cited by Security Affairs, the data is genuine and recent but is not the product of a hack [2] - and the most interesting thing in it is not the user data at all, but the advertising cohort labels riding alongside it [8].
The mechanics first. The archive is a single 744 MB 7-Zip file that expands into a tab-separated table with one header row and 7,337,395 records of 38 fields each [3]. The schema is chess.com-specific end to end: email, partial email, username, user ID, UUID, first and last name, country, location and locale, plus platform state such as chess title, points, skill level, premium status and label, verification and activation flags, best rating and rating type, official rating, and member-since and last-login timestamps [4]. Roughly three-quarters of records include an email address [5], which puts the exposed mailbox count somewhere around 5.5 million [6]. There are no passwords, no password hashes and no payment data anywhere in the file [7].
The authenticity check did not require touching chess.com. Every account UUID is a version-1 identifier, which embeds the timestamp of its own generation; researchers decoded that timestamp across a 200,000-record sample and compared it to each account's registration date, getting a 100% match [9]. That is hard to forge without real chess.com-issued identifiers.
The same analysis argues the collection was scraped rather than dumped. Records are stamped across nine consecutive days in daily batches, the signature of a scheduled job rather than a single export [10], and about 7.4% of accounts appear twice, revisited on different days [11] - roughly 543,000 duplicated records, which a database export does not produce [12]. This is familiar ground. In November 2023, 828,000 chess.com records surfaced with a near-identical field set, and the company told Hackread flatly, "This was NOT a data breach," adding that "Our infrastructure, member accounts, and data such as passwords are secure" [13]. That set was harvested by abusing the platform's find-friends feature to resolve externally sourced email lists against real accounts [14], and a second scrape of roughly 476,000 users followed [15]. Ransomnews describes the new file as the same technique at roughly nine times the scale [16], which the record counts support [17].
One field pair breaks the public-scrape story. Every record has gam_audiences and audiences_member_of populated with Google Ad Manager audience segments, carrying values including coach-nudge experiment groups, trial eligibility, lapsed-user cohorts and rating-band targeting [8]. Those are marketing-stack attributes, not profile data, and they do not appear in chess.com's public API [18]. Ransomnews reads that as evidence the collector was hitting an authenticated or internal-facing endpoint rather than public developer tools [19]. The operational lesson is not about chess ratings. It is that response payloads assembled for a logged-in client tend to include whatever the ad and experimentation stack attached upstream, and a scraper gets the A/B bucket and the churn label for free.
The account distributing the file, using the handle V0idix, is not selling anything; the same handle has posted dozens of free database dumps from unrelated companies [20].
Watch whether chess.com names the endpoint that returned audience segments, and whether it repeats the 2023 line that this was not a breach [13].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
A 15.5 GB file containing over 7.3 million chess.com user records appeared on two data-leak forums this week, offered free with no ransom demand.
The archive is a single 744 MB 7-Zip file that expands to a 15.5 GB tab-separated table with one header row and 7,337,395 records, each with 38 fields.
The schema is chess.com-specific throughout: email, partial email, username, user ID, UUID, first and last name, country, location and locale, plus platform state including chess title, points, skill level, premium status and label, verification and activation flags, best rating and rating type, official rating, member-since and last-login timestamps.
The file contains no passwords, no password hashes and no payment data.
Every record has the gam_audiences and audiences_member_of fields populated with Google Ad Manager audience segments, with values including coach-nudge experiment groups, trial eligibility, lapsed-user cohorts and rating-band targeting.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Strong artifact forensics, single relayed source, no vendor confirmation
The technical grounding is unusually concrete for a leak story: exact archive and record counts, a 38-field schema enumeration, a version-1 UUID timestamp decode across 200,000 sample records matching registration dates at 100%, a nine-day daily batch pattern and a 7.4% duplicate rate. All of it, however, reaches us through one publisher relaying one analysis, and the pivotal question of how non-public ad-audience fields ended up on every row is left explicitly unresolved with no chess.com response reported.
Free, unpaywalled circulation on two forums with a repeat pattern
Distribution is observed rather than inferred: the archive is posted on two data-leak forums at no cost by a handle that has published dozens of other free dumps, meaning spread is not gated by extortion economics. The pattern also repeats a documented 2023 event and a follow-on scrape, so the collection method has demonstrated durability. What is not observed is any downstream misuse, phishing wave, victim notification or takedown, so real-world impact remains unmeasured.
Careful on user risk, ahead of the evidence on cause
The reporting is deliberately deflationary where it matters most to users, stating plainly that no passwords, hashes or payment data are present and steering readers toward phishing caution rather than account panic. The overstatement sits in the causal framing: the headline treats 7.3 million users as exposed, the analysis asserts on the evidence this is not a hack, and the same article then concedes a detail that contradicts a purely public scrape, while the authenticated-endpoint theory is presented as a strong suggestion without vendor confirmation or a second source.
Reputation-building distributor, vendor with denial precedent, analysis-promoting reporter
Incentives are visible on three sides in the supplied material. The distributor gains status rather than revenue, giving them reason to maximize record counts and visibility rather than verify provenance. chess.com has a documented record of publicly framing similar collections as 'NOT a data breach,' which shapes what the vendor is likely to concede about a non-public field appearing in the dump. And the reporting relays a named security-vendor analysis, whose visibility benefits from the file being genuine and novel.
Well-instrumented artifact, one publisher, unanswered causal question
Confidence is moderate. The artifact-level facts (size, record count, schema, email coverage, absence of credentials, UUID authenticity check, batching and duplication) are internally consistent and independently checkable in principle. But the cluster rests on a single publisher relaying a single analysis, chess.com has not been reported as responding to this file, the nine-day collection window is undated, and the mechanism that placed Google Ad Manager segments on every row is unexplained, which is precisely the claim with the greatest consequence.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026