Security1 distinct publisher2 min readPublished
A Samsung and University of Warsaw preprint tracks the names large language models keep inventing, Elena Vasquez and Marcus Chen among them, into hundreds of AI-generated papers that scholarly aggregators index without checking whether the author exists.
The Watch · Security desk
Compiled by The WatchSomething wrong?How this is made
Zenodo mints real DataCite DOIs, and anyone with a free account can create one [6][4]. Once the identifier exists the harvest is automated: the preprint's authors write that these records carry real DOIs harvestable by any scholarly aggregator, and that the infrastructure for large-scale contamination of the scholarly record is already in place [7]. Many of the records were also backdated, with publication dates that do not match the dates they were uploaded [5]. That detail does more work than the volume does, since a backdated record reads as prior work, the kind a later, genuine paper should have cited.
The finding beyond the name lists is correlation. Models do not only repeat single names, they produce what the authors call correlated character ensembles, names that tend to appear together [9]. So a fabricated record arrives with a co-author list that hangs together rather than one obviously invented byline, which is what gets it past a skim.
The same regularity is the only cheap detection signal on offer. Elena Vasquez was particularly common in Claude Sonnet 4 output, so the researchers treat her presence on a paper as evidence that model produced it [12]. Michal Brzozowskim, the paper's lead author, told 404 Media that this accuracy may not hold, because content carrying these names is flooding the web and feeding back into the models that scrape it for training data [13][16]. An indicator that erodes as the contamination spreads carries its own expiry date, and that limits how far it can serve as a control.
The reporting names seven identities: Elena Vasquez, Marcus Chen, Elias Thorne, Elena Amara Okafor, Aris Thorne, Lena Petrova and Elara Voss, attributed across ChatGPT, Gemini and Claude [18]. Elias Thorne was the June case, a lighthouse keeper that three separate models kept generating in fiction [11]. The behaviour appears across the model class, not in one vendor's tuning alone.
The pattern is already load-bearing outside academia. Snopes traced a Facebook rumour that Alex Pretti, killed by US Border Patrol agents in January, had been fired from a nursing job over misconduct allegations; the claim was attributed to an executive director named Dr. Elena Vasquez, who does not exist [15].
What is published and what is inferred are different things. The preprint, its Zenodo count, and the aggregator behaviour it describes are published [1][4]. Which model produced any given document is inferred [12]. For anyone running a retrieval pipeline or a systematic review, the operational change is narrow. A DOI certifies that an account existed at mint time, and author existence is a separate check, because the aggregators are not performing it [8].
Ranked by verification strength, evidence, and original report placement.
A new preprint from Samsung and the University of Warsaw is titled "The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing".
The preprint states: "Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived."
The paper identified a number of names that co-authored hundreds of AI-generated academic papers, articles and books; the authors do not exist but are names large language models repeatedly produce when asked to generate experts in certain fields.
The paper states: "On Zenodo, a CERN operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates."
The researchers saw that many of the papers authored by these AI names were backdated, meaning their publication dates were different from the date they were uploaded to Zenodo.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Scalable Capital puts ChatGPT, Claude and Grok inside the European order ticket2 distinct publishers
product
A school agenda shipped with "Vitoiis" and a planet named Marc, and no one read it first1 distinct publisher
product
Gamma says it hit $100m ARR with 50 people, and 15 months of profit to go with it1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified preprint findings, single outlet, no peer review or platform response
The central assertions are specific and quantified (1,655 Zenodo records, named identities per model family, backdated metadata) and are quoted directly from the preprint, with corroborating third-party anchors in the Snopes debunk and the outlet's own prior Research Gold reporting. But the cluster contains one publisher, the paper is a preprint with no stated peer review or disclosed detection method, the aggregator-indexing claim rests on the authors' own assertion, and Zenodo, DataCite, ResearchGate, Google Scholar and Semantic Scholar are not quoted.
Contamination measurably present in live repositories; downstream harvest volume unknown
The phenomenon is already instantiated in production scholarly infrastructure rather than hypothetical: 1,655 DOI-bearing records exist on Zenodo, ghost names appear on ResearchGate profiles, a commercial vendor fronted a fabricated founder until exposed, and arXiv has tightened submission policy. What is not measured is how much of this material has actually been harvested, cited or ingested by aggregators and downstream users, which caps the score.
Slightly overstated: haunting framing outruns measured downstream harm
The framing - "the academic record is being quietly haunted" and "infrastructure for large-scale scholarly record contamination is already in place" - reaches further than the supplied measurements, which establish record counts and indexing presence but no demonstrated citation uptake, retraction cascade or corrupted downstream research. The provenance-attribution pitch is also presented alongside its own expiry date, since the lead author expects accuracy to decay via training feedback. The overstatement is modest because the core numbers are specific and the infrastructure gap (free-account DOI minting) is concrete.
Researchers promoting an unreviewed preprint; outlet building on its own prior scoops
The findings come from authors publicising their own preprint, which also advances a detection method they propose, and the reporting outlet repeatedly cites its own earlier Elias Thorne and Research Gold stories, giving it continuity interest in the narrative. These are ordinary publication and coverage incentives rather than commercial or funding conflicts, and no vendor, sponsor or product is being sold, so distortion pressure is moderate.
Coherent single-source account of specific findings, awaiting external verification
Confidence is moderate: the quantified claims are internally consistent, quoted directly, partly corroborated by an external fact-check and by a documented corporate removal, and the structural enabler (open DOI minting) is easy to verify. It is held down by having one publisher, an unreviewed preprint whose detection methodology is not described, no platform or aggregator comment, and no independent replication of the model-attribution result.