Build1 distinct publisher2 min readPublished
The one unlearning method that comes with a proof costs a full training run, and the benchmark that defines "efficient" rules it out. Compliance teams are promising deletion anyway.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The gap an operator has to manage sits between two sentences that read the same. "We have deleted your data" can describe a store operation with a log entry to prove it, or it can describe removing a record's influence from a set of weights. According to the dev.to explainer, dropping the document from the training corpus leaves an already-trained model exactly as it was [4], and the erasure right was drafted on the assumption that personal data sits somewhere findable [1]. The two readings share a verb and nothing else.
Nothing in the parameters is addressable as any one person: a document's contribution is smeared across weights that also encode a great many other things [3]. That is why the obvious audit request, show me that it is gone, has no artefact to point at.
Cost is what forces the substitution, and it does not amortise. Each requester implies a different held-out dataset, so honouring R requests exactly means R full training runs [12]. At the price range the piece cites for a single run [6], that arithmetic is the reason the one method with a real guarantee is also the one no lab will run at request volume [7].
What replaces the proof is a test. Approximate unlearning is judged empirically, by whether the data can still be extracted or the model still completes the forgotten passage, and there is no settled definition of what "forgotten" means in a probabilistic system [9]. That evidence is negative by construction: it can show that a method failed to find something, not that nothing is there. The cheap route also has a side effect worth pricing. A heavy-handed gradient ascent pass degrades the model's general ability [8], so one person's erasure request gets paid for in quality by everybody else using the product.
The benchmark language is the clearest admission of where this actually stands. By the efficiency bar used in the 2023 challenge, exact unlearning cannot qualify [11], so the field's own scoring rules treat the only provable answer as out of scope. That is a reasonable way to organise a research programme. It is a thin basis for a privacy notice that tells a user their data will be removed from the model itself.
Ranked by verification strength, evidence, and original report placement.
Approximate unlearning buys efficiency by giving up guarantees: it is typically judged not by a proof but by empirical tests, such as whether the data can still be extracted or the model still completes the forgotten passage, and there is no settled definition of what counts as successfully forgotten in a probabilistic system.
The GDPR right to erasure is built on the mental model that your data sits somewhere as a discrete record that can be located and destroyed, as in a normal database operation.
A trained AI model does not store your data as a record; it is dissolved into the model's parameters, billions of numbers each nudged a little during training by every example the model saw.
A single document does not live in one identifiable place in the weights; its contribution is smeared across many parameters that also encode a great many other things, so you cannot point at the part of the model that is you.
Deleting the original document from the training set does nothing to a model that has already trained on it: the data is gone, the influence remains.
Exact unlearning, meaning retraining the model from scratch on the dataset with the person's data left out, is the gold standard because the resulting model provably never saw that data.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Sound mechanism, single unsourced explainer
The core technical claims — distributed influence in weights, no effect on a trained model from deleting source data, exact unlearning as the only proof-carrying method — are internally coherent and consistent with well-known training mechanics, and the article defines its terms precisely. But the cluster contains exactly one source, a community-platform explainer, and every quantitative or third-party anchor (training-run cost band, the 2023 Google-organised challenge and its efficiency criterion, 2025 extraction and half-cost results) is asserted without a link or named paper. Nothing here is independently corroborated.
No adoption signal in cluster
The cluster contains no release, deployment, benchmark result, pricing or usage disclosure that can be dated and attributed. The only real-world hook is a passing, undated reference to Atlassian pledging to remove opted-out data and retrain, with no scope, timeline or confirmation, and the 2023 challenge and 2025 research results are mentioned without identifiers. That is not enough to measure whether unlearning methods are in production use.
Deflationary framing, slightly overstated absolutes
The article's rhetorical direction is deflationary — it argues deletion promises are weaker than they sound — so it is not inflating a product story. Mildly positive rather than zero because two moves reach past the evidence supplied: the cost band and the 2025 results are asserted without citation, and the derived per-request accounting treats erasure as strictly one training run per request, ignoring that requests could be batched into a scheduled retrain, which would soften the impossibility framing the dek leans on.
Self-published critique channel, no disclosed vendor stake
The single source is a self-published post on a developer community platform under an editorial identity oriented to AI's downsides ('theaidownside' in the URL path), which creates a mild directional pull toward findings that deletion promises fail. Offsetting this, the piece names no product it sells, discloses no vendor relationship, criticises a named company only in passing, and explicitly warns readers not to conclude erasure is fake. No sponsorship, affiliation or commercial interest is disclosed in the supplied material, so the reading is limited to observable framing.
Confident on mechanism, thin on specifics
Confidence is moderate: the structural claims that carry the story would survive corroboration because they restate standard training mechanics and a well-known research gap, and the ledger's derivations mostly follow from the stated facts. It is capped by single-publisher sourcing, absent primary citations for every number and benchmark, no adoption evidence, and an unresolved legal question — whether weights are in scope for an erasure request — that the cluster never touches.
build
Atlassian's full AI opt-out costs an Enterprise upgrade: three of four tiers can't switch metadata off1 distinct publisher
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
build
"The model does not retain training data" is a testable claim, and someone else runs the test1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026