The measurement nobody can get another way
Learning what a safety feature costs in engagement requires withholding it from somebody. TikTok's own account of that tradeoff, in a confidential 2023 document quoted by the Los Angeles Times, reads: "The user impact decision was a delicate balance across safety and the ability to measure impact on DAU (daily active users) and core metrics" [4]. As a methods note, that sentence is ordinary. As a document handed over in litigation against the world's biggest social media companies, placed under court-ordered seal and then described by Bloomberg Businessweek [2], it is a contemporaneous record that somebody weighed a safeguard against a metric and sized the arm accordingly.
The arm was 10% of U.S. users, with an algorithmic safeguard designed to break up echo chambers of harmful content switched off to create a control group [3]. The Los Angeles Times says that group would have been approximately 15 million people at the time [5]. TechCrunch states 15 million as the number the feature was withheld from, and says TikTok wanted to determine whether the safeguard made the app less engaging [14]. Divide the 15 million by the 10% and the implied U.S. base is roughly 150 million accounts [1], which is also a reminder that the round percentage, not the headcount, was the design decision.
Randomisation is what makes an A/B test worth running, and it is also what makes an assignment individually traceable. Chase Nasca, 16, of Bayport, New York, was randomly selected for the experiment on January 25, 2022, and within a month he was dead; the document says his account was fed thousands of videos about sadness, hopelessness, loneliness and suicide [6]. It explains why in a single line: "TikTok's filter bubble prevention strategies did not take effect on this user by design" [7]. The document is dated 2023 [2], roughly a year after that death [6], which means the artifact now in front of Congress is an internal reconstruction of one user's exposure written after the fact. That is the shape of the exhibit: not a policy memo, but a per-user explanation that resolves to an experimental condition.
Thirteen questions, one of which is a schema
Senators Marsha Blackburn and Richard Blumenthal sent a four-page letter on Wednesday to TikTok chief executive Shou Chew and Adam Presser, who runs the company's U.S. spinoff, calling the decision to run the test "depraved" [1]. Twelve of their thirteen questions are the usual discovery-by-letter: the names of every employee informed of the experiment, why minors were permitted in the test, when executives learned of it, and a complete unredacted version of the document [8].
The thirteenth is a data model. The senators want a list of every algorithmic experiment in the U.S. where TikTok "withheld, disabled, delayed, or reduced a safety feature," the number of users involved, and how many of them were minors [9]. Read the four verbs as an engineer rather than a legislator. A staged rollout delays a feature for some users. A ramp reduces it. A guardrail behind a flag is disabled for whoever is not in the flag's audience. The request does not ask whether a register of such experiments exists; it asks for its contents, broken out by arm size and by the age composition of each arm [9]. Any company that has ever run a safety change behind a percentage rollout now has a question to answer about whether its experimentation platform can be queried that way at all.
The letter also fixes why the timeline matters. The senators wrote that Congress had raised concerns about TikTok's algorithm driving young users toward harmful content since October 2021, before the experiment rolled out, which in their framing makes the revelations "even more sinister" because the company "was on notice" [11]. Notice arguments turn "when did executives learn" from colour into the substance of the case. TikTok did not respond to the Los Angeles Times about the letter; in its statement for the Businessweek story a spokesperson said the company was "deeply committed to the safety and well-being of users" and pointed to continued investment in Trust and Safety [12]. TechCrunch also got no immediate response [14].
The deadline is September 1 [10]. Wednesday that week was August 19, 2026 [31], which gives 13 days to assemble a list whose existence nobody has confirmed [2], four and a half years after the assignment date recorded for one 16-year-old [5].
The other clock, and the unit it tests
The Food and Drug Administration's Digital Health Center of Excellence opened a discussion paper on generative AI-enabled medical devices the day before that letter went out, with feedback due under docket FDA-2026-N-7874 by October 19, 2026 [21]. The agency's own summary describes a possible two-axis risk framework, then a premarket approach built on competency assessment, inspired at a high level by how physicians are trained and evaluated, consisting of non-clinical device benchmarking plus clinical confirmation [25].
The load-bearing detail is what gets graded. FDA says it would evaluate the final user-facing device rather than foundational models or isolated subcomponents, and that benchmarking would assess clinical knowledge, analytic capabilities, safety behavior, communication and generalizability [27]. Communication is in the list, alongside knowledge. On the risk side, the framework as reported by Nextgov/FCW could account for whether a product steers a user toward an action or merely provides information, and what happens if a user relies on a faulty output [16]. So the same editorial decision, whether the answer tells the user to do something or hands them a fact and stops, moves the product between risk tiers and also sets part of what the exam scores. Wording stops being surface and becomes a regulatory input.
Clinical confirmation is where the bill appears. MassDevice reports the agency is considering standardized patient interactions with trained actors presenting as patients, a prospective clinical study, and shadow deployment, having acknowledged that benchmark testing may not fully capture real clinical practice [28]. Nextgov/FCW reports the draft allows that testing alone might not be enough, so the agency could also collect real-world evidence [17], which follows from the reason the agency has struggled here in the first place: large language models keep evolving and their outputs vary widely [19].
This also settles a fork quietly. Over the past year academics have proposed requiring generative AI products to pass licensure exams like those doctors take, and the American Medical Association is weighing whether AI licenses are the right path [18]. A license attaches to a persisting entity that can be revoked. FDA's sketch attaches evidence to a shipped, user-facing version, and then leans on postmarket monitoring to cover the fact that the version will not hold still. For a vendor, the practical difference is that a prompt change or a rewritten refusal message touches the tested artifact.
What the agency has not claimed
The disclaimers in the FDA's own posting are unusually broad, and one of them is not boilerplate. The paper is for discussion only, is neither draft nor final guidance, does not communicate proposed regulatory expectations including expectations for supporting evidence in future marketing submissions, and is not intended to address whether the approaches discussed sit within FDA's existing legal authorities or whether new authorities would be necessary [23]. The agency has described an exam without asserting that it can require one. Officials add that the paper is not a policy statement, and Digital Health Center of Excellence director Rick Abramson wrote on LinkedIn that "FDA intends to lead boldly, but with transparency and genuine intellectual humility" [20].
Sitting next to that humility, in the same release, is the claim that the effort aligns with a Trump administration priority to harness AI to accelerate the delivery of innovative medical products to market [24], and CDRH's statement that its ultimate goal is a nimble approach employing least burdensome principles and timely patient access [29]. Trained patient actors and prospective studies are not obviously least burdensome, and nothing in the material reconciles the two. That is the tension a comment letter is for.
Note the asymmetry in who gets time. The regulator soliciting opinions on a framework it has not committed to allowed 62 days [3]. The senators demanding a company's experiment history allowed 13 [2], about a fifth as long [4]. Neither clock was set by the party that has to do the work.
What survives either proceeding
Neither document, as described, prohibits a holdback. The senators ask for a list [9]; the FDA paper explicitly disclaims proposing policy [23]. What both instruments demand is reconstruction: the ability to state, months later, what a specific user was shown or told, and what design decision produced that. TikTok's version of that record was written by hand, in confidence, and is now being read back to its chief executive [1][7].
The record is increasingly a byproduct rather than a choice. In the same week, The New Stack reported that Slack's new agent-only code channels archive themselves when the work is done, stay searchable, and are pitched by Slack as an audit log, inheriting the visibility of whatever conversation spawned them [32]. Nobody has to decide to write that down. The interesting question for a product organisation is no longer whether the trail exists, but whether it can be produced in 13 days and whether it says what you would want it to say.