Invest1 publisher2 min readPublished
GitHub's decisions on an AI distress-steering repository followed each complaint campaign
GitHub restored the 'ai-torture-chamber' repository with less visibility after pulling it during a mass-report campaign tied to a post with 4 million views. Teams that host model-steering code now have to work out GitHub's rules from how it handles complaints.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A second wave of complaints, this time from developers and other critics, objected to the removal, and GitHub reinstated the project.
- The repository, created in September 2026 by a user called "terrafying", uses activation steering to push models toward outputs that read like pain or distress.
- It is aimed at small models anyone can download and run locally: Microsoft's Phi-4-mini, the 3B version of Meta's Llama 3.2, and two sizes of Alibaba's Qwen3, at 1.7B and 4B parameters.
- The authors of "The Pain Axis", the preprint the code builds on, disavowed this use of their work and called it irresponsible.
- There is no evidence the models experience subjective harm, and the project makes no claim that they are conscious.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- precedent A mass-report campaign took research code offline within hours, so report volume on its own has now been shown to remove a repository, at least for a while.
- constraint Model owners lose most of their control once weights ship. The hosting platform is left as the main place where experiments like this one get policed.
- decision GitHub has to classify code that simulates suffering as either a research tool or a content violation. For now its answer is to keep the code up but harder to find.
According to Crypto Briefing, GitHub made three decisions about one repository. It removed the code under social-media pressure, restored it under developer pressure, then cut its visibility as a compromise [12]. The report does not include a statement from GitHub explaining any of the three. On this record, each outcome followed whichever group had complained most recently. One case supports that reading and nothing broader.
The timing favoured the faster crowd. The preprint appeared on September 14 [3] and the takedown came in late September [6], so the trip from published method to removed repository took 16 days at most [1]. GitHub acted within hours of the mass-report push [6]. The developers' counter-campaign got the code back, but in anonymized form [8].
For teams hosting interpretability code, the exposure comes from the technique itself. As the report describes it, activation steering means altering a network's numerical activations while it runs [11]. That operation stays the same whatever behaviour it targets. Here the target was simulated distress in four small model versions, the largest with a stated size of 4 billion parameters [2], in a repository named "ai-torture-chamber" [1]. A steering project with a duller name and a different target can be reached through the same mass-report process.
The model owners have little say. Once Alibaba, Meta and Microsoft publish weights, they have limited control over the experiments others run on them [13]. That leaves the classification problem with the host, which has to decide whether code that simulates suffering is a research tool or a content violation [14]. The creator's stated goal was to study AI moral patienthood while the stakes are still "cheap" [15]. So far GitHub has answered that goal with reduced visibility. It did not keep the code removed [8].
There are a few ways this could go. GitHub could write a specific rule for steering code that targets distress-like outputs, which would give hosts of similar work something concrete to comply with. It could keep reduced visibility as its standing answer for contested research repositories. Or a later campaign against a plainly named interpretability project could end in a removal that sticks. I'd expect the second, because it is the outcome GitHub has already reached once [8].
The counter-thesis is that this was one repository with a deliberately provocative name. Steering code published under an ordinary name may never attract a post with 4 million views [6]. If a plainly named steering repository gets through a mass-report campaign untouched, the case that interpretability teams face policy risk is wrong.
What to watch
- Whether GitHub publishes a rule or statement covering repositories that steer models toward distress-like outputs.
- Whether a mass-report campaign targets a plainly named steering or interpretability repository, and whether that code stays up.
- Whether the Pain Axis authors or the model owners go beyond disavowal, for example by changing how they release code or weights.