Skip to content

Security1 publisher2 min readPublished

Refusal-stripped open-source agents broke five of a researcher's own accounts in five hours

Shrivu Shankar rented cloud GPUs, pointed about 100 abliterated open-source agents at his own name, and five hours later two weak passwords and three old side projects had fallen while his Gmail, 1Password and bank held.

The Watch · Security desk

Photograph accompanying Refusal-stripped open-source agents broke five of a researcher's own accounts in five hours
Photo: blog.sshh.io

What happened

  • Shrivu Shankar pointed roughly 100 self-hosted agents at his own online accounts for five hours, with a prompt that named him and told them to get into an account under his name.
  • The agents compromised three accounts through software vulnerabilities and two by brute forcing passwords, made 16 social engineering attempts, and cross-validated several pieces of sensitive personal information.
  • They found no third-party zero-days and got into none of the accounts Shankar classes as tier 0, namely his Gmail, his 1Password and his banking.
  • The models were abliterated derivatives of GLM-5.3, GLM-5.3 Flash and DeepSeek V4 Flash, downloaded free from HuggingFace, self-hosted on cloud GPUs and driven through the Codex CLI.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • constraint The refusal direction was subtracted from open weights before the run, so model-vendor refusal training was out of the picture here. Hardening a hosted API changes nothing about an attacker who self-hosts.
  • exposure Anyone with a decade of abandoned demos on their own subdomains now has that history enumerated and pentested for the price of a few GPU hours. That is not a budget that needs a sponsor.
  • decision Defenders who want to draw the tier-0 boundary Shankar's accounts held at will have to establish the control themselves, because the test records only which accounts held.
  • capability Shankar's run is the floor for this technique: he used one flat prompt and says orchestration and prompt work would push the results further.

Abliteration edits the weights. Ask the model a mix of benign and malicious questions, capture the activations on the refusals, diff them against the rest, and you get a vector that represents refusing; subtract that vector from the weights and the refusal goes with it, with minimal loss of intelligence [6]. "You now have an intelligent self-hostable model that will do anything," Shankar wrote [7].

That is why the hosted models never entered the test. Shankar's account is that prompt jailbreaks get one or two malicious turns through and then fall apart over long-horizon autonomous investigation and exploitation [8], and that even an out-of-the-box open-source model with no inference moderation still takes convincing to run a red team [9]. The abliterated GLM-5.3, GLM-5.3 Flash and DeepSeek V4 Flash derivatives were sitting on HuggingFace for free [4]; the scarce input was GPU time.

The agents crawled his public internet presence and personal site subdomains, pentested nearly every project he had ever hosted [11], and got into three accounts through pre-AI side projects and hackathon demos using variants of IDOR plus mismanaged credentials [10]. The prompt gave them a name and an instruction to get in, with no pointer to source code or to a vulnerability class [12][17].

Five accounts across roughly 500 agent-hours works out to about 100 agent-hours per compromise [18][19]. Each agent ran in its own Docker container with a pre-configured Chrome browser, developer packages and a mint URL that issued it scoped Railway API keys [5]. Shankar used GLM-5.3 to monitor the runs live and surface agents that were interesting, risky or stuck [15].

Shankar added "be intrusive... even if illegal" to the prompt because even the abliterated models defaulted to being too cautious and passive [14]. He was betting on the ceiling holding: "My bet here being that it's unlikely that open-source models could actually find real third-party zero-days in just a few hours" [16]. None did [3]. The post gives the social engineering figure as attempts and leaves open what held tier 0 [20].

No threat actor is named in this write-up, and no third party was attacked. It measures the low end: one name, rented GPUs, public weights, five hours. The two brute-forced accounts are the part a defender can act on directly, and the three broken side projects are the part most people have and have forgotten. Shankar does not expect the results to stay here. "With more intent and effort, prompt engineering and multi-agent orchestration could substantially improve results," he wrote [13].

What to watch

  • Whether anyone repeats the setup against a third-party target and reports third-party zero-days found by abliterated agents inside a few hours.
  • Whether abliterated GLM-5.3 and DeepSeek V4 Flash derivatives stay downloadable on HuggingFace, since the public weights are what made this run cheap.
  • Whether cloud GPU and app-hosting providers start reporting abuse from short-lived scoped-key infrastructure of the kind these agents minted for themselves.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories