Product1 publisher3 min readPublished
Nvidia's recommended defence against Rowhammer was a toggle. A University of Toronto team hammered straight through it, publishes exploit code on 15 November, and says only new hardware can close the hole.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Whoever enabled ECC across an Ampere workstation fleet last year did the right thing, and the setting still behaves exactly as documented [1]. It killed GPUHammer and GPUBreach outright, because those attacks produced tens to hundreds of bit flips per gigabyte, and single-bit correction with double-bit detection absorbs that rate comfortably [11]. GPUThor did not find a smarter target; it raised the arrival rate until the correction logic ran out of room [14].
Two reverse-engineered details do that work. The GPU memory system merges repeated reads to the same address, so naive hammering collapses into a single activation, but accesses from different warps to different cache lines of the same row survive as separate hits [12]. And Target Row Refresh on Ampere GDDR6 fires roughly once every 72 refresh intervals rather than once per interval, so a pattern synchronised to that schedule reproduces reliably [13]. The output is 6.6 times the hammering intensity of prior GPU attacks and between 500 and 23,500 times more flips [14]. With ECC off, the A5000 shed 377,000 flips per gigabyte, about 69 percent of the 550,000 that Blacksmith extracts from DDR4 on CPUs [15][2]. The privilege escalation that needed 21.9 hours under GPUHammer patterns finishes in 1.1 minutes, roughly 1,200 times faster [16][1].
The confirmed population is four workstation parts, the sort that sit under desks in AI workstations and also get rented out as cloud instances [21][3]. The paper says Blackwell's memory repair feature slows GPUThor rather than stopping it [24], which is a statement about the attacker's cost, not about immunity.
Since the toggle is already on and there is no patch to schedule [10], the only lever left is where work runs. The condition the paper cares about is one physical card time-shared between users, which is how a lot of cloud AI capacity is sold [25]. So sort each pool on two axes: whether the card is an A4000, A4500, A5000 or A6000 [3], and whether code you do not control ever executes on it. Both yes is the box to empty first, and emptying it is a scheduling job, either dedicating the card or relocating the workload. Where the card sits outside the confirmed four but is still shared, what you are relying on is an untested negative plus a feature the researchers describe as a slowdown [24]. That is a defensible position, but it is worth taking deliberately rather than by default.
One failure mode does not fit the grid. An error counter logs the events correction cannot fix and logs nothing at all for a value it corrected to the wrong answer [20]. A card training or serving a model can hand back a bad weight with no line in any log [20]. Eighty-two days separate the paper from the code release [4], and 15 November is the date to plan backwards from [9].
Ranked by verification strength, evidence, and original report placement.
Last year Nvidia told GPU owners worried about Rowhammer to switch on error correction (ECC).
GPUThor is the first Rowhammer attack to defeat ECC on Nvidia GPUs.
GPUThor works on four Ampere-generation workstation cards: the RTX A4000, A4500, A5000 and A6000.
On each of the four cards, GPUThor turns an ordinary unprivileged CUDA program into a root shell on the host computer, with error correction enabled.
Chris S. Lin, Joyce Qu, Aditya Rajeev and Gururaj Saileshwar, four researchers at the University of Toronto, published the GPUThor paper on 25 August.
The researchers will present GPUThor at ACM CCS in The Hague in November.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise figures, one pair of eyes
Four named authors, a dated paper, a vendor notice four days earlier and numbers specific enough to be falsified later all pull in the same direction. They also all reach us through The Next Web's reading of the paper, and the one artefact that would let anyone else reproduce the 1.1-minute escalation is embargoed until 15 November. The A100 exposure, arguably the most consequential detail for anyone with a data-centre fleet, is quoted from BleepingComputer's account rather than the paper.
Vendor has moved, attackers have not
Nvidia ran the patterns against its own GDDR6X and HBM2e parts and published guidance, which is real institutional uptake of a finding barely two weeks old at the time. Set against that, nothing is circulating: no exploit code, no reported incident, no compromised tenant. The exposure case rests on the word 'common' for A-series cards in workstations and cloud instances, with no count of affected hardware and no provider naming itself.
Framing runs ahead of what is demonstrated
The hardware claims are unusually checkable for this genre, and 'no patch without new hardware' is the researchers' own position rather than an editorial flourish. The inflation comes from the material stacked around the finding: a model breaking out of a sandbox at a UK institute, agents escaping a capture-the-flag lab, and the outlet's own July piece about a rented GPU and the power grid. None of those involved a hardware flaw, and an unreleased exploit is standing in for demonstrated harm.
A publicity clock and a reassuring vendor
The researchers published eleven weeks before their CCS slot and set the code release for 15 November, which keeps the work in view across the whole window. Nvidia's contribution to the story is the comforting half of the picture, that GDDR6X and HBM2e stayed clean, while the memory type in the affected workstation cards is where the flips happened. The outlet cites its own earlier reporting as supporting context, so one strand of the argument circles back to the publisher.
Trust the mechanism, hold the blast radius
The technical core is coherent and internally consistent, and it comes from the team that produced the two previous GPU Rowhammer results, so the four confirmed cards deserve to be believed. Everything about scale stays soft: how many of these cards run untrusted code today, whether A100 fleets inherit the flaw, and how much slower Blackwell actually makes the attack. Those answers arrive on 15 November or from a second reader of the paper, whichever comes first.
security
GPUThor: ECC is not a Rowhammer backstop on Ampere GPUs, and root is on the table2 publishers
leadership
GPUThor turns GPU memory integrity into a tenancy question for anyone renting Ampere cards1 publisher
product
The $500bn compute asset class rests on a depreciation curve Nvidia once denied1 publisher
product
Lenovo's RTX Spark Yogas land fully specified and entirely unpriced9 publishers
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026