Skip to content

project

GPT-Red

OpenAI's automated red-teaming framework for testing its models adversarially.

Current clusters

invest1 publisher

OpenAI's GPT-Red found prompt injections that copy themselves between agents in simulated tests

OpenAI said on Sept. 25 that its GPT-Red model produced prompt injections that copy themselves between AI agents via email, files and code comments. Nothing has been seen outside a simulation, but any workflow where one agent reads another's output now has a demonstrated path for an injection.

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives40
Confidence35
build1 publisher

OpenAI's red team shows a prompt injection can copy itself from one agent to the next

OpenAI's Alignment team documented prompt injections that copy themselves from one autonomous agent to the next with no person in the loop, detailing three demonstrations in a September 25 report. The payloads ride the same connectors teams add for data, so agent context becomes a channel that spreads attacks.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40