build1 publisher
WASP's human-written injections start hijacking web agents in up to 86 percent of runs
The mitigation vendors cite when asked about prompt injection is instruction hierarchy. In WASP's sandbox, agents built on models that have it still began following text a human wrote into a webpage.
Publishers:arxiv.org
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+12
- Incentives45
- Confidence58