Skip to content

other

AgentPoison

Earlier published attack on language model agents that poisons a retrieval memory bank with records pairing a trigger in the agent input to an adversarial target in the output.

Current clusters