Skip to content

benchmark

Who&When

Dataset for attributing failures in multi-agent LLM trajectories to the step and agent responsible, with labels requiring that an error was never corrected.

Known aliases

  • Who and When

Relationships

No evidence-backed relationships are recorded.

Current clusters