build1 publisher
WorkspaceBench grades activation-to-text tools on 3,356 questions one 27B model can answer
A new 27-family benchmark scores how well a tool reads a model's intermediate variables during a forward pass. Its authors say it does not fully rule out tools that infer the answer from the prompt.
Publishers:lesswrong.com
Reality
- Evidence40
- Adoption10
- Hype gap+10
- Incentives55
- Confidence45