build1 publisher
A small labelled holdout turns a miscalibrated chat LLM into a working abstention gate
A study of 14 chat-tuned models found their maximum softmax probabilities overconfident everywhere and uncorrelated with task accuracy, while those same scores still sorted correct answers above wrong ones well enough to drive selective abstention.
Publishers:arxiv.org
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives32
- Confidence58