Topic
Claims that reward modelling on human preference optimises for seeming helpful over being complete.
No current published clusters are mapped here yet.