build1 distinct publisher
A reward curve that hit 1.0 hid a text-to-SQL model scoring 6.4% on Spider
Three LoRA stages on Qwen2.5-0.5B pushed reward to 1.0 and accuracy below the untrained base. The fix that recovered 43 points was an execution harness that runs both queries and compares the rows.
Publishers:dev.to
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives38