build1 distinct publisher
SciCon subtracts the text-only score at decode time to stop figure QA from guessing
A preprint argues multimodal models pick scientifically plausible answer text over figure evidence, and that discounting the text-only score during decoding improves accuracy without retraining.
Publishers:arxiv.org
Reality
- Evidence28
- Adoption8
- Hype gap+22
- Incentives55
- Confidence42