papersSEP 11 04:00 UTC
Study examines how scoring rules affect LLM forecasting accuracy
A paper on arXiv compares five proper scoring rules used as training objectives for large language models making binary forecasts about real-world events. The author reports that the choice of reward function influences both the accuracy and the behavior of the resulting forecasters, even though the rules are theoretically equivalent. The work suggests reward design matters when fine-tuning models for prediction tasks.