
When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing
arXiv:2609.28475v1 Announce Type: new Abstract: Forecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a historical analog as an observable…
Read original article on cs.AI updates on arXiv.org →