Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but…

Read original article on cs.AI updates on arXiv.org →