Tropical Reinforcement Learning

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the…

Read original article on cs.AI updates on arXiv.org →