Reinforcement Learning with Decomposed Subtasks

cs.AI updates on arXiv.org · 1h ago
Research Papers

arXiv:2609.27035v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy:…

Read original article on cs.AI updates on arXiv.org →