StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.10942v1 Announce Type: new Abstract: Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which…

Read original article on cs.AI updates on arXiv.org →