AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.11050v1 Announce Type: new Abstract: Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots…

Read original article on cs.AI updates on arXiv.org →