More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

cs.AI updates on arXiv.org · 1h ago
Research Papers

arXiv:2609.35873v1 Announce Type: new Abstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from…

Read original article on cs.AI updates on arXiv.org →