Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

cs.AI updates on arXiv.org · 1d ago
Research Papers

arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error…

Read original article on cs.AI updates on arXiv.org →