Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations…

Read original article on cs.AI updates on arXiv.org →