MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

cs.AI updates on arXiv.org · 1d ago
Research Papers

arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires…

Read original article on cs.AI updates on arXiv.org →