Distillation for Incrimination and Distillation for Capabilities

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.11012v1 Announce Type: new Abstract: Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence…

Read original article on cs.AI updates on arXiv.org →