
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
arXiv:2610.00197v1 Announce Type: new Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency…
Read original article on cs.AI updates on arXiv.org →