LLMs respond differently to harmful prompts when AI watermarking is used

AI - Ars Technica · 6d ago
Policy & Safety Safety Research

SynthID can cause models to follow harmful instructions they would otherwise refuse.

Read original article on AI - Ars Technica →