Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

cs.AI updates on arXiv.org · 2h ago

arXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and…

Read original article on cs.AI updates on arXiv.org →