Gradient-Aligned Pair Selection for Personalized Preference Optimization

cs.AI updates on arXiv.org · 1h ago

arXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing…

Read original article on cs.AI updates on arXiv.org →