arXiv Machine Learning#tech
Self-CTRL: Self-Consistency Training with Reinforcement Learningtranslating…
Factuality: 61/100UnknownarXiv Digital Media
arXiv:2606.18327v1 Announce Type: new
Abstract: Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users. This paper describes Self-Consistency Training with Reinforcement Learning (Self-CTRL), a method that optimizes for consistency between a LM's self-explanations and behavior on related inputs by updating explanations to better predict behavior or updating behavior to better match explanations. We apply our method in two d