arXiv NLP#tech
SHARD: Safe and Helpful Alignment via Self-Reframing Distillationtranslating…
Factuality: 62/100UnknownarXiv Press Group
arXiv:2606.15517v1 Announce Type: new
Abstract: Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We introduce SHARD, a self-reframing distillation method to improve safe-helpfulness. It first rewrites sensitive prompts to surface benign intent using philosophical guidelines, then reframes its original responses into safe, more helpfu