arXiv Machine Learning
techCenter
Privacy Without Regret: Differentially Private Inference-Time Alignmenttranslating…
1 min readUnknownarXiv Digital Media
arXiv:2608.26324v1 Announce Type: new
Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward…