arXiv Machine Learning
techCenter
Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signalstranslating…
1 min readUnknownarXiv Digital Media
arXiv:2608.26571v1 Announce Type: new
Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established…