arXiv CS AI
techCenter
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Samplingtranslating…
arXiv:2608.27982v1 Announce Type: new
Abstract: Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most…
Keywords#Dynamic Sampling#Advantage#Advantage Amplification#Difficulty-aware#Difficulty-aware Advantage