arXiv CS AI
techCenter
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requeststranslating…
arXiv:2608.27831v1 Announce Type: new
Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and…
Keywords#Coding Agents#User Requests#A Compositional#A Compositional Evaluation#Compositional Evaluation