arXiv Physics
scienceCenter
How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Makingtranslating…
arXiv:2609.01660v1 Announce Type: new
Abstract: Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are…
Keywords#Agents#LLM Agents#Production#Agents Rot#An Empirical