arXiv CS AI
techCenter
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agentstranslating…
arXiv:2609.06059v1 Announce Type: new
Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact…
Keywords#Agents#Evaluation#Models#DAREBench#Deployment-Aware