arXiv CS AI
techCenter
Rating the Raters: Rasch Measurement Theory for LLM Evaluationtranslating…
arXiv:2608.27463v1 Announce Type: new
Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of…
Keywords#Evaluation#Measurement#Raters#LLM#LLM Evaluation