arXiv NLP
techCenter
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluationtranslating…
arXiv:2609.02942v1 Announce Type: new
Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants…
Keywords#LLM-as-a-Judge#Rubric#Text#Artifacts#Automated
Comments
Sign in to join the discussion.
Loading comments…