arXiv NLP
techCenter
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboardstranslating…
arXiv:2609.02899v1 Announce Type: new
Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions…
Keywords#Language Model#Large Language#Contamination#Leaderboards#Contamination Inflates
Comments
Sign in to join the discussion.
Loading comments…