papersTODAY 04:00 UTC
arXiv paper extends adaptive testing to continuous-score LLM evaluation
A revised arXiv paper proposes an adaptive evaluation method that applies computerized adaptive testing ideas to generation tasks, where model outputs receive continuous scores instead of binary or multiple-choice marks. The approach aims to reduce the number of items needed while maintaining confident ranking of models. It targets LLM benchmarking beyond traditional multiple-choice setups.