papersSEP 12 04:00 UTC
Benchmark Radar Offers Searchable Database of AI Evaluation Benchmarks
A new arXiv paper introduces Benchmark Radar, a searchable database and engine intended to help model developers locate relevant evaluations along with their datasets and code. The system also aims to make the conditions behind reported benchmark scores easier to understand. It is targeted at researchers working on large language models and other AI systems.