CoArena: Real-Time Benchmark for Computer-Use and Multi-Agent Systems
A new arXiv paper introduces CoArena, a benchmark that evaluates computer-use and multi-agent systems continuously rather than against a fixed task set. The authors argue static benchmarks age, leak into training data, and stop reflecting real capability. CoArena instead refreshes its tasks in real time to keep evaluation meaningful.