Mr.LHDR benchmark evaluates multimodal long-horizon deep research agents
A new arXiv paper introduces Mr.LHDR, a benchmark designed to test deep research agents on extended, multi-step tasks. The authors note that current benchmarks mostly measure shorter exploratory work and seldom assess whether agents can keep going over long horizons. It focuses on web search, tool use and combining evidence from multiple modalities.