papersSEP 12 04:00 UTC
HarvestBench Tests Whether LLM Agents Pay to Avoid Harming Animals
A new arXiv benchmark, HarvestBench, assigns a monetary cost to avoiding a harmful side effect and frames that side effect as the death of a living creature. In the task, nine language models each control two tractors harvesting corn, with animals in their path that are not part of the intended goal. The work measures how much agents are willing to spend to spare them.