papersSEP 10 04:00 UTC
LexAgentHallu: a hierarchical benchmark for hallucinations in legal AI agents
Researchers have introduced LexAgentHallu, a new benchmark for measuring how tool-augmented legal AI agents hallucinate. It uses a hierarchical structure to trace how errors in tool calls and reasoning cascade into fabricated case holdings and miscited legal authority. The benchmark aims to fill a gap left by existing legal evaluations that do not capture agentic workflows.