papersTODAY 04:00 UTC
Benchmark tests entity extraction accuracy in multi-turn voice agent dialogues
Researchers released tau-Elicitation, a 200-task benchmark that measures how well voice agents capture specific entities such as names, addresses, identifiers, dates, and times across multi-turn conversations. The set spans ten entity types with controlled difficulty levels, aiming to pinpoint the exact turn where information capture breaks down. The authors argue that end-to-end evaluations hide these failure points, making targeted diagnosis difficult.