papersSEP 10 04:00 UTC
ActTraitBench: benchmark quantifies the knowledge-decision gap in LLM persona behavior
Researchers introduced ActTraitBench, a benchmark that checks whether large language models actually behave in line with the personas they describe in explicit self-reports. Using human-grounded behavioral validation, it measures a knowledge-decision gap between what models state and the choices they make implicitly. The arXiv posting is a revised v2 of the paper.