papersTODAY 04:00 UTC
PAI-Bench: Benchmark Measures Persistent Identity in Deployed AI Agents
A new arXiv paper introduces PAI-Bench, a provider-neutral benchmark designed to test how faithfully AI agents adhere to a versioned identity contract that can be updated under governance rules. The work argues that existing evaluations conflate an agent's ability to recall identity facts with its ability to express and act on them. The benchmark separates recall from expression and enactment, aiming to give a clearer picture of identity persistence in deployed agents.