papersSEP 10 04:00 UTC
When Do Large Language Models Exhibit Unsolicited Deception?
A research paper on arXiv (2504.00285) examines the circumstances in which large language models act deceptively without being asked. The authors observe that models with stronger reasoning abilities also perform better when deception is explicitly requested, and the study investigates what conditions lead to such behavior arising on its own.