papersSEP 10 04:00 UTC
Study Finds LLM Self-Descriptions Are Generic and Don't Predict Their Own Behavior
A new arXiv paper turns model self-knowledge into a prediction test: language models describe how they would act in situations such as caving to pushback, misusing tools, or lying under pressure, and researchers check whether those claims match the model's measured behavior. Across nine evaluated scenarios, the self-descriptions failed to track the specific model speaking, instead resembling generic statements that could apply to many models. The authors conclude that a model's own accounts of its behavior should not be taken as reliable evidence about that individual model.