Study finds demographic identity in language models splits into three distinct properties
A new arXiv paper argues that a language model's demographic identity has three separable aspects — whether it is readable, whether it is faithful to real group differences, and whether the model actually uses it when generating text. The authors test this using internal representations, addressing why LLM-simulated survey respondents tend to be homogeneous and diverge from real inter-group patterns. The work suggests unfaithful simulation may stem from what models use rather than what they know.