papersSEP 10 04:00 UTC
Study Finds Data, Not Typology, Shapes Language Models' Word Order Preferences
A new arXiv paper examines word order preferences in decoder-only language models, testing 192 artificial languages alongside typologically diverse natural languages. The authors report a consistent left-branching bias in the models and argue that training data, rather than linguistic typology, shapes these preferences. The study also explores how this bias relates to model performance on right-branching languages.