papersSEP 10 04:00 UTC
Study identifies 'cultural binding heads' shaping cultural context in LLMs
A new arXiv paper examines why large language models tend to respond uniformly to different cultural groups even when context should call for differentiation. The authors combine mechanistic interpretability with a factorial experimental design on a cultural-appropriation benchmark, locating specific attention heads they term 'cultural binding heads' that appear tied to this behavior. The work is cross-listed on arXiv under machine learning, artificial intelligence, and computational linguistics.