Audio language models track speakers via text backbone attention, study finds
A new study examines how audio language models attribute speech to the correct speaker, finding accuracy of only 6 to 16 percent on a six-speaker task, below random guessing. The authors show that speaker tracking relies on attention heads in the model's text backbone, and that altering a subset of those heads shifts which speaker the model retrieves.