Study examines domain-specific pretraining effects on Arabic-English code-switching models
A new arXiv paper looks at how a model's pretraining domain profile affects Transformer performance on digital pragmatics in Arabic-English code-switched text. It compares MARBERT and XLM-R against a general-purpose BERT baseline. The work focuses on whether domain-targeted pretraining yields better results for this kind of mixed-language discourse.