papersSEP 10 04:00 UTC
YallaMorph benchmark evaluates Arabic morphological generation in LLMs
Researchers have released YallaMorph, a benchmark for measuring how well large language models generate morphologically accurate Arabic. It addresses a gap in current Arabic evaluation, which focuses on downstream tasks rather than directly testing whether models can control grammatical forms like inflection and derivation. The work highlights that producing fluent Arabic text does not guarantee correct morphosyntactic output.