papersSEP 10 04:00 UTC
SWORD benchmark probes how consistently LLMs reject false facts across languages
Researchers present SWORD, a benchmark that systematically distorts facts from Wikidata and tests whether large language models notice the resulting errors in different languages. Their experiments reveal that models frequently fail to reject distorted statements consistently across languages, even when they perform well on standard multilingual question-answering benchmarks. The findings suggest existing evaluations can overstate a model's genuine factual understanding outside English.