papersSEP 10 04:00 UTC
MultiSynt/MT: open synthetic corpus offers 4.8T tokens in 36 languages for multilingual pretraining
Researchers have introduced MultiSynt/MT, an open synthetic parallel dataset totaling roughly 4.8 trillion target-language tokens spanning 36 languages. Because web-scale training data is heavily skewed toward English, the corpus is intended to give model builders far more multilingual material for pretraining large language models. The work is described in a paper posted on arXiv.