papersSEP 10 04:00 UTC
New 41B-token European Portuguese web corpus introduced in arXiv paper
Researchers have assembled a 41-billion-token collection of web text focused on European Portuguese, addressing the difficulty of separating it from Brazilian Portuguese in large-scale crawls. The paper describes an efficient processing pipeline that handles dialectal overlap at scale to yield a dataset intended for production training use. The work appears as a preprint cross-listed in language computation and AI categories.