papersSEP 10 04:00 UTC
SEA-LION-Embedding: Open, Reproducible Text Embeddings for Southeast Asian Languages
Researchers have introduced SEA-LION-Embedding, a set of text embedding models built for Southeast Asian languages and released with open, documented training resources. The work addresses a persistent gap in the field, where leading embedding models cannot be independently reproduced because their training corpora remain private. The release aims to support reliable performance on downstream tasks across the region's many languages.