papersTODAY 04:00 UTC
Mimir paper proposes multilingual concept modeling beyond token-based LMs
A revised arXiv paper titled Mimir argues that current language modeling is organized around tokens, where corpora are split into tokens and models are trained on token-level objectives such as next-token prediction. The authors propose an alternative that works with concepts at a large multilingual scale. The submission is a replacement version of a cross-listed paper.