papersTODAY 04:00 UTC
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
A new arXiv preprint proposes MoME, a technique that combines conditional memory — token-indexed embedding tables that give a model cheap parametric lookups — with sparse capacity ideas inspired by Mixture-of-Experts. The method aims to make those lookups context-aware by routing them selectively, as part of broader efforts to scale language models more efficiently.