LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Audio
Researchers present LoSATok, a tokenizer designed to serve both audio understanding and generation tasks within a single framework. The work argues that understanding benefits from high-level semantic features while generation needs both semantic and acoustic detail, and existing unified tokenizers encode both in high-dimensional spaces. LoSATok instead uses a low-dimensional representation to handle cross-domain audio.