arXiv paper proposes first-principles update geometry for language-model output head
A revised arXiv preprint argues that the spectral norm used by the Muon optimizer is not a good measure of functional change for a language model's output head, because softmax removes scale invariance assumptions. The author proposes deriving an update geometry tailored to how that parameter block actually functions. The work is theoretical and has not been peer reviewed.