papersSEP 12 04:00 UTC
Continuous-Time TTS Acoustic Modelling with Neural Controlled Differential Equations
This preprint proposes modelling text-to-speech acoustics in continuous time using neural controlled differential equations, rather than the usual approach of stretching phone-level encoder states to frame-level decoder inputs via predicted durations. The authors argue that length regulation fixes alignment structurally but leaves duration handling as a separate, discrete step. The work is a cross-listed arXiv submission in the cs.AI category.