papersTODAY 04:00 UTC
DiTAR+ Improves Decoding Stability in Autoregressive Diffusion Speech Synthesis
A new arXiv preprint introduces DiTAR+, a dual-optimization approach for continuous-latent autoregressive diffusion transformer models used in zero-shot speech generation. The method targets the limited decoding stability these models show when producing long utterances or handling complex linguistic input. No results beyond the abstract are described in the report.