LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

Text-to-speech

topic9 events
papersTODAY 04:00 UTC

Reference-Free Metric Targets Lexical Tone Evaluation in Multilingual TTS

A new arXiv paper proposes a reference-free way to measure whether text-to-speech systems render lexical tone correctly in languages where pitch changes word meaning. The authors note that standard character error rate scoring misses such errors, using Yorùbá words like "ọkọ" (husband), "ọkọ̀" (vehicle) and "ọkọ́" (hoe), which differ only by tone, as an illustration. The work aims to give a budget-friendly evaluation option for massively multilingual speech systems.

papersTODAY 04:00 UTC

Study Questions Reliability of Reference-Free Speech Quality Metrics for TTS

A new arXiv paper examines whether reference-free speech quality predictors such as UTMOS, DNSMOS and SCOREQ are dependable both as automatic evaluators for text-to-speech systems and as reward signals in preference optimization. The authors argue that these dual roles rest on assumptions about prediction accuracy that may not hold for modern TTS outputs. The work suggests current evaluation practices could mislead comparisons and reward-based training.

papersTODAY 04:00 UTC

ParsVoice: Large Multi-Speaker Persian Speech Corpus for Text-to-Speech

Researchers released ParsVoice, described as the largest publicly available multi-speaker Persian speech corpus, aimed at addressing the scarcity of open speech-text data for the language. The dataset is intended to support work on multi-speaker text-to-speech, speech-language modeling, and other low-resource speech tasks. The paper was posted to arXiv under computer science categories.

papersTODAY 04:00 UTC

Revisable CTMC Inference Stack Proposed for Guided Discrete Flow Matching TTS

A new arXiv paper introduces a mask-sample-revise inference pipeline built on continuous-time Markov chains for guided discrete flow matching in text-to-speech. The approach targets alignment-free non-autoregressive TTS systems that treat synthesis as conditional infilling over neural codec tokens, avoiding separate duration predictors and external aligners. It proposes letting the sampling process revisit and correct earlier token decisions during generation.

papersSEP 12 04:00 UTC

Continuous-Time TTS Acoustic Modelling with Neural Controlled Differential Equations

This preprint proposes modelling text-to-speech acoustics in continuous time using neural controlled differential equations, rather than the usual approach of stretching phone-level encoder states to frame-level decoder inputs via predicted durations. The authors argue that length regulation fixes alignment structurally but leaves duration handling as a separate, discrete step. The work is a cross-listed arXiv submission in the cs.AI category.

papersSEP 10 04:00 UTC

X2-NativeCursor: Token-Level Text Progress Tracking for Streaming Incremental TTS

Researchers introduce X2-NativeCursor, a mechanism that lets incremental-text streaming TTS systems track which part of the input is currently being spoken. Since text usually arrives ahead of the audio, arrival timing alone cannot indicate spoken progress, so the method provides a token-level cursor for alignment. This supports features like synchronized highlighting, interruption handling, and real-time dialogue-history updates in speech applications.

papersSEP 10 04:00 UTC

Deterministic prompting for speaker-stable low-resource Greek TTS

Researchers present a data curation recipe combined with deterministic prompting to keep speaker identity stable in text-to-speech systems trained on limited speech data. Modern Greek serves as the test case, since it lacks the curated corpora that underpin state-of-the-art synthesis for high-resource languages. The work addresses quality degradation that TTS models typically show when clean training speech is scarce.