LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#asr

9 curated events
papersTODAY 04:00 UTC

arXiv paper proposes sequential adapter stacking for low-resource ASR

A new arXiv preprint describes a method that stacks adapters sequentially to help large multilingual speech recognition models handle low-resource languages. The authors note that current systems perform unevenly, favoring high-resource languages and losing accuracy where labeled audio is scarce. The abstract covers the approach at a high level, with results and evaluation details not included in the excerpt.

papersTODAY 04:00 UTC

Controllable Dysarthric Speech Synthesis for Speaker-Diverse ASR Training

Researchers propose a speech synthesis method that separates speaker identity from dysarthric articulation patterns, allowing finer control over generated dysarthric speech. The approach conditions synthesis on individual patients, producing varied synthetic speakers to supplement scarce training data for dysarthric speech recognition. This addresses a field bottlenecked by high speaker variability and limited labeled recordings.

papersTODAY 04:00 UTC

Typhoon ASR Streaming enables low-latency Thai speech recognition

A new arXiv paper introduces Typhoon ASR Streaming, a deployable Thai speech recognition system designed for low-latency use cases such as live captioning and voice agents. Most open Thai ASR models are offline and Whisper-based, which prevents them from transcribing incrementally. The system uses real-time shallow fusion and remains steerable during streaming.

papersTODAY 04:00 UTC

Replay-Based Editing Reduces Timestamp Drift in Autoregressive ASR

A new study examines how autoregressive speech recognition systems that output timestamps as decoded tokens can gradually lose alignment during long stretches without speech. The authors propose a replay-based distribution editing approach that corrects this drift while limiting forgetting of previously learned behavior. The work targets timestamped transcription without relying on frame-level aligners or inference-time fixes.

papersSEP 10 04:00 UTC

Orukeet paper proposes multilingual ASR using frozen Gabor kernels in Parakeet encoder

A new research paper introduces Orukeet, a speech recognition model that replaces half of an adapted Parakeet encoder's temporal filtering layers with 12,288 fitted Gabor kernels, which are then kept frozen. The remaining parameters are trained on multilingual and multi-accent speech data, followed by a final adaptation and checkpoint selection stage. The approach aims to extend strong transcription models to more languages and accents.

modelsSEP 10 04:00 UTC

Qwen-Audio-3.0-ASR technical report details LLM-integrated speech recognition

A technical report published on arXiv introduces Qwen-Audio-3.0-ASR, an automatic speech recognition system that combines scaled training data, larger model architecture, and integration with large language models. The paper outlines the system's design choices and evaluates its performance, situating it within recent progress in ASR research.

papersSEP 10 04:00 UTC

BuzzASR: 100+ monolingual fine-tuned Whisper models for speech recognition in 102 languages

A new arXiv paper introduces BuzzASR, a collection of more than one hundred monolingual Whisper models fine-tuned for automatic speech recognition in 102 languages. The language-specialized models are designed to cover languages that large general-purpose multilingual ASR systems often serve poorly. The release is presented as a resource for practitioners and researchers working on speech technology across many languages.

papersSEP 10 04:00 UTC

Researchers propose fine-grained error correction to improve Korean ASR for consultation services

A newly published arXiv paper tackles recognition errors that occur when automatic speech recognition is deployed in customer consultation settings. The authors introduce a fine-grained error correction method aimed at Korean-language call center dialogue, where even advanced ASR systems continue to make mistakes. The work targets customer service automation and large-scale transcription applications.