LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#speech-recognition

16 curated events
papersTODAY 04:00 UTC

VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition

A new arXiv paper introduces VoiceCodeBench, a benchmark that checks whether speech recognition transcripts reproduce exact written values rather than just scoring well on word error rate. It focuses on structured tokens such as identifiers, file paths, and commands, which voice-driven workflows often need verbatim. The work argues that WER alone does not capture whether these exact tokens survive transcription.

papersTODAY 04:00 UTC

Sparse Autoencoders Applied to Interpret Whisper Speech Encoder Internals

A new arXiv paper examines the internal representations of Whisper, an automatic speech recognition model, by applying sparse autoencoders to its encodings. The authors note that interpretability research has focused mostly on text-based transformers, leaving speech systems comparatively unstudied. Their work aims to make the features learned by Whisper's encoder more understandable.

papersTODAY 04:00 UTC

Typhoon ASR Streaming enables low-latency Thai speech recognition

A new arXiv paper introduces Typhoon ASR Streaming, a deployable Thai speech recognition system designed for low-latency use cases such as live captioning and voice agents. Most open Thai ASR models are offline and Whisper-based, which prevents them from transcribing incrementally. The system uses real-time shallow fusion and remains steerable during streaming.

papersTODAY 04:00 UTC

arXiv paper examines merging LLM knowledge into automatic speech recognition

A new arXiv preprint in the cs.CL category describes work on combining large language models with automatic speech recognition systems. The paper focuses on knowledge-merging techniques related to established LM fusion approaches such as shallow fusion and density ratio methods. It appears to be a research contribution rather than a product or model release.

papersTODAY 04:00 UTC

Study Finds Data Scale, Not Latency, Shapes Cross-Lingual Transfer in Streaming ASR

A new arXiv paper examines whether a multilingual or English-only encoder is the better starting point when adapting streaming speech recognition models to another language. The authors report that the deciding factor is the amount of training data rather than latency considerations, challenging the assumption that multilingual encoders are always the stronger warm start.

papersTODAY 04:00 UTC

Open Persian speech corpus Neyshekar released with 99 hours of audio

Researchers have published Neyshekar, an openly available Persian read-speech corpus intended to cover formal and informal speech, named entities, and longer sentences. Version 6 contains 62,279 validated recordings totaling 99.02 hours, contributed by 190 speakers. The dataset is aimed at supporting automatic speech recognition work in Persian.

papersTODAY 04:00 UTC

One-shot pruning found to act as implicit regularizer for speech recognition models

A study argues that one-shot magnitude pruning does more than compress neural networks, acting as an implicit regularizer for automatic speech recognition. Testing with Whisper-small, the authors combine gradient- and Fisher-based sensitivity measures to guide which weights to remove. The work reframes pruning as a training technique rather than only a efficiency tool.

papersTODAY 04:00 UTC

Language Model Priors Used for Acoustic Adversarial Attacks on ASR

A new arXiv paper examines how language model priors can be leveraged to craft acoustic adversarial attacks against automatic speech recognition systems. It focuses on real-time ASR, where transcription decisions must be made under strict temporal limits using incomplete audio input. The work suggests that this causal constraint creates an exploitable vulnerability in streaming recognition pipelines.

papersTODAY 04:00 UTC

arXiv paper describes Sophea, a production Greek-English speech recognition system

A research team reports on a multi-month engineering effort to build Sophea, a bilingual Greek-English automatic speech recognition system intended for production use. The system was assessed against nine production gates, including word error rates for both languages and language identification performance. The work is presented as a case study in the engineering work required to move speech recognition from research to deployment.

papersTODAY 04:00 UTC

Token Merging for Multilingual Speech Recognition Studied Across Model Scale

A new arXiv study systematically examines token merging as a way to cut the computational cost of large multilingual speech recognition models such as Whisper. The technique dynamically combines token representations during inference, and the authors test how its effectiveness varies with model size and fine-tuning. The work targets deployment efficiency for transcribing low-resource languages without language-specific training.

papersSEP 10 04:00 UTC

BaltiVoice: 16.8-hour speech corpus and fine-tuned Whisper ASR system for Balti

Researchers have released BaltiVoice, a 16.8-hour read-speech corpus with 10,060 validated utterances for Balti, a Tibetic language spoken in Gilgit-Baltistan, Pakistan. The language previously had no publicly available speech recognition resources, making this the first open dataset and ASR model for Balti. The team fine-tuned OpenAI's Whisper architecture on the corpus to enable automatic speech recognition for the language.

modelsSEP 10 04:00 UTC

Qwen-Audio-3.0-ASR technical report details LLM-integrated speech recognition

A technical report published on arXiv introduces Qwen-Audio-3.0-ASR, an automatic speech recognition system that combines scaled training data, larger model architecture, and integration with large language models. The paper outlines the system's design choices and evaluates its performance, situating it within recent progress in ASR research.

papersSEP 10 04:00 UTC

Orukeet paper proposes multilingual ASR using frozen Gabor kernels in Parakeet encoder

A new research paper introduces Orukeet, a speech recognition model that replaces half of an adapted Parakeet encoder's temporal filtering layers with 12,288 fitted Gabor kernels, which are then kept frozen. The remaining parameters are trained on multilingual and multi-accent speech data, followed by a final adaptation and checkpoint selection stage. The approach aims to extend strong transcription models to more languages and accents.

papersSEP 10 04:00 UTC

BuzzASR: 100+ monolingual fine-tuned Whisper models for speech recognition in 102 languages

A new arXiv paper introduces BuzzASR, a collection of more than one hundred monolingual Whisper models fine-tuned for automatic speech recognition in 102 languages. The language-specialized models are designed to cover languages that large general-purpose multilingual ASR systems often serve poorly. The release is presented as a resource for practitioners and researchers working on speech technology across many languages.

productsSEP 10 09:42 UTC

Analysis questions privacy implications of Apple Watch Audio Intelligence

A German tech outlet examines Apple's Audio Intelligence feature, which is designed to detect and interpret speech picked up by the Apple Watch worn by millions of users. The piece frames the always-available listening capability as a potential step toward a surveillance-heavy future. It questions how such ambient audio processing fits with expectations of privacy on a wrist-worn device.

tipsAUG 28 00:00 UTC

Hugging Face Open ASR Leaderboard Adds First Global South Language

Hugging Face's Open ASR Leaderboard has expanded its coverage to include a language from the Global South for the first time. The addition broadens the benchmark's evaluation of automatic speech recognition systems beyond the predominantly high-resource languages it previously tracked. It reflects a wider push to measure model performance on underrepresented languages.

WHY IT MATTERS ↘Benchmarks drive where engineering effort goes, so extending a widely cited ASR leaderboard to a Global South language gives vendors and researchers a shared target for measuring quality on languages that commercial incentives alone have largely ignored. The caveat is that a single added language still reflects an underrepresentative sample, so teams should treat it as a starting signal for data collection and evaluation rather than evidence of broad multilingual coverage.