LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

Whisper

model5 events
papersTODAY 04:00 UTC

Token Merging for Multilingual Speech Recognition Studied Across Model Scale

A new arXiv study systematically examines token merging as a way to cut the computational cost of large multilingual speech recognition models such as Whisper. The technique dynamically combines token representations during inference, and the authors test how its effectiveness varies with model size and fine-tuning. The work targets deployment efficiency for transcribing low-resource languages without language-specific training.

papersTODAY 04:00 UTC

Sparse Autoencoders Applied to Interpret Whisper Speech Encoder Internals

A new arXiv paper examines the internal representations of Whisper, an automatic speech recognition model, by applying sparse autoencoders to its encodings. The authors note that interpretability research has focused mostly on text-based transformers, leaving speech systems comparatively unstudied. Their work aims to make the features learned by Whisper's encoder more understandable.

papersTODAY 04:00 UTC

Typhoon ASR Streaming enables low-latency Thai speech recognition

A new arXiv paper introduces Typhoon ASR Streaming, a deployable Thai speech recognition system designed for low-latency use cases such as live captioning and voice agents. Most open Thai ASR models are offline and Whisper-based, which prevents them from transcribing incrementally. The system uses real-time shallow fusion and remains steerable during streaming.

papersSEP 10 04:00 UTC

BuzzASR: 100+ monolingual fine-tuned Whisper models for speech recognition in 102 languages

A new arXiv paper introduces BuzzASR, a collection of more than one hundred monolingual Whisper models fine-tuned for automatic speech recognition in 102 languages. The language-specialized models are designed to cover languages that large general-purpose multilingual ASR systems often serve poorly. The release is presented as a resource for practitioners and researchers working on speech technology across many languages.

papersSEP 10 04:00 UTC

BaltiVoice: 16.8-hour speech corpus and fine-tuned Whisper ASR system for Balti

Researchers have released BaltiVoice, a 16.8-hour read-speech corpus with 10,060 validated utterances for Balti, a Tibetic language spoken in Gilgit-Baltistan, Pakistan. The language previously had no publicly available speech recognition resources, making this the first open dataset and ASR model for Balti. The team fine-tuned OpenAI's Whisper architecture on the corpus to enable automatic speech recognition for the language.