LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

audio-language-models

topic7 events
papersTODAY 04:00 UTC

arXiv Paper Proposes Machine Unlearning for Speech Question Answering Models

A new arXiv preprint examines how large audio-language models can be made to forget sensitive information they may have memorized during training. The work focuses on the speech question-answering setting, where such models have shown strong performance but also carry privacy risks. The authors frame machine unlearning as a way to reduce unintended retention of private data in these systems.

papersTODAY 04:00 UTC

Audio language models track speakers via text backbone attention, study finds

A new study examines how audio language models attribute speech to the correct speaker, finding accuracy of only 6 to 16 percent on a six-speaker task, below random guessing. The authors show that speaker tracking relies on attention heads in the model's text backbone, and that altering a subset of those heads shifts which speaker the model retrieves.

papersTODAY 04:00 UTC

Study compares end-to-end models for clinical SOAP note generation from audio

A new arXiv paper examines how well audio-language models can turn long doctor-patient conversations into structured SOAP clinical notes. The authors compare lightweight and heavyweight end-to-end approaches, noting that while cascaded speech recognition pipelines remain strong, end-to-end models tend to lose information or produce hallucinations. The work targets the modality gap in long-form clinical audio.

papersTODAY 04:00 UTC

arXiv Paper Proposes Frame-Level Grounding for Audio-Language Model Temporal Perception

A new arXiv preprint addresses the limited ability of large audio-language models to pinpoint when specific sounds occur within a recording. The authors propose adding frame-level grounding during training so these models can localize audio events more precisely rather than only describing clips in broad terms. The work targets fine-grained temporal understanding, a known weak point for current audio-language systems.

papersTODAY 04:00 UTC

arXiv paper proposes emotion regulation framework for empathetic speech dialogue in audio-language models

A new arXiv preprint introduces ER-EDF, a framework that draws on psychological theories of emotion perception and regulation to guide empathetic responses in spoken dialogue systems built on large audio-language models. The work aims to improve how such systems both recognize a speaker's emotional state and regulate their own generated reply. It is a research contribution and has not been presented as a product or model release.

papersSEP 12 04:00 UTC

Paper Proposes Adaptive Perturbation Selection for Contrastive Audio Decoding

A new arXiv paper addresses hallucination in large audio-language models, where models sometimes let language priors override what is actually heard in the audio. The authors propose a method that adaptively selects perturbations for contrastive decoding, a training-free approach, arguing that existing techniques rely on crude perturbations such as masking or added noise. The work aims to improve how reliably these models ground their outputs in acoustic evidence.

papersSEP 12 04:00 UTC

Switch-Aware Evaluation of ASR and Audio Language Models on English-Yoruba Code-Switched Speech

A new arXiv preprint argues that word error rate alone hides important failures when speech recognition systems and audio language models handle code-switched speech. The authors propose an evaluation method that accounts for language switches, and apply it to English-Yoruba audio, a low-resource pair with diacritics. They find that strong monolingual benchmark scores do not carry over to this setting.