LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#audio

21 curated events
papersTODAY 04:00 UTC

AURA: Unified Multimodal Framework for Conversational Music Editing

Researchers present AURA, a multimodal framework that lets users edit music through a back-and-forth conversation rather than one-off commands. Existing instruction-guided editors handle each request in isolation, which makes it hard to iteratively refine a track. AURA is designed to keep track of prior edits so a session can build progressively toward a desired result.

papersTODAY 04:00 UTC

Predictive audio representations for early detection of occluded objects

A new arXiv preprint proposes using predictive audio representations to spot and track hidden dynamic objects before they come into view. The authors argue that occluded traffic agents can appear too late for detection systems, and that audio cues offer an earlier warning signal. The work targets safety-critical settings such as autonomous driving.

papersTODAY 04:00 UTC

SpecAugment-Patch Merging Proposed to Speed Up Audio Spectrogram Transformer Training

A new arXiv paper introduces SpecAugment-Patch Merging, a method that masks input spectrograms at the patch level before positional embeddings are added, then merges patches to cut computation. The authors describe it as a simple approach to accelerate training of Audio Spectrogram Transformers. The work is categorized under machine learning research.

papersTODAY 04:00 UTC

Audio language models track speakers via text backbone attention, study finds

A new study examines how audio language models attribute speech to the correct speaker, finding accuracy of only 6 to 16 percent on a six-speaker task, below random guessing. The authors show that speaker tracking relies on attention heads in the model's text backbone, and that altering a subset of those heads shifts which speaker the model retrieves.

papersTODAY 04:00 UTC

arXiv Paper Proposes MAST for Label-Efficient Biodiversity Sound Detection

Researchers present MAST, a framework for detecting animal vocalizations in passive acoustic recordings while relying on far fewer labeled examples. It combines masked audio pretraining with self-training to improve robustness and transfer across recording sites. The approach aims to make large-scale biodiversity monitoring more practical where expert annotation is costly.

papersTODAY 04:00 UTC

arXiv Paper Proposes Frame-Level Grounding for Audio-Language Model Temporal Perception

A new arXiv preprint addresses the limited ability of large audio-language models to pinpoint when specific sounds occur within a recording. The authors propose adding frame-level grounding during training so these models can localize audio events more precisely rather than only describing clips in broad terms. The work targets fine-grained temporal understanding, a known weak point for current audio-language systems.

papersTODAY 04:00 UTC

Paper Revisits Scaling and Training Objectives for Procedural Audio Pre-training

A new arXiv preprint examines how procedural audio should be scaled when used as a data source for learning transferable audio representations. The authors also ask whether training choices originally developed on natural audio still hold when the source is procedurally generated. The work aims to clarify design principles for procedural audio pre-training, an area the authors say lacks settled guidance.

papersTODAY 04:00 UTC

LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Audio

Researchers present LoSATok, a tokenizer designed to serve both audio understanding and generation tasks within a single framework. The work argues that understanding benefits from high-level semantic features while generation needs both semantic and acoustic detail, and existing unified tokenizers encode both in high-dimensional spaces. LoSATok instead uses a low-dimensional representation to handle cross-domain audio.

papersTODAY 04:00 UTC

BGM2Pose Estimates 3D Human Pose Using Background Music as Sensing Signal

Researchers propose BGM2Pose, a method that estimates a person's 3D pose without physical contact by using ordinary music playing nearby as an active sensing signal. The approach aims to avoid the intrusive chirp signals and impractical setups used by earlier acoustic pose-estimation systems. A revised version of the paper is available on arXiv.

papersTODAY 04:00 UTC

arXiv Paper Models Music-Taste Correspondences with Normalized Dataset

A new arXiv preprint describes a multimodal dataset covering crossmodal correspondences between music and taste, a topic tied to intangible cultural heritage. The authors apply dataset normalization and perceptual validation so that computationally modeled links between sound and flavor hold up against human perception. The work targets applications in museums, exhibitions and gastronomic tourism.

papersTODAY 04:00 UTC

DuoTok Paper Proposes Dual-Track Music Tokenization for Vocal-Accompaniment Generation

Researchers present DuoTok, a tokenization method that separates vocal and accompaniment streams while keeping source information intact for multi-track music generation. The approach aims to balance acoustic detail, sequence modelability, and cross-track structure that existing codecs trade off against each other. A revised version of the preprint is now available.

papersSEP 11 04:00 UTC

Active noise cancellation adapted for open-ear smart glasses

A preprint describes an active noise cancellation approach designed for open-ear smart glasses, where the usual in-ear error microphone cannot be used. Conventional ANC depends on measuring residual sound at the ear canal, which the authors say is not feasible for this form factor. The work is posted on arXiv as a cross-list replacement and has not been peer reviewed.

papersSEP 11 04:00 UTC

PitchFlower: Flow-Based Neural Audio Codec With Pitch Controllability

Researchers introduce PitchFlower, a flow-based neural audio codec designed to separate pitch information from other audio content. The method applies a training perturbation in which F0 contours are flattened and randomly shifted, encouraging the model to disentangle pitch from the rest of the signal. This yields explicit, adjustable pitch control when synthesizing or compressing audio.

papersSEP 10 04:00 UTC

Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

A new arXiv paper addresses voice-cloning fraud in which just a sentence or two of a genuine multi-speaker conversation is swapped for synthetic audio. Rather than giving one real-or-fake verdict for an entire recording, the proposed method pinpoints the exact time spans of fake speech without needing labelled examples of the targeted fakes.

papersSEP 10 04:00 UTC

Annotator disagreement in temporal laughter localization found to be structured, not random

A new paper studies how human labelers disagree about the precise onset and boundaries of laughter when annotating audio. The authors show that this disagreement follows systematic patterns rather than acting as random noise, challenging the common practice of scoring temporal laughter localization against a single reference annotation. They argue that evaluation protocols should instead account for the structured nature of annotator disagreement.

papersSEP 10 04:00 UTC

Bayesian Tracking Guides Deep Spatial Filters to Extract Moving Speakers in Real Time

Deep spatially selective filters deliver high-quality, real-time speech enhancement for stationary speakers whose directions are known, but they struggle when sources move. A new arXiv paper introduces an autoregressive guidance scheme built on Bayesian speaker tracking that lets these filters follow moving speakers from only their initial positions. The approach preserves the efficiency and enhancement quality of the underlying architecture in dynamic scenes.

papersSEP 10 04:00 UTC

PRISM-Bench: Audio-Centric Benchmark for Evaluating Text-to-Audio-Video Generation

Researchers have introduced PRISM-Bench, a diagnostic benchmark for text-to-audio-video generation models that places the audio modality at the center of evaluation. The paper argues that prior benchmarks tend to treat sound as a minor add-on to video quality metrics or test it separately from the combined audiovisual output. The new benchmark is intended to give a fuller picture of how generative systems handle audio together with visuals.

papersSEP 12 04:00 UTC

Study compares methods to reduce catastrophic forgetting in sound event classification

A new arXiv paper examines ways to keep machine learning models from forgetting previously learned classes when trained incrementally on sound event classification. It evaluates architectural and regularization strategies using the FSD50K dataset and related benchmarks. The work is a research investigation rather than a released product.

papersSEP 12 04:00 UTC

Paper Proposes Adaptive Perturbation Selection for Contrastive Audio Decoding

A new arXiv paper addresses hallucination in large audio-language models, where models sometimes let language priors override what is actually heard in the audio. The authors propose a method that adaptively selects perturbations for contrastive decoding, a training-free approach, arguing that existing techniques rely on crude perturbations such as masking or added noise. The work aims to improve how reliably these models ground their outputs in acoustic evidence.