LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#dataset

15 curated events
papersTODAY 04:00 UTC

Paper Maps Uthmani Quranic Script to Standard Arabic for NLP

A new arXiv paper addresses the mismatch between the Uthmani orthography used in printed Qurans and the Standard (Imla'i) Arabic that most Arabic NLP tools support, since the two forms differ at the byte level. The authors present a corpus-aligned mapping between Uthmani and Standard word forms, along with a deterministic validator for recitation. The work targets improved text processing for Quranic Arabic within existing Arabic language tooling.

papersTODAY 04:00 UTC

Paper Tackles Lost-in-the-Middle Problem in Long-Text Generation

A new arXiv paper addresses how large language models tend to ignore information placed in the middle of long contexts, a problem studied mostly for retrieval tasks rather than long-input-to-long-output generation. The authors introduce a synthetic dataset and evaluation framework for this setting and propose a mitigation approach. The work is a revised cross-listing (v2) on arXiv cs.AI.

papersSEP 11 04:00 UTC

arXiv Paper Presents Dataset and Model for Imputing River Water Surface Elevation

Researchers posted an arXiv paper describing a dataset and machine learning model for estimating water surface elevation across river networks where gauge coverage is sparse. The work frames river systems as a large spatiotemporal graph and targets applications such as flood forecasting and water resource management. The paper appeared as a new submission and has since been revised.

papersTODAY 04:00 UTC

SomBench: Benchmark Dataset for Machine Learning in Lunar Science

Researchers introduced SomBench, a benchmark dataset aimed at machine learning applications in lunar science. The work addresses the difficulty of combining observations from multiple lunar orbital missions, which differ in sampling, projection, and instrument characteristics. It is intended to give researchers a common basis for evaluating models on lunar data.

papersTODAY 04:00 UTC

arXiv Paper Models Music-Taste Correspondences with Normalized Dataset

A new arXiv preprint describes a multimodal dataset covering crossmodal correspondences between music and taste, a topic tied to intangible cultural heritage. The authors apply dataset normalization and perceptual validation so that computationally modeled links between sound and flavor hold up against human perception. The work targets applications in museums, exhibitions and gastronomic tourism.

papersTODAY 04:00 UTC

Transformer Model and Mandarin Speech Dataset Target Audio-Based Kinship Verification

Researchers introduce CONVTRAP-TN, a transformer-based architecture designed to determine whether two speakers share a first-order family relationship using only audio. The work also releases a new uncontrolled Mandarin kinship speech dataset and reports an ablation study on the model's components. Kinship verification from voice is a relatively underexplored area, and the Mandarin dataset addresses a gap in multilingual resources.

papersTODAY 04:00 UTC

Open Persian speech corpus Neyshekar released with 99 hours of audio

Researchers have published Neyshekar, an openly available Persian read-speech corpus intended to cover formal and informal speech, named entities, and longer sentences. Version 6 contains 62,279 validated recordings totaling 99.02 hours, contributed by 190 speakers. The dataset is aimed at supporting automatic speech recognition work in Persian.

papersSEP 10 04:00 UTC

BaltiVoice: 16.8-hour speech corpus and fine-tuned Whisper ASR system for Balti

Researchers have released BaltiVoice, a 16.8-hour read-speech corpus with 10,060 validated utterances for Balti, a Tibetic language spoken in Gilgit-Baltistan, Pakistan. The language previously had no publicly available speech recognition resources, making this the first open dataset and ASR model for Balti. The team fine-tuned OpenAI's Whisper architecture on the corpus to enable automatic speech recognition for the language.

papersSEP 10 04:00 UTC

SloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset

SloMoDeblur is a large-scale dataset designed to advance motion blur removal in smartphone photography. Its creators argue that current deblurring benchmarks fall short in size, resolution, and relevance to real-world mobile imaging conditions. The paper has been posted on arXiv and cross-listed under the AI and machine learning categories.

papersSEP 10 04:00 UTC

V2TATC: Voice-Trajectory Embedding and Dataset for Air Traffic Control Situational Awareness

Researchers present V2TATC, a machine-learning approach that jointly embeds controller voice communications and aircraft trajectory data to support situational awareness in air traffic control. The work also introduces a new dataset to enable development of scalable decision-support tools as traffic, particularly at low altitudes, grows in the US National Airspace System.

papersSEP 10 04:00 UTC

SAFER-Activities Dataset Introduced for Fall Detection and Routine Activity Recognition

Researchers have released SAFER-Activities, a dataset aimed at improving action recognition in smart healthcare monitoring systems, with a focus on detecting falls among people with mobility challenges. It addresses limitations of existing clip-based datasets by supporting assessment of both fall events and routine activities for timely intervention.

papersSEP 12 04:00 UTC

ChronoBerg corpus aims to give language models long-term temporal structure

A new arXiv paper introduces ChronoBerg, a resource designed to capture how language changes over time and to help foundation models reason about temporal context. The authors argue that while existing training corpora are broad, they often lack the long-term chronological structure needed for time-aware language understanding. The work is posted as a cross-list replacement on arXiv cs.AI.

papersSEP 9 13:22 UTC

DeepMind releases AlphaGenome Atlas covering all 9 billion human DNA letter changes

Google DeepMind has published a dataset called the AlphaGenome Atlas that predicts the possible consequences of roughly nine billion single-letter variations in the human genome. The collection is about one petabyte in size, which the outlet notes is more than 30 times larger than the AlphaFold database. A case involving epilepsy is cited as an example of how the resource was used.