LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#datasets

16 curated events
papersTODAY 04:00 UTC

Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text

Researchers present a parallel corpus pairing Arabic and Russian scientific writing, addressing a gap in resources for these two major research languages. The work also introduces a benchmark for evaluating large language models on the corpus, aimed at supporting cross-community knowledge exchange on sustainability topics. It is a revised arXiv submission in the computation and language category.

papersTODAY 04:00 UTC

arXiv Paper Probes Dataset Biases Behind Phantom Transfer

A new preprint on arXiv studies why a teacher model's bias can still pass to a student model even when the training data has had all overt mentions of that bias removed. The authors report that no data-level defense tested so far reliably detects or eliminates this residual, or "phantom," transfer. The work frames the phenomenon as a dataset-level problem rooted in subtle statistical traces rather than explicit labels.

papersTODAY 04:00 UTC

ParsVoice: Large Multi-Speaker Persian Speech Corpus for Text-to-Speech

Researchers released ParsVoice, described as the largest publicly available multi-speaker Persian speech corpus, aimed at addressing the scarcity of open speech-text data for the language. The dataset is intended to support work on multi-speaker text-to-speech, speech-language modeling, and other low-resource speech tasks. The paper was posted to arXiv under computer science categories.

papersTODAY 04:00 UTC

CVSS-X corpus adds English-to-28-language speech translation data

Researchers released CVSS-X, a large synthetic speech-to-speech translation corpus that reverses the direction of the earlier CVSS dataset. While CVSS translated 21 languages into English, CVSS-X supports translation out of English into 28 target languages. The corpus is described as a resource for training and evaluating multilingual speech translation systems.

papersSEP 10 04:00 UTC

AtlasNLP tracks which countries are represented in NLP datasets

Researchers have released AtlasNLP, a resource that maps how well individual countries are covered across NLP training datasets. The work addresses the fact that geographic metadata is seldom attached to datasets, leaving gaps in coverage hard to detect. The authors position the atlas as a tool for guiding data collection and informing AI policy decisions.

papersSEP 10 04:00 UTC

TimeCues Studio: A Workspace for Music Annotation and Algorithm Prototyping

A new arXiv paper introduces TimeCues Studio, a workspace for placing precise annotations such as points, segments, and loops on music recordings. The tool supports both hand labeling and algorithmic prototyping, targeting the shortage of annotated training data that limits machine-learning approaches in multimedia applications.

papersSEP 10 04:00 UTC

Study proposes perturbation-sensitive selection for medical QA rationales

A new arXiv paper addresses the scarcity of high-quality rationales in medical question-answering datasets, where answer labels are plentiful but explanations are expensive to validate. The authors reframe the data acquisition problem as deciding which already-labeled questions warrant rationales, using a perturbation-sensitive selection criterion. The approach aims to improve QA robustness by targeting rationale annotation where it has the greatest effect.

papersSEP 10 04:00 UTC

Dataset of 1,482 Annotated Tweets Released for Enthymeme Detection in Political Discourse

Researchers have released an annotated collection of 1,482 tweets on divisive political topics aimed at identifying enthymemes, which are arguments that leave a premise or conclusion unstated. Since labeling such incomplete arguments is widely considered a subjective task, the resource is intended to support more consistent research into detecting implicit reasoning in persuasive online text.

papersSEP 10 04:00 UTC

New 41B-token European Portuguese web corpus introduced in arXiv paper

Researchers have assembled a 41-billion-token collection of web text focused on European Portuguese, addressing the difficulty of separating it from Brazilian Portuguese in large-scale crawls. The paper describes an efficient processing pipeline that handles dialectal overlap at scale to yield a dataset intended for production training use. The work appears as a preprint cross-listed in language computation and AI categories.

papersSEP 10 04:00 UTC

Minimal-pair dataset probes whether language models grasp light-verb constructions

Researchers built a minimal-pair dataset that tests whether language models can tell light-verb constructions like 'make a decision' apart from full predicate uses of the same verbs, such as 'make a cake'. The work targets phraseological competence, examining how verb meaning shifts in multiword expressions. By comparing model behavior on nearly identical sentence pairs, the benchmark aims to reveal how deeply models represent this distinction.

papersSEP 10 04:00 UTC

MultiSynt/MT: open synthetic corpus offers 4.8T tokens in 36 languages for multilingual pretraining

Researchers have introduced MultiSynt/MT, an open synthetic parallel dataset totaling roughly 4.8 trillion target-language tokens spanning 36 languages. Because web-scale training data is heavily skewed toward English, the corpus is intended to give model builders far more multilingual material for pretraining large language models. The work is described in a paper posted on arXiv.

papersSEP 12 04:00 UTC

Paper Examines How Prevalence Drives Precision in Detector-Built Datasets

A new arXiv paper argues that when datasets are built by running a detector, heuristic, or model over candidate pools, the resulting label precision depends on the true-positive rate within each pool rather than on detector quality alone. The authors apply Bayes' rule to show how this prevalence effect introduces hidden contamination into detector-defined datasets, a problem they describe as silent. The work suggests dataset builders should account for pool-level prevalence when estimating or reporting precision.

papersSEP 12 04:00 UTC

OpenResearcher: Open Pipeline for Deep Research Agent Trajectory Synthesis

A new arXiv paper introduces OpenResearcher, a fully open pipeline for generating the long-horizon training data needed by deep research agents, which must interleave search, evidence collection, and multi-step reasoning. The authors note that current data collection approaches depend on proprietary web APIs, which restricts how far such datasets can scale. The work is an updated cross-list submission on arXiv cs.AI.

papersSEP 12 04:00 UTC

E-CONAN Benchmark Suite Targets Arabic Textual Entailment and Inference

A new arXiv paper introduces E-CONAN, a set of benchmarks covering entailment, contradiction and neutral relations for Arabic natural language inference. The authors frame the work as a response to the limited resources available for Arabic compared with English and other well-served languages, noting that inference models are a component of many downstream NLP applications. The datasets are intended to support training and evaluation of Arabic inference systems.

papersSEP 12 04:00 UTC

arXiv paper introduces GitSkills, a dataset of agent skills collected from GitHub

A new arXiv paper presents GitSkills, a dataset built from GitHub repositories that package agent skills as folders containing a SKILL.md instruction file, sometimes with helper scripts and reference material. The work focuses on skills that language-model agents load when they decide a task matches a skill's description. It is a replacement submission (v2) in the cs.AI category.