LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#nlp

40 curated events
papersTODAY 04:00 UTC

Study Analyzes Self-Reported Limitations in NLP Research

A new arXiv paper examines the Limitations sections that top-tier NLP conferences have required since late 2022, noting that the volume of accepted papers has produced a corpus too large for manual review. The authors analyze these self-reported limitations to characterize what researchers themselves identify as the constraints of their work. The study aims to make this body of disclosures more tractable to assess at scale.

papersTODAY 04:00 UTC

Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text

Researchers present a parallel corpus pairing Arabic and Russian scientific writing, addressing a gap in resources for these two major research languages. The work also introduces a benchmark for evaluating large language models on the corpus, aimed at supporting cross-community knowledge exchange on sustainability topics. It is a revised arXiv submission in the computation and language category.

papersTODAY 04:00 UTC

Framework Proposed to Study Interdisciplinary Discourse in Scholarly Publications

A new arXiv paper presents a framework for examining how scholarly publications meaningfully integrate ideas from multiple disciplines. Motivated by the growth of interdisciplinary work and institutional incentives supporting it, the authors argue that existing computational approaches are insufficient for capturing such cross-disciplinary engagement. The paper aims to offer a structured way to analyze interdisciplinary discourse in research writing.

papersTODAY 04:00 UTC

arXiv paper studies lexicon structure and compositionality in evolutionary semantics

A revised preprint on arXiv examines how the structure of a lexicon relates to the compositional way sentence meanings are built from word meanings. The author notes that much existing work on semantic universals assumes either fixed signal structures in lexicons or holistic composition that cannot be interpreted. The work frames these questions within evolutionary semantics.

papersTODAY 04:00 UTC

Conformance-Driven Iterative Refinement for Natural-Language to SysMLv2 Translation

A new arXiv paper proposes a method for converting natural-language specifications into SysMLv2, the textual modeling language standardized for model-based systems engineering. The approach refines candidate translations iteratively, using conformance checks to guide corrections. It aims to lower the barrier to producing formal system models that capture requirements, structure, and behavior.

papersSEP 11 04:00 UTC

Paper Proposes Training-Free Method for Analyzing SEC Filings

A new arXiv preprint describes a method for corporate and financial-disclosure analysis that avoids training and cross-model alignment steps. The authors argue that dense text embeddings and large language models struggle with context limits, hallucination risk, compute cost, and inconsistent vector spaces across independently trained models. The work is demonstrated on SEC filings, with a revised version now posted.

papersTODAY 04:00 UTC

Paper Proposes Method to Restore Zipfian Frequency Patterns in Unsupervised Term Discovery

A revised arXiv paper examines how unsupervised term discovery systems segment unlabelled speech and group the resulting units into candidate word or syllable types. The authors note that real lexicons follow a Zipfian frequency distribution, but the widely used centre-based clustering approach does not reproduce it. Their work introduces a method aimed at recovering that distribution in the discovered lexicon.

papersTODAY 04:00 UTC

arXiv Paper Examines Factual Errors in Human-Written Text for Detection

A new arXiv study looks at how factual mistakes appear in text written by people, aiming to inform automatic detection of incorrect spans. The authors argue that factual error detection has long been a key research problem, but interest has shifted with the rise of large language models. The work connects analysis of human-written errors to building systems that can flag factual inaccuracies.

papersTODAY 04:00 UTC

Quantum-Classical Hybrid Model Tested for Paraphrase Detection

Researchers evaluated a 10-qubit hybrid quantum-classical variational circuit with 2,148 parameters on paraphrase detection tasks, using MRPC and Quora Question Pairs among three benchmarks. The work reports performance, robustness, and entanglement results, aiming to fill a gap in empirical validation of quantum machine learning for natural language tasks. The paper is an arXiv preprint and has not been peer-reviewed.

papersTODAY 04:00 UTC

Study Extends Retrieval-Head Analysis to Multilingual Language Models

Researchers extend prior work on retrieval heads — attention heads that pull information out of context — from English to multilingual models. They identify retrieval heads and a separate class of retrieval-transition heads, and report that behavior differs across languages. The work is a revised arXiv preprint in computation and language.

papersTODAY 04:00 UTC

arXiv paper proposes document topic alignment metrics for topic models of health social media

A new arXiv preprint argues that existing evaluations of topic models for short social media texts rely almost entirely on measures of the generated topics themselves. The author proposes metrics that also assess how well documents align with their assigned topics. The work targets public health communication data drawn from social media.

papersTODAY 04:00 UTC

Automated Speaking Assessment Model Adds Relevance and Grammar Error Cues

A new arXiv paper proposes improvements to automated speaking assessment systems, which currently underuse content relevance signals such as image or exemplar prompts. The authors also address shallow grammar analysis by incorporating more granular grammar error type information into the evaluation. The work targets multi-aspect scoring of spoken language for assessment purposes.

papersTODAY 04:00 UTC

arXiv Paper Proposes Better Event Candidate Acquisition for Event Linking

A new arXiv preprint addresses event linking, the task of matching event mentions in text to knowledge base entries or flagging them as absent from the KB. The authors argue that existing architectures still suffer from weak candidate acquisition, especially when mentions are short or ambiguous. The paper introduces an approach aimed at improving how candidate events are gathered before linking.

papersTODAY 04:00 UTC

arXiv paper reviews barriers to trustworthy AI use in cancer genomics

A revised arXiv paper examines how AI and natural language processing are used to extract and interpret biomedical knowledge in cancer genomics. It argues that clinical adoption has lagged and lays out the barriers, risks and possible pathways needed for trustworthy translation into routine oncology. The work is a review and framing contribution rather than a report of new model results.

papersTODAY 04:00 UTC

WMT26 Builds Pseudo-References for 10 Reference-Free MT Language Pairs

Researchers describe the process used to create pseudo-references for the WMT26 General Machine Translation task, where ten language pairs lack any human translations or post-edited outputs. The work also covers six additional language pairs that do have references, aiming to give systems a consistent basis for automatic scoring. The paper details how these synthetic references were constructed and evaluated.

papersTODAY 04:00 UTC

Hybrid 1D-CNN-BiLSTM Framework Proposed for Biomedical Extractive Summarization

Researchers present a hierarchical hybrid model that combines one-dimensional convolutional networks with bidirectional LSTMs to produce extractive summaries of biomedical and clinical text. The approach avoids text generation entirely, sidestepping the factual hallucination risks that make abstractive large language models unreliable in medical settings. The work is posted as a preprint on arXiv.

papersTODAY 04:00 UTC

arXiv Paper Audits Commonsense Reasoning Benchmarks Used for LLM Evaluation

A new arXiv paper argues that commonsense reasoning in language models is typically measured with multiple-choice benchmarks such as HellaSwag and PIQA, yet these benchmarks themselves are rarely scrutinized. The authors propose a more comprehensive approach to evaluating the benchmarks, questioning how well they actually capture the capability they claim to test. The work is framed as a meta-evaluation of standard commonsense reasoning tests.

papersTODAY 04:00 UTC

ScorePrompts system lets users query symbolic music scores in natural language

Researchers present ScorePrompts, an interactive system that accepts an uploaded musical score and returns natural-language descriptions of its structure. Users can ask questions about specific passages, and the system displays the matching analysis results in staff notation. The work was posted as an arXiv preprint in the computation and language category.

papersTODAY 04:00 UTC

arXiv preprint links emotional grounding to task-oriented dialogue learning

A revised preprint on arXiv examines task-oriented dialogue systems, which help users complete goals through natural language. The authors argue that emotional factors shape learning in these systems, and that success depends not only on completing tasks but also on sustaining positive emotional exchanges and conveying information accurately.

papersTODAY 04:00 UTC

Controlled Study Reexamines What Drives Coreference Resolution Performance

A new arXiv paper revisits comparisons between state-of-the-art coreference resolution systems. Because every leading system fine-tunes a pretrained language model, the authors ask whether differences in scores come from the underlying language model or from task-specific design choices. The work presents a controlled reevaluation to separate those factors.

papersTODAY 04:00 UTC

arXiv Paper Analyzes Topology of Dependency Trees Across 124 Languages

A research paper examines syntactic dependency trees drawn from 124 languages chosen for typological, genetic and geographic diversity. The author investigates the structural properties of these trees, noting that the organizing principles behind their topology are still not well understood. The work aims to identify regularities that hold across languages.

papersTODAY 04:00 UTC

IndicQE-APE Benchmark Consolidates Quality Estimation and Post-Editing for Indic Languages

Researchers have assembled IndicQE-APE, a single benchmark that brings together scattered Indic-language resources for quality estimation and automatic post-editing. The dataset draws on WMT shared task data from 2020 through 2024, allowing models to be trained and evaluated across multiple tasks and language pairs under consistent conditions. The goal is to remove the fragmentation that previously made cross-task and cross-language comparison difficult.

papersTODAY 04:00 UTC

Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection

A revised arXiv preprint proposes a mechanism-oriented taxonomy of indirect linguistic expressions, the disguised phrasing such as algospeak and euphemisms that users adopt to hide sensitive meaning from platforms. The work organizes these encoding strategies by how they work rather than how they look, aiming to give LLM-based detection systems a more general basis for spotting obfuscated content. It targets the gap between surface-form moderation filters and adversarial evasion in social media text.

papersTODAY 04:00 UTC

Paper Examines Adequacy-Fluency Tradeoff in MT Meta-Evaluation

A new arXiv paper analyzes how meta-evaluation of machine translation must balance alignment with adequacy versus fluency, noting that the preferred balance shifts depending on which translation systems are included in the evaluation set. Because those system sets are typically small and filtered, the authors propose parameterizing this balance explicitly. The work appears in the cs.AI and cs.LG cross-listings.

papersTODAY 04:00 UTC

arXiv paper proposes automatic pregroup supertagging to scale Hindi quantum NLP

A new arXiv preprint describes an approach to scaling quantum natural language processing for Hindi by automating pregroup supertagging. The work builds on prior Hindi-specific pregroup grammars that map grammatical structure into diagrammatic and quantum circuit representations. Automating the tagging step is intended to reduce the manual effort needed to apply these grammars at larger scale.

papersTODAY 04:00 UTC

Reward-Guided Self-Training Improves Pronoun Translation in Context-Aware MT

A new arXiv paper examines how context-aware machine translation systems handle pronouns, which depend on discourse information that ordinary fine-tuning tends not to emphasize. The authors propose ProNMT, a self-training approach that uses reward signals to iteratively refine these sparse, context-sensitive decisions while keeping overall translation quality balanced. The work targets the trade-off between general fluency and accurate pronoun-specific output.

papersTODAY 04:00 UTC

Pipeline builds citation-grounded causal graphs from humanitarian reports

Researchers describe a processing pipeline that turns large volumes of heterogeneous humanitarian reporting into structured disaster storylines and causal knowledge graphs, with each extracted claim tied back to its source citation. The work targets the first hours of a crisis, when the amount of incoming information typically outstrips what human analysts can review. The approach is presented as a way to make automated synthesis of relief-related reporting more traceable and verifiable.

papersTODAY 04:00 UTC

Study Finds Affixal Negation Cues Improve Language Model Negation Understanding

A new arXiv paper argues that research on negation in language models has focused too narrowly on a few common single-word cues such as "not" and "never". The authors examine a wider range of negation signals and report that affixal negations, formed through prefixes and suffixes, lead to better negation understanding than the cues usually studied. The work points to broader cue coverage as a way to address a persistent weakness in both LMs and LLMs.

papersTODAY 04:00 UTC

arXiv Paper Uses Transformer Ensembles to Detect Schwartz Values in News Sentences

A new arXiv study tackles multi-label classification of the 19 refined Schwartz human values across roughly 74,000 English news and manifesto sentences from the ValueEval'24 corpus. The authors focus on extreme label imbalance, where some values rarely appear, and combine transformer ensembles with an analysis of value hierarchies and moral presence. The work is framed as a methodological contribution to sentence-level value detection rather than a deployed product.

papersTODAY 04:00 UTC

Interpretable Recognition of Cognitive Distortions in Natural Language Texts

A new arXiv paper proposes classifying natural language texts along multiple factors using weighted structured patterns such as N-grams, while accounting for heterarchical rather than strictly hierarchical links between those patterns. The authors apply the method to detecting cognitive distortions, framing it as a socially impactful task, and emphasize that the approach keeps the decision process interpretable. The work appears as a cross-listed replacement submission in arXiv cs.AI and cs.LG.

papersTODAY 04:00 UTC

arXiv Paper Proposes Reward-Based Outcome Evaluation for Grammatical Error Correction

A new arXiv preprint argues that grammatical error correction systems are typically judged by how closely their edits or outputs match reference corrections, which can unfairly penalize valid alternative rewrites. The author proposes evaluating GEC by the outcome instead, using reward-based scoring that judges whether the resulting text is fluent and correct rather than counting edits. This reduces reliance on gold references and aims to better reflect real-world usefulness.

papersTODAY 04:00 UTC

Study Examines Reliability of LLM and Rule-Based Annotation on Turkish Narrative Corpus

A new preprint evaluates whether automatically generated narrative feature labels would be endorsed by human annotators. The authors compare LLM-based and rule-based annotation against human judgments across three studies using the Turkish-language Objective Projection corpus. The work contributes inter-rater reliability evidence for datasets that ship machine-generated annotations.

papersTODAY 04:00 UTC

Bangla Sentence Function Classification Corpus and Benchmark Released

Researchers present a new annotated corpus for classifying sentence functions in Bangla, a resource previously lacking for the language. The work benchmarks several models on the task and adds interpretability analysis of their predictions. Such sentence-type identification supports dialogue systems, speech synthesis, and machine translation.

papersTODAY 04:00 UTC

arXiv paper proposes method to classify generalisation claims in NLP research

A new arXiv preprint argues that generalisations are widespread in scientific writing yet carry ambiguous meaning, making them hard for readers and automated systems to interpret consistently. The authors propose an automated approach for identifying and sorting such claims by how broad they are, with the goal of surfacing research that leans too heavily on sweeping statements. The work sits at the intersection of natural language processing and research evaluation.

papersTODAY 04:00 UTC

arXiv paper proposes attention calibration for position-fair dense retrieval

A revised arXiv paper addresses a known weakness in dense retrieval: compressing a passage into a single embedding tends to weight early text more heavily, so retrieval quality drops when the relevant span sits later in the passage. The authors propose attention calibration as a way to reduce this positional bias, building on earlier inference-time approaches. It is a research contribution rather than a released product.