LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#research

40 curated events
papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

Paper Argues Formal Language Properties Should Constrain Neural Models

A new arXiv preprint argues that current neuroscience and language-model research mostly checks whether brain signals or model layers can predict annotated linguistic variables, which shows correlation but not how language is actually implemented. The author proposes instead deriving what a neural system must be capable of from the formal properties of language itself, then treating those requirements as constraints on neural dynamics. This reframes the goal from prediction accuracy toward identifying the mechanisms a system needs in order to support language.

papersTODAY 04:00 UTC

arXiv Paper Proposes Formal Theory of Mind Based on Information Access History

A revised arXiv preprint introduces a formal framework for theory of mind that derives an agent's beliefs from the history of information they encountered, rather than assuming those beliefs are already known. The account models factors such as the order of exposure, the source of information, and its perceived credibility. This extends prior formal treatments that generally take beliefs as given inputs.

papersTODAY 04:00 UTC

Machine Learning Method Screens Point Defects in Semiconductors

A new arXiv paper describes a machine-learning approach for prescreening point defects in semiconductors, aimed at applications in power electronics and quantum technologies. The work positions itself as an alternative to the high-throughput density-functional theory calculations that have traditionally dominated defect exploration. The authors frame the method as part of a broader shift away from conventional simulation workflows.

papersTODAY 04:00 UTC

VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition

A new arXiv paper introduces VoiceCodeBench, a benchmark that checks whether speech recognition transcripts reproduce exact written values rather than just scoring well on word error rate. It focuses on structured tokens such as identifiers, file paths, and commands, which voice-driven workflows often need verbatim. The work argues that WER alone does not capture whether these exact tokens survive transcription.

papersTODAY 04:00 UTC

arXiv Paper Proposes Homeostatic Continual Learning for AI Agents

A new arXiv preprint introduces a method called Homeostatic Continual Learning that aims to let an AI agent keep learning as its environment changes without losing previously acquired knowledge. The approach targets catastrophic forgetting, a long-standing problem in continual learning research. The abstract provides only a brief description of the method's core mechanism.

papersTODAY 04:00 UTC

Study Separates Inference Topology From Diversity in Multi-Agent LLM Emotion Detection

A new arXiv paper examines multi-agent LLM pipelines by treating two design choices as independent variables: how agent calls are wired together and where the differences between agents come from. The authors evaluate this on multilingual, low-resource emotion detection, where labeled data is scarce. The goal is to clarify which gains come from the structure of the agent network versus from the diversity introduced between agents.

papersTODAY 04:00 UTC

Modular tokenizers proposed for efficient multilingual LLMs

A new arXiv paper argues that multilingual LLMs suffer from using one shared vocabulary across all supported languages, which produces uneven compression rates between languages. The authors also note that large embedding and output matrices raise memory demands and slow processing. Their proposed modular tokenizer design assigns separate tokenization components per language to address both issues.

papersTODAY 04:00 UTC

arXiv paper studies lexicon structure and compositionality in evolutionary semantics

A revised preprint on arXiv examines how the structure of a lexicon relates to the compositional way sentence meanings are built from word meanings. The author notes that much existing work on semantic universals assumes either fixed signal structures in lexicons or holistic composition that cannot be interpreted. The work frames these questions within evolutionary semantics.

papersTODAY 04:00 UTC

Paper Argues LLM-Judge Calibration in Biomedical ML Needs Four Separate Ledgers

A new arXiv paper examines how synthetic perturbations are often used as cheap calibration data for LLM evaluators in biomedical machine learning, where expert review is limited. The authors argue that a planted mutation key should not be treated as either a detector output or automatically as human ground truth. They propose formalizing four distinct ledgers to make reporting of calibration results more responsible.

papersTODAY 04:00 UTC

Paper proposes agent-controlled goal selection and termination in hierarchical RL

A new arXiv paper examines the agent-centric general value function (ACGVF) approach, which shifts two design choices from the environment or system designer to the learning agent itself. Under this construction, the agent decides both which goal to pursue and when to treat a goal as completed. The note builds on prior work by Tasse et al. (2026) on goal-based hierarchical reinforcement learning.

papersTODAY 04:00 UTC

Knowledge-enhanced approach proposed for single-cell foundation models

A new arXiv paper examines how single-cell foundation models depend on large transcriptomic pretraining datasets, noting that adding more data brings diminishing returns at rising computational cost. The authors' data scaling analysis suggests incorporating structured biological knowledge could improve efficiency instead of relying on scale alone. The work points toward knowledge-enhanced pretraining as an alternative direction for the field.

papersTODAY 04:00 UTC

Study questions realism of language-model agents in farming decision simulations

A new arXiv paper examines whether language-model agents can credibly stand in for human respondents in surveys and social simulations. The authors argue that judging realism from population averages or distributional similarity can be misleading, an effect they call the "average-farmer illusion." Their experiments test what such aggregate evidence actually demonstrates about individual-level behavior.

papersTODAY 04:00 UTC

VisInteract benchmark targets interactive text-to-visualization under flawed queries

A new arXiv preprint introduces VisInteract, an approach and benchmark aimed at text-to-visualization systems that must cope with ambiguous, incomplete, or factually wrong user requests. The authors note that current systems typically assume well-specified inputs and generate a chart in a single pass. Their work instead frames chart creation as a dynamic, interactive process that can correct and refine imperfect queries.

papersTODAY 04:00 UTC

Paper Proposes Framework for Judging When Synthetic Survey Data Is Trustworthy

A new arXiv paper argues that the debate over synthetic data in marketing research has been stuck between two extremes: treating large language models as a replacement for human survey respondents, or rejecting them outright. The authors say the more useful question is when synthetic respondents can be trusted, and they outline how that reliability should be evaluated. The work focuses on marketing research but touches on broader issues of validating model-generated data.

papersTODAY 04:00 UTC

SkillAtlas: An Attack Trace Library for Agent Skills

Researchers present a library of attack traces aimed at reusable skills for language-model agents. The work argues that risks in agent skills surface through model decisions, user context, tool calls, and execution feedback rather than through fixed signatures or a single sandboxed run, which limits existing static and dynamic analysis methods. The library is intended to help catalog and study these behaviors.

papersTODAY 04:00 UTC

arXiv Paper Proposes Accountability Engineering Approach for AI Deployment

A new arXiv paper argues that current evaluation practices for AI systems are too focused on models themselves, which is insufficient for systems used in healthcare, finance, and public services. The authors outline a vision for "AI deployment accountability engineering," aimed at embedding accountability into how safety-critical socio-technical systems are built and assessed. The work is a position/vision paper rather than an empirical study.

papersTODAY 04:00 UTC

LoRA Study Maps Asymmetric Transfer Across Tasks and Languages

Researchers ran a controlled LoRA fine-tuning experiment to see how gains from training on one task or language carry over to others. The work finds that transfer between tasks and languages is uneven rather than symmetric, meaning improvements in one setting do not reliably help elsewhere. The findings point to limits in assuming that fine-tuning benefits generalize broadly across multilingual, multi-task models.

papersTODAY 04:00 UTC

Study traces how large language models represent animacy

A arXiv paper examines where the concept of animacy is encoded inside large language models. The authors trace internal circuits tied to the animate/inanimate distinction, which involves verb-argument constraints and contextual cues beyond simple word-level features. The work is a revised version of a preprint in the cs.CL category.

papersTODAY 04:00 UTC

Paper proposes evolving context parameterization for large language models

A new arXiv paper addresses a limitation in context parameterization, a technique that lets language models absorb context into reusable parameters instead of reprocessing it for every query. The authors note that current approaches treat context as static and are therefore ill-suited to settings where information changes over time. Their work introduces a method for keeping those internalized parameters up to date as contexts evolve.

papersTODAY 04:00 UTC

arXiv Paper Proposes Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

A new arXiv paper argues that judging multi-agent LLM systems only by their final answers misses how those systems actually reach their results. The authors propose a diagnostic method grounded in human group research to distinguish process losses from assembly bonuses when LLM teams collaborate. This matters both for building better agent pipelines and for using LLM groups as stand-ins for human group behavior.

papersTODAY 04:00 UTC

arXiv Paper Benchmarks Model-Agnostic Keyframe Selection for Long Video MLLMs

A new arXiv preprint evaluates keyframe selection techniques that can be plugged into existing multimodal large language models without modifying them. The work targets the constraint that MLLMs cannot ingest every frame of a long video due to visual-token and compute limits, and compares the main families of approaches proposed to address this. The study positions keyframe selection as a model-agnostic add-on for improving long-video understanding.

papersTODAY 04:00 UTC

arXiv Paper Examines Capacity Limits of Reasoning via Superposition

A new arXiv preprint studies how much intermediate computation a single vector can carry when language models reason through continuous or recurrent methods rather than token-by-token chain-of-thought. The work frames multi-step reasoning as superposition, where partial computations are packed into hidden states, and analyzes the resulting capacity limits. It offers a theoretical lens on the trade-offs between explicit token-based reasoning and continuous latent approaches.

papersTODAY 04:00 UTC

Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election

A new preprint looks at how generative AI writing tools and the language models behind them may shape the political information voters encounter ahead of Sweden's 2026 election. The authors argue that as these assistants become a common way to gather information, their tendency to favor certain issues or viewpoints deserves closer scrutiny. The work adds to a growing body of research on how model behavior can sway user opinions.

papersTODAY 04:00 UTC

BEACON: Behavior and Appearance Control for Subject-Specific Video Generation

A new arXiv paper introduces BEACON, a method for generating videos of a specific person that retains both their visual identity and their individual expressive behavior. The authors argue that beyond matching appearance, such models must also capture the facial mannerisms that distinguish how a given subject acts on camera. The work targets human-centric video synthesis where subject-specific fidelity is the main challenge.

papersTODAY 04:00 UTC

WAVIE Method Aims to Improve Deepfake Detection on Unseen Manipulations

A new arXiv paper introduces WAVIE, a lightweight approach to detecting deepfake faces that combines wavelet-based augmentation with intermediate vision embeddings. The authors target a common weakness in existing detectors, which often lose accuracy when they encounter manipulation techniques not seen during training. The work is positioned as a step toward more reliable face forgery detection in real-world settings.

papersTODAY 04:00 UTC

ZAPS Method Aims to Speed Up Neural Architecture Search with Zero-Cost Proxies

A new arXiv preprint introduces ZAPS, a technique for selecting zero-cost proxies that estimate how well a neural network architecture will perform before any training takes place. Because evaluating candidates normally requires full training runs, such proxies could make architecture search far cheaper. The paper focuses on how to identify which proxy to use for a given search task.

papersTODAY 04:00 UTC

DA-DLM Models Token Dependencies in Diffusion Language Models

A new arXiv paper introduces DA-DLM, a method for diffusion language models that explicitly captures relationships between tokens during generation. Existing diffusion models denoise masked text by predicting several tokens at once under an assumption of conditional independence, which the authors say weakens coherence. The proposed approach aims to restore those inter-token dependencies.

papersTODAY 04:00 UTC

AURA: Unified Multimodal Framework for Conversational Music Editing

Researchers present AURA, a multimodal framework that lets users edit music through a back-and-forth conversation rather than one-off commands. Existing instruction-guided editors handle each request in isolation, which makes it hard to iteratively refine a track. AURA is designed to keep track of prior edits so a session can build progressively toward a desired result.

papersTODAY 04:00 UTC

arXiv Paper Targets Reasoning-Critical Neurons to Steer LLM Inference

A new arXiv preprint proposes locating the specific neural components that matter most for reasoning tasks, then modifying model activations to steer outputs accordingly. The authors argue this approach can make inference on hard problems more dependable without extra post-training or costly sampling. The work is presented as a way to improve reliability and efficiency during deployment.

papersTODAY 04:00 UTC

PPDL framework combines physical priors with deep learning for user retention forecasting

Researchers present PPDL, a framework for predicting channel-level user retention ratios in multi-channel paid user acquisition. The method integrates physical priors with deep learning to capture the sharp early churn typical of retention curves. Accurate early forecasts are intended to help marketers allocate advertising budgets more effectively.

papersTODAY 04:00 UTC

Closed-Form Occlusal Geometry Proposed for Orthodontic Report Generation

This arXiv paper notes that intraoral scan datasets such as Bite2Text arrive already aligned in occlusion, which means key occlusal measurements can be calculated directly rather than inferred by a learned captioning model. The authors argue for deriving these quantities in closed form as a basis for automatically generating orthodontic reports. The work falls under computation and language research and has not been peer reviewed.

papersTODAY 04:00 UTC

Perceptual Reality Transformer Explores What Illustrations Must Preserve

A new arXiv paper introduces the Perceptual Reality Transformer, a model aimed at helping people convey atypical perceptual experiences while keeping their intended meaning intact. The work argues that recognizable imagery alone is insufficient, since such accounts also carry vividness, duration, uncertainty, and emotional tone. It examines what an illustration needs to retain so those qualities survive translation into a generated image.

papersTODAY 04:00 UTC

arXiv paper outlines neuromorphic design automation flow bridging neuroscience and EDA

A preprint proposes a framework for electronic Neuromorphic Design Automation, described as a pipeline connecting computational neuroscience models with established electronic design automation practices. The authors argue for a unified, end-to-end approach rather than treating the two domains separately. The work is a revised cross-listing on arXiv and presents a conceptual design flow rather than a released tool.

papersTODAY 04:00 UTC

arXiv paper argues foundation models should move toward open-ended discovery

A new arXiv preprint proposes "Discovery Foundation Models," framing open-ended discovery as the next stage for AI systems. The authors argue that models have moved from recalling and reasoning over existing knowledge to acting with tools and learning from outcomes, and that the next step is generating genuinely new findings. The paper is a position piece rather than an experimental release.

papersTODAY 04:00 UTC

MOSCOPT Method Optimizes Multiple LLM Agent Skills Together

A new arXiv paper introduces MOSCOPT, an approach that jointly optimizes collections of prompts and skills for LLM agents rather than refining a single text template. The authors argue that existing prompt and skill optimization methods miss beneficial interactions between multiple skills used by an agent. The work is a research preprint and has not yet been peer reviewed.

papersTODAY 04:00 UTC

Explainable Hybrid Feature Selection Proposed for Intrusion Detection in IoMT

A new arXiv paper describes an intrusion detection system designed for Internet of Medical Things networks, where devices are diverse and computing power is limited. The approach combines hybrid feature selection with explainability so that real-time traffic can be screened while keeping the model's decisions interpretable. The authors frame resource constraints and the need for timely analysis as the main obstacles the method targets.

papersTODAY 04:00 UTC

Rater Ising-Potts Model Derives Weights From LLM Embeddings

A new arXiv paper introduces a Rater Ising-Potts model, an extension of the Ising model designed for multinomial rating data. The approach builds pairwise agreement indicators and category labels into the model, drawing its weights from large language model embeddings. The authors position the work as a link between network psychometrics and AI methods.

papersTODAY 04:00 UTC

arXiv Paper Proposes BusMA, a Shared Bus Communication Layer for Multi-Agent AI Systems

A new arXiv preprint introduces BusMA, a communication substrate intended to coordinate multi-agent systems that handle planning, tool use, and evidence synthesis. The authors argue that current designs, which rely on hierarchical manager-worker structures or router-based message passing, have limitations that a bus-style architecture could address. The paper is a preprint and has not yet been peer reviewed.

papersTODAY 04:00 UTC

Hybrid Machine-Learning Model Adds Uncertainty Estimates to Irrigation Scheduling

A new arXiv paper presents a hybrid mathematical and machine-learning approach for irrigation decision support that also quantifies confidence in its soil-moisture forecasts. The authors note that irrigation is typically scheduled reactively, even though agriculture uses about 70% of global freshwater withdrawals. The model aims to give growers both a forward-looking moisture prediction and a measure of how reliable that prediction is.