LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#llm

40 curated events
papersTODAY 04:00 UTC

Study locates and steers opportunity-recognition behavior inside LLMs

A new arXiv paper examines how entrepreneurial cognition research can be extended to large language models, which are increasingly used in entrepreneurial tasks. The authors identify an internal representation tied to opportunity recognition and show they can causally steer it, effectively turning the behavior up or down. The work sits at the intersection of entrepreneurship theory and interpretability research on model internals.

papersTODAY 04:00 UTC

Study Separates Inference Topology From Diversity in Multi-Agent LLM Emotion Detection

A new arXiv paper examines multi-agent LLM pipelines by treating two design choices as independent variables: how agent calls are wired together and where the differences between agents come from. The authors evaluate this on multilingual, low-resource emotion detection, where labeled data is scarce. The goal is to clarify which gains come from the structure of the agent network versus from the diversity introduced between agents.

papersTODAY 04:00 UTC

arXiv paper analyzes how AI-generated data affects dataset decomposition

A new preprint examines what happens to batch decomposition and downstream model performance when training sets mix human data with text or images produced by existing large language models. Using random datasets containing anomalies, the authors study criticality in dissimilar decomposition and undersampling techniques. The work aims to clarify the statistical behavior of datasets that are increasingly populated with synthetic samples.

papersTODAY 04:00 UTC

arXiv paper studies how evidence ordering shapes streaming test-time compute

A new arXiv paper argues that reasoning policies for a fixed task and compute budget should change depending on the order in which evidence arrives. It frames this as an "information-slack dilemma": computing early leaves more time to finish but relies on incomplete evidence, while waiting yields better information at the cost of fewer remaining steps. The authors contend that depth and width alone do not capture this tradeoff, which streaming settings must account for.

papersTODAY 04:00 UTC

Paper Proposes Framework for Judging When Synthetic Survey Data Is Trustworthy

A new arXiv paper argues that the debate over synthetic data in marketing research has been stuck between two extremes: treating large language models as a replacement for human survey respondents, or rejecting them outright. The authors say the more useful question is when synthetic respondents can be trusted, and they outline how that reliability should be evaluated. The work focuses on marketing research but touches on broader issues of validating model-generated data.

papersTODAY 04:00 UTC

LLM-Assisted Multi-Agent RL Framework Coordinates EV Charging, Stations and Grid

A new arXiv paper proposes combining large language models with multi-agent reinforcement learning to jointly optimize electric vehicle charging scheduling in public charging systems. The approach targets three competing goals at once: driver charging satisfaction, charging station profitability, and stability of the smart grid. It is positioned as a unified optimization method for connected EV infrastructure in IoT settings.

papersTODAY 04:00 UTC

DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern

Researchers have introduced DiffuTester, a technique to accelerate unit test generation using diffusion large language models. The method extracts structural patterns from code to guide parallel generation, aiming to make automated testing faster and more scalable. The work addresses the need for efficient large-scale software testing.

papersTODAY 04:00 UTC

Paper proposes evolving context parameterization for large language models

A new arXiv paper addresses a limitation in context parameterization, a technique that lets language models absorb context into reusable parameters instead of reprocessing it for every query. The authors note that current approaches treat context as static and are therefore ill-suited to settings where information changes over time. Their work introduces a method for keeping those internalized parameters up to date as contexts evolve.

papersTODAY 04:00 UTC

arXiv Paper Examines Capacity Limits of Reasoning via Superposition

A new arXiv preprint studies how much intermediate computation a single vector can carry when language models reason through continuous or recurrent methods rather than token-by-token chain-of-thought. The work frames multi-step reasoning as superposition, where partial computations are packed into hidden states, and analyzes the resulting capacity limits. It offers a theoretical lens on the trade-offs between explicit token-based reasoning and continuous latent approaches.

papersTODAY 04:00 UTC

arXiv study explores using LLMs to simplify medical information for diabetes patients

A new arXiv paper examines how large language models can be used to make complex medical information easier for patients to understand, using diabetes as a case study. The authors argue that clearer simplification supports patient comprehension, informed decision-making, and better health outcomes. The work focuses on the challenges of translating clinical knowledge into patient-friendly language.

papersTODAY 04:00 UTC

Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election

A new preprint looks at how generative AI writing tools and the language models behind them may shape the political information voters encounter ahead of Sweden's 2026 election. The authors argue that as these assistants become a common way to gather information, their tendency to favor certain issues or viewpoints deserves closer scrutiny. The work adds to a growing body of research on how model behavior can sway user opinions.

papersTODAY 04:00 UTC

MOSCOPT Method Optimizes Multiple LLM Agent Skills Together

A new arXiv paper introduces MOSCOPT, an approach that jointly optimizes collections of prompts and skills for LLM agents rather than refining a single text template. The authors argue that existing prompt and skill optimization methods miss beneficial interactions between multiple skills used by an agent. The work is a research preprint and has not yet been peer reviewed.

papersTODAY 04:00 UTC

SlopShape: Identifying AI-Generated Commercial Web Content

A new arXiv paper introduces SlopShape, a method aimed at spotting AI-written commercial web content such as marketing or product pages. The authors note that word-level AI-text detectors perform well on unedited output but break down when text is reworded, and that a single word-level score says little about a document's character or which model produced it. The work explores whether document-level analysis can better characterize automatically generated commercial text.

papersTODAY 04:00 UTC

Study traces how large language models represent animacy

A arXiv paper examines where the concept of animacy is encoded inside large language models. The authors trace internal circuits tied to the animate/inanimate distinction, which involves verb-argument constraints and contextual cues beyond simple word-level features. The work is a revised version of a preprint in the cs.CL category.

papersTODAY 04:00 UTC

LLM-based split learning predicts mental distress across heterogeneous surveys

A new arXiv paper proposes a schema-aware split learning approach that uses LLMs to predict mental distress from survey data while keeping sensitive records private. The method is designed to work across surveys with differing structures and questions, which is a common obstacle when pooling mental health data from schools, employers, and clinics. The work targets privacy-preserving collaboration, so data stays local rather than being centralized.

papersTODAY 04:00 UTC

Learning to Coach: Training an LLM to Distill Guidance From Experience

A new arXiv paper introduces Learning to Coach (L2C), a framework that trains a separate LLM acting as a coach to pull actionable guidance out of experience. The motivation is that raw solution trajectories are typically long and noisy, which limits how well language models can learn from them. The approach aims to convert such trajectories into more useful, condensed coaching signals.

papersTODAY 04:00 UTC

arXiv paper examines midtraining stage as a way to control how LLM traits generalize

A new arXiv preprint explores whether midtraining, a stage between pretraining and post-training, can influence which behaviors a large language model carries forward. The authors propose a method called Inoculation Midtraining, which uses invented words to shape how desirable and undesirable properties generalize. The work is a research contribution and has not been peer reviewed.

papersTODAY 04:00 UTC

CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems

A new arXiv preprint introduces CoMem, a memory framework for LLM-driven multi-agent systems that combines collective knowledge with agent-specific memory. The authors argue that most existing approaches rely on flat, unstructured memory, which limits how agents learn and improve over time. The work targets better long-term cooperation and performance in evolutionary multi-agent setups.

papersTODAY 04:00 UTC

ECAS: Edge-Controlled Agentic System for Validation-Gated Scientific Execution

A new arXiv paper introduces ECAS, an agentic system that uses large language models to turn a scientist's high-level goal into correct, target-scale runs on high-performance computing resources. The approach places control at the edge and gates execution behind validation checks, aiming to address how brittle and labor-intensive it currently is to translate research intent into working HPC workflows. The authors position the work as a step toward more reliable LLM-driven scientific computing.

papersTODAY 04:00 UTC

Paper Probes LLM Benchmark Success Using Token-Level Perplexity

A new arXiv paper argues that standard task-performance evaluations of large language models reveal little about whether correct answers stem from the mechanisms researchers assume, which can encourage confirmation bias. The authors propose a simple, principled method that uses token-level perplexity to contrast how models behave on benchmarks with how they distribute probability internally. The work is a replacement submission to arXiv's computation and language section.

papersTODAY 04:00 UTC

Study examines how users iterate prompts to explore narrative space in LLM story generation

A new arXiv paper analyzes public chatbot logs to understand how people write stories with large language models. The authors find that users repeatedly revise their prompts, tweaking characters and other story elements to explore different narrative directions. The work frames this behavior as navigation through a space of possible stories.

papersTODAY 04:00 UTC

arXiv paper uses persona-based distillation to improve LLM humor generation

A new arXiv preprint argues that standard next-token training objectives work against the surprise and incongruity that comedy requires, making humor a hard task for large language models. The authors propose HumorGen, a method that distills knowledge from multiple personas to create a cognitive synergy effect. The work is a revised submission and focuses on generation quality rather than a released product.

papersTODAY 04:00 UTC

arXiv Paper Proposes Self-Orchestrating LLMs to Cut Inference Latency

A new arXiv preprint introduces a method for having language models coordinate their own computation by exploiting semantic dependencies between generated tokens. The authors argue that standard autoregressive decoding is slow and leaves GPUs underused when batch sizes are small, and that their approach improves inference efficiency. The work is currently a research preprint and has not been peer reviewed or released as a product.

papersTODAY 04:00 UTC

CALICO System Aligns LLM Annotation Prompts With Expert Codebooks

Researchers present CALICO, a human-centered system that helps domain experts turn their annotation codebooks into prompts for large language models. The work targets gaps in existing pipelines, which offer little support for producing prompts that stay reliable, easy to revise, and auditable. It is described in a paper posted to arXiv under the cs.CL category.

papersTODAY 04:00 UTC

Study Tests Limits of Linear Truth Directions in LLM Activations

A new arXiv paper investigates the linear directions in a large language model's activation space that prior work associates with statement truth. It questions how universal or generalizable these truth directions are, building on earlier claims about their consistency across contexts. The work falls within ongoing research on interpreting and steering model internals.

papersTODAY 04:00 UTC

Study proposes vulnerability modeling and execution-based benchmark for secure code generation

A new arXiv paper addresses the gap between code that runs correctly and code that is secure when generated by large language models. The authors argue that progress has been limited by existing benchmarks that are small and not executable, making security flaws hard to measure reliably. Their approach combines task-adaptive modeling of vulnerabilities with an execution-based benchmark intended to evaluate both functional correctness and security.

papersTODAY 04:00 UTC

CoTAL: Human-in-the-Loop Prompt Engineering for Formative Assessment Scoring

Researchers present CoTAL, a human-in-the-loop prompt engineering method for using large language models to score formative assessments and generate feedback for students. The work examines how well such prompting approaches generalize across educational contexts, with teachers involved in refining the prompts. It is published as an arXiv preprint in the computation and language category.

papersTODAY 04:00 UTC

Position paper argues anthropomorphism hinders LLM research

A new arXiv position paper contends that attributing human-like traits to AI systems is an automatic habit that persists even among technical experts, and that it skews how researchers frame and evaluate language models. The authors review a large body of published work to show how anthropomorphic language shapes experimental design, interpretation of results, and safety claims. They call for alternative conceptual frameworks that describe model behavior without implying human-like minds.

papersTODAY 04:00 UTC

MUSE: Theory-Guided Story Engine for LLM Narrative Generation

A new arXiv paper in computational linguistics introduces MUSE, a story-generation engine that draws on narrative theory to steer how plot, character, and language choices fit together across planning, drafting, and revision. The authors frame story guidance as facing two bottlenecks, pointing to the quality of the guidance and how it is supplied. The abstract as posted is truncated, so full details of the method and results were not available.

papersTODAY 04:00 UTC

Study Tests Whether LLMs Can Simulate Individual Financial Decisions

A preliminary arXiv paper examines whether large language models can stand in for people as user simulators in financial settings. The authors ran a controlled paper-trading experiment with 120 volunteers to see how well model behavior tracks evolving individual investing choices. They conclude that current LLM simulation of such decisions is not yet reliable and call for further work.

papersTODAY 04:00 UTC

Lightning Weave: Capability Composition for More Efficient Reasoning Models

A new arXiv paper introduces Lightning Weave, a method aimed at pushing the accuracy-efficiency frontier of reasoning models. The authors argue that accuracy and inference efficiency often pull toward different reasoning behaviors, making joint improvement difficult. Their approach relies on composing capabilities rather than optimizing the two objectives independently.

papersTODAY 04:00 UTC

arXiv paper examines merging LLM knowledge into automatic speech recognition

A new arXiv preprint in the cs.CL category describes work on combining large language models with automatic speech recognition systems. The paper focuses on knowledge-merging techniques related to established LM fusion approaches such as shallow fusion and density ratio methods. It appears to be a research contribution rather than a product or model release.

papersTODAY 04:00 UTC

New Paper Proposes Detecting LLM Hallucinations via Feed-Forward Neurons

A preprint introduces NeuroActiSep, a method that aims to spot factual hallucinations in large language models by inspecting feed-forward neurons, according to its abstract. The approach is described as working in a single pass, making it cheaper than methods requiring repeated sampling. The paper frames this as an under-explored alternative to existing white-box truthfulness detection techniques.

papersTODAY 04:00 UTC

Func-R1: Method Aims to Improve Mathematical Function Reasoning in Multimodal LLMs

A new arXiv paper introduces Func-R1, an approach aimed at strengthening mathematical function reasoning in multimodal large language models. The work targets the challenge of combining visual perception with symbolic logic when solving math problems from images. The abstract frames deliberate mathematical reasoning in visual settings as an indicator of advanced multimodal model capability.

papersTODAY 04:00 UTC

Study examines hindsight bias in clinical LLM temporal reasoning

A new arXiv paper argues that clinical language models are frequently assessed on retrospective patient records that already contain the eventual diagnosis, treatment response and outcome. Because those records expose information a real prospective decision-maker would not have, such benchmarks may reward models for exploiting future data instead of genuine reasoning. The work examines how this exposure shapes model judgments in clinical temporal tasks.

papersTODAY 04:00 UTC

Self-Demonstrations Improve LLM Schema-Ontology Mapping, arXiv Paper Finds

A new arXiv preprint examines how self-demonstrations, where a model generates its own worked examples before answering, affect the task of mapping relational database schemas onto a shared ontology. The authors frame the problem around semantic heterogeneity, obscure table and column names, missing metadata, and the abstraction gap that complicates enterprise knowledge integration. Results reported in the abstract suggest the self-demonstration approach is unexpectedly effective for this mapping task.

papersTODAY 04:00 UTC

THEMIS Workflow Records Requirement-to-Repair Traces for LLM Program Repair

A new arXiv paper introduces THEMIS, a staged workflow for repairing code in repositories using large language models. The system publishes observable traces that document how issue requirements are turned into edits, along with evidence collected after those edits. The authors argue such inspectable records are needed alongside correct patches for repository-level repair.

papersTODAY 04:00 UTC

TyPatch turns historical Linux kernel patches into typestate rules for bug detection

A new arXiv paper introduces TyPatch, a method that converts historical Linux kernel patches into typestate rules usable by static analysis. Building on prior work where large language models generate checkers from past patches, the approach aims to extend the reach of defect knowledge beyond the sites where bugs were originally fixed. The technique targets automated kernel bug detection.

papersTODAY 04:00 UTC

T-SMART: Mechanism-Level Attribution for Tool-Augmented Time-Series QA

A new arXiv paper introduces T-SMART, a method that attributes how tool-augmented language models arrive at answers for time-series question answering. The work targets a known weakness: LLMs handle poorly the case where numerical signals are turned into text and require explicit computation. It offers analysis at the level of internal mechanisms rather than just final outputs, though the abstract provided is truncated.