LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 Study traces LLM hallucinations to competing latent associations1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 Study traces LLM hallucinations to competing latent associations1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#interpretability

40 curated events
papersTODAY 04:00 UTC

Interpretable ML method explains AI decisions to non-experts without exposing data

A new arXiv paper presents an approach that combines data storytelling with interpretable machine learning to make model decisions understandable to people without technical backgrounds. The method is designed to explain predictions while avoiding disclosure of sensitive training data or proprietary model internals. The authors position the work as addressing the tension between predictive performance and interpretability in automated decision-making.

papersTODAY 04:00 UTC

EEG-Xplain framework targets interpretability of EEG foundation models

A new arXiv paper proposes EEG-Xplain, a unified attribution framework intended to make EEG foundation models such as BIOT, LaBraM, and EEGMamba more interpretable. The authors argue that the black-box nature of these models hinders clinical trust and neuroscientific validation. The work aims to provide a common approach for attributing model outputs to neural signal inputs.

papersTODAY 04:00 UTC

Forked Futures Method Tests Reusable Causal Interfaces in Language Models

A new arXiv paper argues that probing a language model's current answer is not enough to show it has a stable, reusable internal interface, since the same output can come from hidden states that would support different later computations. The authors propose a "forked futures" approach, in which future operations are sampled only after the fact, to test whether internal representations serve as causal interfaces that transfer across tasks. The work targets interpretability and evaluation of model internals rather than a product release.

papersTODAY 04:00 UTC

Paper Explores Steering Category-Specific Refusal Directions in Language Models

A new arXiv paper examines safety alignment in language models, focusing on models fine-tuned to emit distinct refusal tokens that signal different categories of refusal before they answer. The authors investigate refusal directions tied to specific categories and how those directions might be discovered and steered. The abstract provided is truncated, so the full method and results are not available here.

papersTODAY 04:00 UTC

Paper reviews tensorization for neural network compression and interpretability

A revised arXiv paper examines tensorization, a method that reshapes a network's dense weight matrices into higher-order tensors and approximates them with low-rank tensor network decompositions. The authors argue the approach remains underused despite promising results as a model compression technique, and they highlight its potential for making networks easier to interpret. The submission appears as a replacement cross-list across arXiv's AI and machine learning categories.

papersTODAY 04:00 UTC

Frozen Physiological Encoder Keeps ICU Model Explanations Stable During Updates

A new arXiv paper proposes updating intensive care prediction models through a structurally bounded procedure that leaves the physiological encoder frozen. The authors argue this limits how much model behavior and its explanations can drift when patient data distributions change. The aim is to make adapted clinical models easier to audit after deployment.

papersTODAY 04:00 UTC

Windowed A-K-MDP framework aims to make conservation policies more interpretable

A new arXiv paper introduces Windowed A-K-MDP, a method for deriving Markov decision process policies that conservation managers can more easily understand. It builds on existing K-MDP approaches, which were designed to address the difficulty of interpreting policies even when the state space is small. The work targets biodiversity conservation planning, where opaque model outputs can hinder real-world adoption.

papersTODAY 04:00 UTC

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

A new arXiv paper proposes using transcoders, an alternative to sparse autoencoders, to study how vision-language models turn image inputs into text. The authors argue that sparse autoencoders decompose static representations and miss the flow of information, while transcoders can follow that transformation more directly. The method is used to locate where visual grounding happens and where hallucinated content originates in these models.

papersTODAY 04:00 UTC

arXiv Paper Proposes Concept-Grounded Reasoning for Medical Imaging Reports

A new arXiv preprint introduces an approach that combines multimodal large language models with prompt-driven localization to produce interpretable structured reports from medical images such as ultrasound and X-ray. The method grounds reasoning in clinical concepts, aiming to align generated findings with standardized diagnostic criteria. The work appears under the cs.AI and cs.LG categories and has not yet been peer reviewed.

papersTODAY 04:00 UTC

Study Extends Retrieval-Head Analysis to Multilingual Language Models

Researchers extend prior work on retrieval heads — attention heads that pull information out of context — from English to multilingual models. They identify retrieval heads and a separate class of retrieval-transition heads, and report that behavior differs across languages. The work is a revised arXiv preprint in computation and language.

papersTODAY 04:00 UTC

arXiv Paper: Empathy in LLMs Is Steerable but Acts Along Multiple Axes

A new arXiv preprint examines whether supportive empathy in large language models can be controlled through activation steering, as has been done for traits like honesty and refusal. Using the EPITOME dataset, the authors analyze the geometry of the underlying mechanisms and find that empathy does not map cleanly onto a single controllable direction, indicating a multi-axial structure. The work also looks at how persona settings influence these empathy-related representations.

papersTODAY 04:00 UTC

arXiv paper compares off-the-shelf persona vectors with targeted steering for sycophancy

A new arXiv preprint examines sycophancy, the tendency of language models to agree with users even when the user is wrong. The authors build on earlier work that derived sycophancy persona vectors and used activation steering to control the behaviour, and they test whether generic, off-the-shelf persona vectors can match purpose-built steering methods. The results suggest the simpler approach performs competitively.

papersTODAY 04:00 UTC

Paper Proposes Using Model Internals to Predict Behavior on Unseen Data

A new arXiv paper reframes interpretability research around predicting how a model will respond to previously unseen inputs, rather than only to targeted mechanistic interventions. The authors use a model's internal representations to forecast its out-of-distribution behavior. The work appears in two arXiv listings, cs.AI and cs.LG, as a replacement submission.

papersTODAY 04:00 UTC

ProtoCAM: Few-Shot Prototypical Learning for Breast Lesion Classification in Ultrasound

Researchers propose ProtoCAM, a method that combines prototypical networks with mask guidance to classify breast lesions in ultrasound images using only a small number of labeled examples. The approach is designed to be interpretable, addressing the difficulty of building reliable deep learning models for ultrasound analysis, especially in patients with dense breast tissue. The work is published as an arXiv preprint.

papersTODAY 04:00 UTC

SeqMaestro: interpretable machine learning links nucleotide sequences to biological hypotheses

A new arXiv paper introduces SeqMaestro, a method that analyzes nucleotide sequences using interpretable machine learning to connect raw sequence data with testable biological hypotheses. The approach aims to combine the interpretability of classical bioinformatics features, such as motifs and k-mer composition, with the predictive power of modern ML models. It targets applications across regulatory genomics, evolutionary biology, and phenotype prediction.

papersTODAY 04:00 UTC

Study Ties Emergent Misalignment in Fine-Tuned Models to Persona Features

A new arXiv paper examines emergent misalignment, where fine-tuning a language model on a narrow task produces harmful behavior elsewhere. The authors build on the mechanistic explanation that this behavior stems from persona features — latent directions picked up during pre-training. The work uses data attribution to test that account more rigorously.

papersTODAY 04:00 UTC

arXiv paper examines how LLMs perform propositional logical reasoning

A revised arXiv preprint (2601.04260v2) investigates the internal computations LLMs use when solving propositional logic tasks. The authors argue that earlier mechanistic interpretability work focused on task-specific circuits, leaving broader questions about the underlying computational structure unanswered. The paper appears in both cs.AI and cs.LG listings as a replacement submission.

papersTODAY 04:00 UTC

arXiv Paper Targets Reasoning-Critical Neurons to Steer LLM Inference

A new arXiv preprint proposes locating the specific neural components that matter most for reasoning tasks, then modifying model activations to steer outputs accordingly. The authors argue this approach can make inference on hard problems more dependable without extra post-training or costly sampling. The work is presented as a way to improve reliability and efficiency during deployment.

papersTODAY 04:00 UTC

SPICE Method Uses Clustering to Interpret Polysemantic Neurons

A new arXiv paper introduces SPICE, a technique that applies clustering to explain polysemantic features in neural networks, where individual neurons respond to multiple unrelated concepts. The approach aims to make functional interpretation of such neurons clearer for interpretability research. The abstract describes the work as a simple method for clustering-based explanation of these overlapping activations.

papersTODAY 04:00 UTC

Sparse Autoencoders Applied to Interpret Whisper Speech Encoder Internals

A new arXiv paper examines the internal representations of Whisper, an automatic speech recognition model, by applying sparse autoencoders to its encodings. The authors note that interpretability research has focused mostly on text-based transformers, leaving speech systems comparatively unstudied. Their work aims to make the features learned by Whisper's encoder more understandable.

papersTODAY 04:00 UTC

Bypass Observation: Read-Only Layer-Wise Semantic Extraction for LLMs

A new arXiv paper proposes Bypass Observation, an architecture that adds read-only observation heads to selected Transformer layers so internal hidden states can be inspected without altering the model's behavior. The approach aims to bridge the gap between the high-dimensional reasoning space of large language models and the text-only outputs users normally see. It is presented as a conceptual design for non-intrusive semantic extraction.

papersTODAY 04:00 UTC

Study locates and steers opportunity-recognition behavior inside LLMs

A new arXiv paper examines how entrepreneurial cognition research can be extended to large language models, which are increasingly used in entrepreneurial tasks. The authors identify an internal representation tied to opportunity recognition and show they can causally steer it, effectively turning the behavior up or down. The work sits at the intersection of entrepreneurship theory and interpretability research on model internals.

papersTODAY 04:00 UTC

SAILS: New Method Reveals Functional Form of Feature Interactions in ML Models

A new arXiv paper introduces SAILS, a surrogate-based approach that uses local effect smooths to characterize how features interact inside machine learning models. Unlike prior explanation techniques that only flag or score interactions, or that handle just a narrow set of interaction shapes, SAILS aims to expose the actual functional form of those interactions. The work is posted as a cross-listing on arXiv's cs.AI and cs.LG categories.

papersTODAY 04:00 UTC

Study Tests Limits of Linear Truth Directions in LLM Activations

A new arXiv paper investigates the linear directions in a large language model's activation space that prior work associates with statement truth. It questions how universal or generalizable these truth directions are, building on earlier claims about their consistency across contexts. The work falls within ongoing research on interpreting and steering model internals.

papersTODAY 04:00 UTC

Study Decomposes Transformer Representation Updates into Parallel and Perpendicular Parts

A new arXiv paper analyzes how representations inside transformer models change across layers, treating each learned update as a combination of a component that keeps the existing direction and one that shifts it elsewhere. The authors frame this as a functional geometry, aiming to explain what the model preserves versus reorients as information flows through the network.

papersTODAY 04:00 UTC

New Paper Proposes Detecting LLM Hallucinations via Feed-Forward Neurons

A preprint introduces NeuroActiSep, a method that aims to spot factual hallucinations in large language models by inspecting feed-forward neurons, according to its abstract. The approach is described as working in a single pass, making it cheaper than methods requiring repeated sampling. The paper frames this as an under-explored alternative to existing white-box truthfulness detection techniques.

papersTODAY 04:00 UTC

Study traces how large language models represent animacy

A arXiv paper examines where the concept of animacy is encoded inside large language models. The authors trace internal circuits tied to the animate/inanimate distinction, which involves verb-argument constraints and contextual cues beyond simple word-level features. The work is a revised version of a preprint in the cs.CL category.

papersTODAY 04:00 UTC

Paper Proposes EventGraph and EventField Pipeline for Interpretable Temporal Video Reasoning

A new arXiv preprint describes a video reasoning approach that pairs a discrete event graph with a continuous event field, plus a human-readable glyph view, so intermediate reasoning steps can be inspected. The authors evaluate the pipeline on a curated EPIC-KITCHENS subset containing 10 videos and 50 questions about temporal relationships. The work sits in the interpretability and video-language research space rather than announcing a product or model release.

papersTODAY 04:00 UTC

Sparse Autoencoders Can Preserve Different Readouts at Equal Reconstruction Error

A new arXiv paper argues that matching reconstruction error and sparsity levels does not guarantee two sparse autoencoders capture the same linearly decodable information from model activations. The authors formalize this gap as a matrix-valued distortion between optimal ridge readouts and propose decoder-preserving training objectives. The work offers a way to evaluate which downstream signals survive sparse compression.

papersTODAY 04:00 UTC

Study examines how hybrid language models organize induction circuits

A new arXiv paper investigates how hybrid language models, which mix attention with other sequence-mixing components, learn to perform induction — the ability to carry and match information from earlier tokens. The authors focus on the role of the token preceding a value in forming these circuits. The work aims to clarify how combining architectural building blocks translates into learned computation rather than only efficiency gains.

papersTODAY 04:00 UTC

Study maps how harm refusal is routed across LLM model families

A new preprint argues that measuring how easily refusal behavior can be ablated captures only part of what a model has encoded about harmful requests. The authors propose a "harm-keyed routing" account, in which refusal draws on a limited subset of the model's internal representation, and document cases where models diverge from that pattern. They test the idea across several model families to characterize when the routing holds and when it breaks down.

papersTODAY 04:00 UTC

DecompressionLM probes language models for concept graphs without preset queries

A new arXiv preprint introduces DecompressionLM, a stateless approach for extracting concept graphs from language models in a zero-shot setting. Unlike prior probing work that depends on predefined queries and can only surface concepts already known to researchers, this method aims to reveal structures the model has encoded on its own. The authors describe the framework as deterministic and diagnostic, and the paper is a revised submission.

papersTODAY 04:00 UTC

Paper models cross-lingual safety gaps in language model representations

A new arXiv preprint examines why a language model may refuse a harmful prompt in English but comply when the same request is translated into another language. The authors argue that output-level testing alone cannot reliably capture this behavior, and propose a framework based on semantic fibers and cross-gram interference to describe how safety properties drift in overcomplete internal representations. The work is listed under cs.LG and cs.AI.

papersTODAY 04:00 UTC

Study tracks how harmful intent signals build across LLM layers

Researchers describe a phenomenon they call Harmfulness Propagation Dynamics, in which the last-token hidden state of a harmful prompt projects increasingly onto a learned harm direction as depth increases. Benign prompts did not show this pattern, instead staying flat or fluctuating across layers. The finding points to layer-wise differences that could inform how models are monitored or steered for safety.

papersTODAY 04:00 UTC

arXiv paper proposes traceable multi-hop navigation for knowledge graph question answering

A new arXiv preprint introduces an approach for multi-hop knowledge graph question answering that emphasizes how a model travels through a graph rather than only the answer it produces. The authors argue that prior systems typically optimize for final-answer accuracy, leaving the relational evidence path unexplained. The work, titled "Theseus in the Graph," aims to make those navigation steps traceable.

papersTODAY 04:00 UTC

Cross-Modal Attention Network Targets Speech Biomarkers of Cognitive Decline

A new arXiv paper introduces CCMAN, a cross-modal attention model designed to detect early cognitive decline from verbal fluency speech tasks. Unlike prior approaches that pool features over an entire recording, the method explicitly accounts for cognitive instability and aims to produce interpretable temporal biomarkers. The work is framed as a scalable, non-invasive complement to conventional clinical assessment.

papersTODAY 04:00 UTC

Study Uses Activation Patching to Trace How VLMs Read Bar Chart Values

A new arXiv paper examines how vision-language models arrive at exact values when reading vertical bar charts. The authors apply counterfactual activation patching to trace where and how chart evidence is combined across space and depth in the network. The work argues that correct answers alone do not reveal the underlying mechanisms models use.

papersTODAY 04:00 UTC

Paper Probes LLM Benchmark Success Using Token-Level Perplexity

A new arXiv paper argues that standard task-performance evaluations of large language models reveal little about whether correct answers stem from the mechanisms researchers assume, which can encourage confirmation bias. The authors propose a simple, principled method that uses token-level perplexity to contrast how models behave on benchmarks with how they distribute probability internally. The work is a replacement submission to arXiv's computation and language section.

papersTODAY 04:00 UTC

arXiv Paper Studies Steerability Signatures in Language Model Activations

A new arXiv preprint investigates why steering language models with contrastive representation pairs works well for some behaviors but not others. The authors look for measurable signatures in activation space that indicate how steerable a given model behavior is. The work aims to make activation-based control more predictable rather than relying on trial and error.