LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

ai-interpretability

topic11 events
papersTODAY 04:00 UTC

arXiv paper invites mathematicians to develop new math for AI safety

A newly posted arXiv paper argues that current AI systems risk outpacing human understanding and control, and that fresh mathematical work is needed to make them legible, steerable, and cooperative. The author structures the call by mathematical subfield so researchers can identify where their expertise applies, and frames it as an open invitation to the mathematics community.

papersTODAY 04:00 UTC

Sparse Autoencoders Applied to Interpret Whisper Speech Encoder Internals

A new arXiv paper examines the internal representations of Whisper, an automatic speech recognition model, by applying sparse autoencoders to its encodings. The authors note that interpretability research has focused mostly on text-based transformers, leaving speech systems comparatively unstudied. Their work aims to make the features learned by Whisper's encoder more understandable.

papersTODAY 04:00 UTC

arXiv Paper: Empathy in LLMs Is Steerable but Acts Along Multiple Axes

A new arXiv preprint examines whether supportive empathy in large language models can be controlled through activation steering, as has been done for traits like honesty and refusal. Using the EPITOME dataset, the authors analyze the geometry of the underlying mechanisms and find that empathy does not map cleanly onto a single controllable direction, indicating a multi-axial structure. The work also looks at how persona settings influence these empathy-related representations.

papersTODAY 04:00 UTC

IMPACT-VLA attributes robot policy behavior using counterfactual trajectories

A new arXiv paper introduces IMPACT-VLA, a method for tracing how much each input modality — camera images, proprioceptive state, and language instructions — contributes to a vision-language-action policy's decisions at different points during task execution. The approach relies on counterfactual trajectories to isolate the effect of individual inputs, addressing the difficulty of interpreting these multimodal robot policies. The abstract excerpt does not detail experimental results or benchmarks.

papersTODAY 04:00 UTC

MACCHIATO training method targets certifiably interpretable ReLU-MLP Boolean models

A new arXiv paper introduces MACCHIATO, a specialized training algorithm for ReLU-based multilayer perceptrons that solve Boolean tasks. The method is designed to produce networks that are certifiably interpretable while guaranteeing correct generalization across the full truth table. The authors frame the work as a response to the growing gap between AI capability and the ability to explain model behavior.

papersSEP 12 13:56 UTC

Anthropic Paper Proposes Mathematical Framework for Analyzing Transformer Circuits

Anthropic researchers published a paper outlining a mathematical approach to reverse-engineering how transformer models compute internally, treating attention heads and MLP layers as composable circuits. The framework aims to make the internal mechanisms of these models more tractable to study and explain. It is intended as a foundation for interpretability work rather than a description of any specific deployed system.

papersSEP 12 04:00 UTC

Paper proposes semantic framework for judging AI system representations

A new arXiv paper argues that an AI system's output should not be read as a description of a fact or world state, but as an engineered representation. The authors propose a semantic framework for describing such systems so their representations can be checked for correctness. The work is a conceptual contribution rather than a model release or benchmark.

papersSEP 10 04:00 UTC

CT-SAFR framework proposes multi-layered verification for chain-of-thought reasoning in robots

Researchers present CT-SAFR, a multi-layered verification framework intended to make chain-of-thought reasoning by large language models safe and interpretable in autonomous robots. The framework responds to findings that reasoning models often do not verbalize their true decision processes, adding verification layers to support trustworthy AI-driven robotic decisions.

papersSEP 10 04:00 UTC

New arXiv paper proposes calibrating AI agent confidence from internal representations

A newly released arXiv paper addresses how to measure the confidence behind agentic AI actions, arguing this is essential as such systems enter safety-critical applications. The authors note that agentic workflows fail in more complex ways than traditional machine learning systems and propose deriving calibrated confidence estimates directly from the model's internal representations.

papersSEP 10 04:00 UTC

Position Paper Argues LLM Self-Explanations Must Move From Plausible to Actionable

A position paper examines how large language models generate natural-language accounts of their own decisions, a practice known as self-explanation. The authors argue that current explanations are often merely plausible-sounding rather than genuinely useful, and they outline what would be needed to make them actionable for real-world use.

papersSEP 9 11:14 UTC

GPT-6 Astra spurs research interest in looped transformers and hidden reasoning

A new analysis examines the ideas behind GPT-6 Astra, focusing on transformer architectures that reuse the same blocks across multiple passes rather than adding more layers. It reviews recent work on recurrent depth, where looping blocks can increase effective model depth and enable internal computation that is not exposed in the visible output. The piece also discusses hidden chains of thought and what this implies for interpreting model reasoning.