LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 Study traces LLM hallucinations to competing latent associations1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 Study traces LLM hallucinations to competing latent associations1 src1.3 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.3 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

activation steering

topic5 events
papersTODAY 04:00 UTC

arXiv paper compares off-the-shelf persona vectors with targeted steering for sycophancy

A new arXiv preprint examines sycophancy, the tendency of language models to agree with users even when the user is wrong. The authors build on earlier work that derived sycophancy persona vectors and used activation steering to control the behaviour, and they test whether generic, off-the-shelf persona vectors can match purpose-built steering methods. The results suggest the simpler approach performs competitively.

papersTODAY 04:00 UTC

arXiv Paper Studies Steerability Signatures in Language Model Activations

A new arXiv preprint investigates why steering language models with contrastive representation pairs works well for some behaviors but not others. The authors look for measurable signatures in activation space that indicate how steerable a given model behavior is. The work aims to make activation-based control more predictable rather than relying on trial and error.

papersTODAY 04:00 UTC

Study locates and steers opportunity-recognition behavior inside LLMs

A new arXiv paper examines how entrepreneurial cognition research can be extended to large language models, which are increasingly used in entrepreneurial tasks. The authors identify an internal representation tied to opportunity recognition and show they can causally steer it, effectively turning the behavior up or down. The work sits at the intersection of entrepreneurship theory and interpretability research on model internals.

papersTODAY 04:00 UTC

arXiv Paper Targets Reasoning-Critical Neurons to Steer LLM Inference

A new arXiv preprint proposes locating the specific neural components that matter most for reasoning tasks, then modifying model activations to steer outputs accordingly. The authors argue this approach can make inference on hard problems more dependable without extra post-training or costly sampling. The work is presented as a way to improve reliability and efficiency during deployment.

papersTODAY 04:00 UTC

arXiv Paper: Empathy in LLMs Is Steerable but Acts Along Multiple Axes

A new arXiv preprint examines whether supportive empathy in large language models can be controlled through activation steering, as has been done for traits like honesty and refusal. Using the EPITOME dataset, the authors analyze the geometry of the underlying mechanisms and find that empathy does not map cleanly onto a single controllable direction, indicating a multi-axial structure. The work also looks at how persona settings influence these empathy-related representations.