papersTODAY 04:00 UTC
arXiv Paper: Empathy in LLMs Is Steerable but Acts Along Multiple Axes
A new arXiv preprint examines whether supportive empathy in large language models can be controlled through activation steering, as has been done for traits like honesty and refusal. Using the EPITOME dataset, the authors analyze the geometry of the underlying mechanisms and find that empathy does not map cleanly onto a single controllable direction, indicating a multi-axial structure. The work also looks at how persona settings influence these empathy-related representations.