papersTODAY 04:00 UTC
Study tracks how harmful intent signals build across LLM layers
Researchers describe a phenomenon they call Harmfulness Propagation Dynamics, in which the last-token hidden state of a harmful prompt projects increasingly onto a learned harm direction as depth increases. Benign prompts did not show this pattern, instead staying flat or fluctuating across layers. The finding points to layer-wise differences that could inform how models are monitored or steered for safety.