papersSEP 12 13:27 UTC
Study links reasoning models' internal states to distinct thought steps
A new study finds that operations such as arithmetic, recalling formulas, and logical deduction show up as separate patterns inside reasoning models, most visibly in their middle layers. This suggests models carry out more processing than their published chain-of-thought text discloses, which researchers flag as relevant to AI safety and oversight. The findings could inform how developers monitor or audit model reasoning.