papersTODAY 04:00 UTC
Paper Proposes Using Model Internals to Predict Behavior on Unseen Data
A new arXiv paper reframes interpretability research around predicting how a model will respond to previously unseen inputs, rather than only to targeted mechanistic interventions. The authors use a model's internal representations to forecast its out-of-distribution behavior. The work appears in two arXiv listings, cs.AI and cs.LG, as a replacement submission.