papersTODAY 04:00 UTC
Forked Futures Method Tests Reusable Causal Interfaces in Language Models
A new arXiv paper argues that probing a language model's current answer is not enough to show it has a stable, reusable internal interface, since the same output can come from hidden states that would support different later computations. The authors propose a "forked futures" approach, in which future operations are sampled only after the fact, to test whether internal representations serve as causal interfaces that transfer across tasks. The work targets interpretability and evaluation of model internals rather than a product release.