papersSEP 12 04:00 UTC
Paper Identifies 'Perfect Aliasing' Failure in Compliant-Context Truth Probes
A new arXiv paper examines a problem it calls "perfect aliasing," in which a truthfulness probe trained on data where honest reporting and the task's prescribed action line up cannot tell those two targets apart from the labels alone. The authors argue this amounts to a failure of semantic identification, and they study it using a controlled binary reporting setup.