papersSEP 10 04:00 UTC
Reference-based method audits LLM bias via relative representations of hidden states
An arXiv paper in cs.AI introduces a technique for auditing bias in large language models by analyzing internal hidden states instead of relying on generated outputs. By comparing a model's representations against those of a reference model using relative representations, the approach aims to detect internal bias shifts that output-based benchmarks or judge models could miss. The authors frame it as a cheaper alternative to benchmark-heavy or judge-dependent auditing pipelines.