arXiv Paper Argues Fairness Benchmarks Like BBQ Are Too Easy to Pass
A new arXiv preprint examines how fairness benchmarks such as BBQ are used to evaluate aligned language models and argues that a single example can be sufficient to pass them. The author contends this makes current evaluation methods unreliable for judging how fair a model actually is, and calls for rethinking how fairness is measured. The paper notes it uses stereotyped and offensive examples only for illustration.