papersSEP 12 04:00 UTC
Paper Examines How Prevalence Drives Precision in Detector-Built Datasets
A new arXiv paper argues that when datasets are built by running a detector, heuristic, or model over candidate pools, the resulting label precision depends on the true-positive rate within each pool rather than on detector quality alone. The authors apply Bayes' rule to show how this prevalence effect introduces hidden contamination into detector-defined datasets, a problem they describe as silent. The work suggests dataset builders should account for pool-level prevalence when estimating or reporting precision.