papersSEP 10 04:00 UTC
Paper audits and mitigates bias in protein-protein interaction datasets for ML
A new machine learning study argues that protein-protein interaction databases carry study and technical biases that skew protein and interaction attributes, letting models succeed by exploiting shortcuts rather than genuine biological signals. The authors propose methods to audit these datasets for such biases and to mitigate them, aiming for models that learn real biology instead of dataset artifacts.