papersSEP 11 04:00 UTC
Survey maps privacy and extraction attacks that abuse ML explanations
A systematization of knowledge paper reviews 25 studies showing how explainable AI outputs can be turned against the models they describe. The authors group these attacks into model extraction, membership inference and model inversion, and note that explanations widen the confidentiality and privacy risks of deployed systems. The work calls for treating explanation interfaces as part of the attack surface.