papersSEP 11 04:00 UTC
arXiv paper examines in-context multimodal jailbreaks in AI models
A new arXiv preprint analyzes how harmful examples placed in a prompt can make multimodal large language models produce unsafe outputs without any change to their weights. The authors frame this in-context jailbreak behavior as a vulnerability and propose a posterior reweighting approach to explain it. The work falls under cross-listed machine learning submissions.