arXiv paper proposes predictive likelihood ratios for LLM watermark detection
A new arXiv preprint presents a method for detecting watermarks in language model output by constructing predictive likelihood ratios. The approach builds on prior work by Li et al. (2025) and averages over uncertainty in the detection test. Watermark detection works by testing whether observed tokens depend on pseudorandom values derived from a secret key.