Study Finds AI Text Watermarks Fail Legal Evidence Standards
Study finds AI watermarking schemes KGW, Unigram, and SynthID fail Daubert legal criteria after simple paraphrase attacks strip out marks.
A new empirical study tests three widely used LLM watermarking schemes—KGW, Unigram, and SynthID-Text—against the legal admissibility bar set by the Daubert standard and the forensic rigor required by NIST SP 800-86. The researchers built a 60-point Forensic Readiness Score framework with mandatory gates to structure the evaluation, focusing on meaning-preserving paraphrase as the most realistic and hardest-to-dismiss attack vector.
The results undercut regulatory assumptions embedded in laws like the EU AI Act and California's SB 942, both of which require watermarks to be robust and hard to remove. Across 846 paraphrase runs, every initially detected KGW and Unigram watermark disappeared after paraphrasing, and SynthID lost its mark in 98.3% of cases. Even without any attack, false-negative rates were already high across all three methods. SynthID additionally misidentified human-written text as AI-generated in over 5% of cases and produced ambiguous 'paradox' results for nearly a fifth of its own outputs.
None of the three schemes satisfied more than two of the five Daubert factors, meaning current watermarking implementations fall short of the evidentiary threshold courts require. For engineers building compliance or provenance tooling, the study is a clear signal that watermark-based detection is not yet a dependable substitute for stronger content-authentication approaches.