Watermarks for AI-generated text are easy to remove and could be stolen and copied, rendering them useless, researchers have found. They are saying these sorts of attacks discredit watermarks and may idiot people into trusting text they shouldn’t.
Watermarking works by inserting hidden patterns in AI-generated text, which permit computers to detect that the text comes from an AI system. They’re a reasonably latest invention, but they’ve already grow to be a preferred solution for fighting AI-generated misinformation and plagiarism. For instance, the European Union’s AI Act, which enters into force in May, would require developers to watermark AI-generated content. But the brand new research shows that the innovative of watermarking technology doesn’t live as much as regulators’ requirements, says Robin Staab, a PhD student at ETH Zürich, who was a part of the team that developed the attacks. The research is yet to be peer reviewed, but will likely be presented on the International Conference on Learning Representations conference in May.
AI language models work by predicting the subsequent likely word in a sentence, generating one word at a time on the premise of those predictions. Watermarking algorithms for text divide the language model’s vocabulary into words on a “green list” and a “red list,” after which make the AI model select words from the green list. The more words in a sentence which might be from the green list, the more likely it’s that the text was generated by a pc. Humans tend to write down sentences that include a more random mixture of words.
The researchers tampered with five different watermarks that work in this fashion. They were in a position to reverse-engineer the watermarks by utilizing an API to access the AI model with the watermark applied and prompting it repeatedly, says Staab. The responses allow the attacker to “steal” the watermark by constructing an approximate model of the watermarking rules. They do that by analyzing the AI outputs and comparing them with normal text.
Once they’ve an approximate idea of what the watermarked words may be, this enables the researchers to execute two sorts of attacks. The primary one, called a spoofing attack, allows malicious actors to make use of the data they learned from stealing the watermark to supply text that could be passed off as being watermarked. The second attack allows hackers to wash AI-generated text from its watermark, so the text could be passed off as human-written.
The team had a roughly 80% success rate in spoofing watermarks, and an 85% success rate in stripping AI-generated text of its watermark.
Researchers not affiliated with the ETH Zürich team, resembling Soheil Feizi, an associate professor and director of the Reliable AI Lab on the University of Maryland, have also found watermarks to be unreliable and vulnerable to spoofing attacks.
The findings from ETH Zürich confirm that these issues with watermarks persist and extend to essentially the most advanced varieties of chatbots and enormous language models getting used today, says Feizi.
The research “underscores the importance of exercising caution when deploying such detection mechanisms on a big scale,” he says.
Despite the findings, watermarks remain essentially the most promising method to detect AI-generated content, says Nikola Jovanović, a PhD student at ETH Zürich who worked on the research.
But more research is required to make watermarks ready for deployment on a big scale, he adds. Until then, we should always manage our expectations of how reliable and useful these tools are. “If it’s higher than nothing, it remains to be useful,” he says.