The True Motive Behind Watermarking: To Avoid AI-generated Text During Training
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Did no one notice? This solves in an elegant way the well known problem: if the internet will be full of AI slop how can the AI companies train their model without poisoning the data set and avoid the Ouroboros problem - AI eating its own generated text during the training? Well, detect the generated text and omit it from training. We know that Claude watermark texts but the fact other LLMs did not publish they do it, they are maybe, just maybe, doing it anyways. Still I think the problem is the quality will be worse for high quality texts because it essentially changes the NATURAL word frequency (so the result inevitably will be UNNATURAL).