SeriousM
4 hours ago
So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet (from now on), yet people start complaining that "their" output is marked as spam.
Well, the solution is quite easy: start thinking on your own again and write the lines yourself.
I welcome this watermarking. Finally it's an easy detectable signal that someone just generated some request/answer to waste my time by forcing me thinking for the other person too. Now that time is over.
(It never was hard to spot this type of texts but now there is evidence)
sixtyj
4 hours ago
The reason why Anthropic did it is imho the different - they don’t want to ingest their own output again (or somehow differentiate from already stored data) so they try to sign it.
The data volume that is generated daily is staggering so every petabyte counts :)
dgellow
4 hours ago
> So we finally tackle the issue of ai-text pollution and probably found a way to clean up the internet
We are very, very far from that! So far we have a single ai vendor introducing a statistical bias to their generation that can make it simpler to identify genAI in some (longer) texts. We don’t know yet how effective that will be in practice, and if other models will follow suit
Eddy_Viscosity2
4 hours ago
If other AI vendors don't follow suit, then people will move over to them to escape detection. Then of course Claude will stop watermarking.
AureliusMA
3 hours ago
Watermarking is probably done by other vendors covertly as well. See my other comment.
dgellow
4 hours ago
Yep. Also, the vast majority of online content is very short (twitter, Reddit comments, HN comments, etc). It’s not something where SynthID type approach will be effective i assume. Still a good first baby step