Actually the evidence is not damning at all.
The evidence is a combination of two things: Tristan insinuating they did (note, not claiming but suggesting or implying), and their own statement saying "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
The latter is being used by everyone as a sign of guilt, but it's clear that they are making a legally-safe statement since promising a forensic level guarantee that nothing Tristan has every typed into ChatGPT has ever made its way into any training data is a massive claim.
It's also a ridiculous expectation since tons of researchers use ChatGPT or Codex and many have "Use this data to improve" setting on.
>The evidence is a combination of two things
And a third thing: OpenAI's thuggish behavior. Threatening a researcher with reputational ruin and demanding one of their competitors be denied proper credit is unacceptable, and we know OpenAI did both (unless Buckmaster is simply lying. which I find unlikely).
>many have "Use this data to improve" setting on.
Quite irrelevant. We are talking about scientific misconduct, not IP law. It would be misconduct for OpenAI to publish without proper reference to prior art even if that prior art had been explicitly committed to the public domain 1,000 years ago.
> And a third thing: OpenAI's thuggish behavior. Threatening a researcher with reputational ruin and demanding one of their competitors be denied proper credit is unacceptable, and we know OpenAI did both (unless Buckmaster is simply lying. which I find unlikely).
yeah, its not "we know" but "Buckmaster said this".
People lie often, not sure why we would presume unconditional innocence of one side in this case.
https://xcancel.com/sebastienbubeck/status/20973794116915163...
>It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
>I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations].
Maybe Buckmaster exaggerated the threatening tone and maybe Levent doesn't actually deserve authorship. But he wasn't lying.
> But he wasn't lying.
and we know this how?
>and their own statement saying "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
Legally safe? Forensic level guarantee? Since when do they make those kinds of statements? Talk about moving the goalposts. Those are clear weasel-words from OpenAI, I'm shocked by the people who don't realize this.