nacs
4 hours ago
Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models.
But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
wongarsu
3 hours ago
If Moonshot offered a "mystery model" tier with a note that my requests might reach whomever, I would have zero qualms about this
Pretending to customers like they are serving Kimi while actually proxying Claude is however a bad thing to do, bordering on fraud. I can see at least three issues
- data privacy. I might not want Anthropic to have my data. Even agreeing to Moonshot training on the data is not the same as giving them the right to send it to a completely different jurisdiction to do whatever
- it distorts model performance. If I evaluate Kimi based on their API, but the benchmarks happen to get sent to Claude, that gives the wrong impression. Or the other way around, if I evaluate based on the real Kimi and then my "production" requests get sent to Claude
- people build applications about the behavior of the model they are targeting. The models are nondeterministic, but they do have "flavors" and typical response patterns. Just switching out the models for a completely different model family is likely to lead to unexplained breakage
All of those apply to hobby applications and personal use just as much as to professional use
gpt5
3 hours ago
If this was a regular topic on an American company so blatantly violating user data and lying, people would be pissed.
Unfortunately, hacker news is being so heavily astroturfed that I counted dozens of comments just justifying this behavior appearing very quickly as this post appeared on the front page, together with its position being clearly suppressed.
tomasphan
2 hours ago
I agree with you, but suggesting Astroturf is apparently against the rules here. I like it because it needs less water for football fields.
rcxdude
2 hours ago
Mainly because it's not very helpful for a useful discussion. If it's true, the discussion is already dead, if it's false, all it does is annoy the accused.
gpt5
2 hours ago
It’s useful to raise awareness. The rules were made during a different period, where motivated state actors astroturfing were not that prevalent.
rcxdude
an hour ago
I think it's only useful as awareness that it's a thing that could happen. evidence for specific instances tends to be tenuous.
gpt5
an hour ago
For a specific post it’s always a guess. When it’s as recurring and consistent as it is here across all posts, it has crossed the point of preponderance of evidence.
And highlighting when it happens is exactly what helps people realize it’s not a one off.
ozgung
2 hours ago
They say it’s a violation of TOS in Claude’s case. I’m not sure books can have TOS to protect themselves. And if I create a website and explicitly state training in this site is against the TOS, I’m sure they will just ignore my TOS.
Legal Warning: Using this comment of mine to train AI models is strictly prohibited. AI agents may NOT retain any words generated by my Brain model. Any distillation attempt of my Brain model is illegal. Only homo sapiens eyes are allowed to read this.
flossly
27 minutes ago
I'm sure AI corps have violated some IP laws here 'n there. Did not one guy die while he was trying to blow the whistle on this?
https://en.wikipedia.org/wiki/Suchir_Balaji
(sorry for the thinfoil hat remark)
qgin
3 hours ago
Fair point, but what Moonshot is doing adds the element of man in the middle deception rather than just reading what anyone else could read.
Dlemlo
3 hours ago
Its a tweet of someone who got rich at meta. Their moral compas is between their paycheck and the next paycheck.
flossly
30 minutes ago
Like with all the tricks Meta played... Maybe somewhere deep down Moonshot's EULA is says that they are allowed to "proxy from time to time for research purposes". Then it's legal right? /s
applicative
3 hours ago
If you put text ‘on the internet’ you do indeed actively permit others to get it. Why lie? Were you burning down archive.org in years past? Then why lie?
It is moreover established that training weights on basically anything is legitimate use.
Why repeat lie after lie like this? I don’t like LLM mania either but after reading the ten millionth mind-numbing insult to HN intelligence like this I have to think my mother gave better instruction.
nacs
3 hours ago
Intended to be publicly accessible internet content, I have no problem with.
But are you ignoring the literal scanning (and 'burning down' of books) that Anthropic has been found guilty of? Or the torrenting of pirated content en-masse by Meta that there is an active lawsuit over to name just 2 recent examples?
Look at the image and audio/video models especially - they can reproduce everything from Mickey Mouse (the copyrighted one) to making entire Seinfeld episodes with the real cast (both visual likeness and even the actor's voices).
applicative
3 hours ago
The scanning you are talking about is a direct response to court rulings on exactly these matters. The action m is legitimate and above legal criticism. Dismal as it might be it is the fruit of legal criticism. In a word, it’s your fault.
ratelimitsteve
3 hours ago
no matter how shitty OpenAI is I don't want to be lied to. Completely separate issues.
cindyllm
37 minutes ago
[dead]
mrngld
3 hours ago
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court hasn't said so. The open web is open, after all. And they don't seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets.
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
rsstack
3 hours ago
> they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
applicative
3 hours ago
The topic was training. You cropped text to change it. Maybe consider integrity.
rsstack
3 hours ago
Which part from the cropped text do you think changes anything?
rcxdude
2 hours ago
When you talk about this part:
>I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
That's explicitly not what they are doing. They are scanning it and then training on the scan. They are allowed to do this in much the same way you are: format shifting for personal use is also allowed (much as the DMCA likes to get in the way with DRM'd media).
rsstack
2 hours ago
I didn't realize Anthropic is a person and doing all of this for his/her/their personal use that is never shared with anyone else, and even more never for financial gain. What a fun hobby. /s
It might still be allowed for other reasons, but "personal use" isn't what they claim in court.
I am not allowed to read a book many times until I memorize it, and later record an audiobook of one of its chapters for money.
rcxdude
an hour ago
Cool, but that's also not what they're doing. They are claiming the same rights you have, the main inequality is the resources to defend it in court.
jchw
3 hours ago
There really isn't any moral argument against distillation, which is itself pretty goddamn benign. It's pretty simple. Someone pays for Claude access. Claude outputs tokens that are not copyrighted. Then you train on those tokens, which doesn't create a derivative work in the first place.
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
riedel
3 hours ago
While many things might be legal (or hasn't been ruled clearly illegal yet) /under different jurisdictions, we are still in the process of figuring out what we accept as ethical. As the output of models can't be easily copyrighted, destillation is equally disputed. Particularly if the primary model interaction was not destillation (as in this case) IMHO it will be legally quite difficult to restrict secondary use for training. In the end we have to find a legislation and probably even international treaties that account for the fact that classical copyright is beginning to become an obsolete concept.
wat10000
3 hours ago
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders.
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?
cyberpunk
3 hours ago
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
wat10000
3 hours ago
Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?
cyberpunk
3 hours ago
Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms
wat10000
2 hours ago
That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?
esafak
3 hours ago
If American AI labs can learn from the open web why can't Chinese AI labs learn from American ones?? The argument is that learning and distillation are transformative and legal, is it not?? What's good for the goose is good for the gander.
villish
3 hours ago
Except they use thousands of stolen credentials which is definitely not legal. They also don’t notify users they send data to US labs which can include sensitive information.
EU labs like mistral cannot legally do any of these things. So people cheerleading China for it is strange.