pandoro
2 hours ago
All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
skybrian
2 hours ago
Does distilling actually violate any IP protection laws? Sure, it’s against their terms of service.
tills13
2 hours ago
I get what you're saying and two wrongs do not make a right but the irony, and why people are even talking about this, is that the thing being distilled clearly, and knowingly, violated copyright & terms across the entire internet.
skybrian
an hour ago
Yeah, everyone is saying that but I don't find the irony very interesting.
Laurel1234
41 minutes ago
If it does then their original training did as well.
applicative
2 hours ago
One mentions distillation to intimate that one is in fact still the leader. It has zero to do with fairness.
parineum
2 hours ago
For me, I don't really care about the theft aspect but when people are claiming that these open models are better value or going to overtake anthropic/openai models, the implication that the open models are training of distilled data means all the "progress" they are making is just mimiced from the closed models.
It's a bit interesting how the open models are able to keep pace with the closed models except whole maintaining a steady following time.
abdullahkhalids
33 minutes ago
It's important to note that even if these open models are distilled, they are showing genuine improvements in their architecture, which enables inference costs to be several factors below what equivalent closed models have.
The interesting question is: will Anthropic release a Fable like model with an architecture similar to Kimi, and get the inference cost gains? They should surely beat Kimi because they can internally distill as much as they want.
orbital-decay
an hour ago
Distillation is absolutely not the reason they're good. It's not necessarily even done on a more capable model. It can even be done on itself and still bring improvement, or on a weaker model as well (see GLM and Gemini, which is definitely true because it repeats Deepmind's injections).
nonethewiser
2 hours ago
Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.
shlewis
2 hours ago
Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.
user
an hour ago
mosura
2 hours ago
What of the costs for creating the data that was used to train the model being distilled?
yandie
2 hours ago
At least 1.5B by stealing books, per a recent ruling.
I don’t feel sorry for the model companies
user
an hour ago
mosura
2 hours ago
They only had to pay for storing the books on a server for later possible use. They did not have to pay anything for the training which was declared fair use.
user
2 hours ago
applicative
an hour ago
Do you people seriously believe that Moonshot and the Chinese personal-cult-state dont possess and train on the same torrents?
mosura
an hour ago
They aren’t hypocritically crying foul about it, so no we don’t care if they do.
user
35 minutes ago
pandoro
2 hours ago
Could you give an example of the value that only training a model can create but none of the rights-holder could? I feel like if you got a direct, instant communication channel to any of the rights-holder that created the content in the training set of those models, you'd get more value than what the LLM could ever give you on any specific subject.
8note
2 hours ago
the outputs of the model have no property protections, and training a model on the outputs of another model does the exact same thing - its expensive and creates new value over what was in the input - a set of documents.
jayd16
2 hours ago
I'm going to hope this was sarcasm and if it is, it's great.
titzer
2 hours ago
An incredibly ironic comment.
rvba
2 hours ago
A lot of those arguments could be said about writing a book, or a decent forum guide.