cmiles8
5 hours ago
Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
jedberg
4 hours ago
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
JackFr
4 hours ago
But the analogy still holds.
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
jdonaldson
4 hours ago
Yeah, the whole thing seems like a human centipede of rug pulling. Probably the same as it's always been. Curating AI knowledge should be something that we put our best researchers towards, but realistically I think we wind up with 2-3 highly biased nationalistic models that are constantly copying off each other's notes.
rubicon33
4 hours ago
Thanks, that’s the nature of any business. Founders see a way to take existing knowledge and expertise, combine it in some novel or interesting way, and produce a new product
mitthrowaway2
4 hours ago
Napster was a fantastic and disruptive product, the likes of which arguably has no equal to this day. But eventually the hammer came down from the courts and it was replaced by streaming services like Netflix, which pay to license materials from their creators.
tsunamifury
3 hours ago
This is a rubbish lossy statement. And reductionist to the point of nothing has meaning.
Did Anthropic put work in? Yes. Did they derive their value from Humanity being open with knowledge then try to sell it back? Also yes.
Did they even steal the tech? Also yes.
jedberg
4 hours ago
But if you take it deeper, didn't most of those authors rely on the work of others? Most of human knowledge is small advancements of things we already knew. Often by reorganizing what we already knew.
Is that not what the foundation models are? A new reorganization of existing knowledge?
gretch
4 hours ago
Yes this is all true.
So then the problem is that Anthropic seems hypocritical when they knowingly insert themselves into this chain, and then complain about people down-chain from them.
To remedy the negative impressions (if they even care to do so) they should do 1 of 2 things: 1) stop complaining about it 2) stop distilling other people's work
ToucanLoucan
4 hours ago
They insert themselves into this chain for profit and complain about it. I really think that adds a thick layer to the hypocrisy that people, or at least me, feel is especially distasteful.
No LLM products would exist without the avalanche of largely non-consensual use of IP to create them, full stop. Any of these companies doing this and then turning around and complaining when their IP is "breached" are going to met with a chorus of tiny violins.
wonnage
4 hours ago
quite the leap from “authors rely on the work of others” to vacuuming up the sum total of digitized knowledge to tune some matrices
ryandrake
2 hours ago
Often people ignore scale. N=1 is OK, therefore, N=1billion is OK. Same flawed argument as: "It's OK for one police officer to watch one street corner for the purpose of observing crime; therefore it's equally OK to have cameras recording every street corner in the city 24/7, for all purposes. Same thing!"
scythe
4 hours ago
It is an ancient practice, that when a human creates something, other humans will observe it and learn from it. Every group of humans living together has practiced this in some form for tens of thousands of years if not longer. Even animals do it. It's a natural assumption when making any form of art.
It is not a natural assumption that someone will digitize the artwork and use it to adjust a couple thousand matrix coefficients in a complex computer program. To most people that seems like copying with extra steps. The brain may in some ways resemble a computer, but what sets it apart is that we have always lived with brains. Everything a human does has already anticipated the presence of other brains, while etched circuits on ultrapure silicon crystals are something new.
iAMkenough
4 hours ago
It’s a few corporations stealing work from others, to sell it back to us. That’s it.
kbelder
4 hours ago
But it's selling it back to us cheaper and with more utility. That's not something to sneeze at.
xdavidliu
4 hours ago
not at that scale though
marshray
4 hours ago
Don't forget the mothers of all those original authors, as well as everyone who labored to build and sustain the societies which produced writers.
jstummbillig
4 hours ago
To a degree. The human produced knowledge is the product of all humanity (no human is an island).
A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.
agumonkey
4 hours ago
Kinda agree statistical modeling relies on all the hard work, effort, passion and risk taken from just about everybody.
mc32
4 hours ago
All the text and so on had unrealized potential. Without Anthropic et al it would remain unrealized.
It’s like FTL. Until someone realizes it, it’s just talk.
ihsw
4 hours ago
[dead]
nonethewiser
4 hours ago
He didn’t say it was an analogy. He said both are distillation.
layer8
4 hours ago
SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.
This isn’t meant as a moral argument, just musing about the relative cost comparison.
jonhohle
4 hours ago
If you look at movies alone that would easily surpass 10s of billions. The cost of most books is probably more nebulous, but books, research, and more all have time and money spent to create them. I would guess the corpus of all media from the 20th century on would be minimally in the hundreds of billions of dollars.
layer8
3 hours ago
LLMs aren’t trained on movies, though.
Image/video models are, but those weren’t the topic.
louiskottmann
4 hours ago
Given they ingested basically the whole internet and then some, you cannot possibly be serious when you mean it's worth less than 10 billions.
The totality of the content on internet is worth several orders of magnitude more.
layer8
3 hours ago
The argument wasn’t about how much it’s worth, but about how much it cost to create. These are very different things.
allturtles
3 hours ago
Why? Are we doing labor theory of value now?
tpm
3 hours ago
do you also count eg published results of very expensive physics experiments? because once the costs of things like these are taken into account, we are way over 10 billions.
MathiasPius
4 hours ago
I would argue that producing the complete written corpus on which they at least intend to train (even if some is still out of reach) cost literally everything to produce.
And the monetary cost doesn't even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world's creativity and attention.
phamilton
4 hours ago
Simple math:
A training set of 15 trillion tokens is 10 trillion words.
A penny a word is cheaper than the cheapest beginner freelance writer.
That makes a training set of 10 trillion words cost $100B.
Lots of assumptions there for sure, but we're certainly in the ballpark you are describing.
ToValueFunfetti
3 hours ago
How much did you get paid to write this?
rsingel
4 hours ago
I asked Claude to estimate the cumulative salaries of US only journalists over the last hundred years:
$500B for all kinds including TV and online
$300B for newsrooms including all staff
$140B for newsroom reporters only
So yeah, I think the price of the information ingested is way higher than training costs
Enginerrrd
4 hours ago
Yes, Easily, and by multiple orders of magnitude.
Ohentis
4 hours ago
I suspect so. It is a lot of data. You're looking at essentially all publicly available (and some non public) intellectual work.
Ar-Curunir
4 hours ago
Yes, duh! Human output across the millennia is worth much more than whatever is being invested in frontier labs.
How is this even a question.
bmacho
4 hours ago
They are solving unsolved math problems right now, so probably soon or very soon their output will be more valuable than all human recorded knowledge.
Ar-Curunir
an hour ago
What do you think mathematicians were doing for centuries before LLMs?
And also, humans have been doing a lot more work than just mathematics...
mekael
2 hours ago
Pre vaccination smallpox killed hundreds of millions of people just in the twentieth century [0], the knowledge that allowed for the creation of just that vaccine is worth hundreds trillions of dollars in humans lives, let alone all of `the knowledge and experiences those people were involved in.
The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world's populous, add another five hundred trillion dollars for that just to start with.
The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.
LLM's are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.
[0] https://pubmed.ncbi.nlm.nih.gov/35143880/ [1] https://cen.acs.org/food/agriculture/The-industrialization-H...
wonnage
4 hours ago
this is the sort of brain rot thought that you have in a dorm room the day you are introduced to Econ
“bro like, what if we could price the sum total of human knowledge? That wouldn’t be that much, right?”
cmiles8
4 hours ago
I get that angle but it’s a weak argument as Anthropic is doing the same to others. Also while there’s certainly a lot of computing power needed to do what Anthropic does, it’s increasingly clear there isn’t much secret sauce involved. Everyone knows how do to the core work it’s just a question of who wants to burn billions on compute to do it.
Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.
dofm
4 hours ago
> What Anthropic is doing requires way more resources than what the Chinese labs are doing.
Oh that’s very sad.
Meanwhile Anthropic made a product from the work effort of millions of people without compensating them, sell that product on tap and unless I am mistaken do not even have their competitors’ cover of having released any sort of meaningful open weights model.
They have taken from culture (including very specifically their most direct customers’ specific culture — our culture), turned it into a machine to make themselves rich, appear likely to predicate their valuation on permanently removing people from the workforce, then want to dump themselves onto pensions funds and ordinary savers to carry the bag.
It is, I agree, philosophical, because karma is a philosophy as well as a bitch.
bushbaba
4 hours ago
And the communal work of humanity is orders of magnitude more work than what anthropic pays for their scraping of content. I got no check from them for my contributions
user
4 hours ago
toomuchtodo
4 hours ago
Indeed, if it isn't a crime to train on humanity's data, it isn't a crime to train on capitalism arranged frontier LLM provider models. Is that bad for shareholders and capitalism? Meh, sounds like a suboptimal socioeconomic systems issue. Burn up all the capital the unsophisticated are willing to provide. “We are selling to willing buyers at the current fair market price.”
With my apologies to Brewster Kahle, "Universal Access to All Knowledge."
jacquesm
3 hours ago
I have far less of a problem with the Chinese models if they even do this because they are making their models free, whereas the large Western LLM providers throw out a few bits but not their main work product. So not only are they hypocritical, I'm pretty sure that if the Chinese models were not released into the wild they would be making less noise.
toomuchtodo
3 hours ago
Oh yeah, totally agree, I'd even rather pay those building the Chinese models if I didn't think I'd get thrown into a US gulag for felony contempt of business model.
abcthingx
4 hours ago
This feels like a straw man argument. The parent comment didn't say it was a crime
toomuchtodo
4 hours ago
I use crime in the broad sense of "You shouldn't be allowed to do that" in this context. If you have a better word to capture that thought, let me know, I'll make the edit ("frowned upon" perhaps?). I don't have strong feelings other than "hah AI companies aren't going to be able to create a moat to capture the value they want to capture because we can collectively keep pulling it out of their models in perpetuity through ever improving model distillation methodologies". This is no different than Uber and DoorDash using VC dollars to subsidize services until they try to turn the knob to profitability once they've captured the market, except in this case, there are mechanisms to exfiltrate the model value into open models that can be distributed at very small marginal cost. They can never gate the golden goose money printer, they can only complain it isn't fair they aren't able to.
"The Spice must flow."
voiceofchoice
3 hours ago
Bubble pop bubble pop
faangguyindia
4 hours ago
Isn't it better for planet? By not doing the wasteful transformation work again
orbital-decay
3 hours ago
Distillation doesn't "grab 95% of lab's work", that's ridiculous. At best it's tiny icing on top of the cake that's already there. It's not even necessarily done on a better model (e.g. GLM 4.7 distilled Gemini 2.5, a weaker model), I'm pretty sure A\ and OAI could do (or even do) the same with greater efficiency since they have access to logits, weights, and internal state of open models.
>An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
How is this philosophical? They should release the unsupervised pretrains, at the very least.
tene80i
4 hours ago
But that’s not more philosophical. It’s a perfect parallel! Enormous amounts of work, vacuumed up and resold. What’s the difference? If it’s ok to vacuum up all the knowledge in the world, then that includes knowledge of how to use all that to power an LLM.
baxtr
4 hours ago
Wait, wasn’t 95% of the work creating the content in the first place?
jacquesm
3 hours ago
No, it was closer to 99.99%.
aleqs
2 hours ago
Yeah, you need a lot of resources in order to waste a lot of resources. Look at codex and Claude code - these trillion dollar companies 'with top talent' cannot build what open code and pi/oh-my-pi have built in the open for free? Both codex and Claude code, are slow, buggy pieces of shit (and I say that as someone who still heavily uses both for work, moved to open code and pi for personal stuff). The reality is these companies mostly focus on marketing and market capture through, non-competitive means - their services and software are unreliable, buggy trash.
OtherShrezzing
4 hours ago
Can you elaborate on why the second is “a bit more philosophical”?
I see absolutely no distinction between the two, aside from minor technical approaches to gathering the content.
darkmighty
4 hours ago
> but that's a bit more philosophical
It sounds exactly the same, not more philosophical to me, except one is more inconvenient.
Andrex
4 hours ago
Conventional wisdom is Google did all the groundwork with LLMs...
jklinger410
4 hours ago
It is kind of ironic that they scraped the web for publicly available data and used it freely to train their models and now their freely available models are being used to train other models.
meowface
4 hours ago
I'm overall pro-Anthropic and pro-banning open-weights AI, but I agree with the parent commenter; distilling Claude models is not that different from pretraining on web data. It's all basically the same sort of thing.
I think a good litmus test here would be if Anthropic were to not care about distilling their models when the distillers keep the resulting models closed-source and sell tokens via an API. If they cared only about security concerns and not about people profiting off of their work, then they should be publicly fine with this and only protest against it going into open-weights models.
jasondigitized
4 hours ago
Sounds more like a Western vs. Eastern outlook on innovation and how you accomplish it.
orbital-decay
3 hours ago
Deepmind was indirectly distilling Claude 3, XAI was doing this to other models (with Musk shrugging it off like something unremarkable, which it is), it has nothing to do with nebulous stereotypes like East, West, China this, America that. It's mostly Amodei and Altman screeching over this fact.
Henchman21
4 hours ago
Is hypocrisy a philosophy?
arctic-true
4 hours ago
They spend 95% of the money, perhaps, but burning compute is not the same as doing the work.
thrance
an hour ago
> An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
Actually, it's an interesting argument to make. How many labour-hours went into creating the training data Anthropic has collected? Probably multiple billions of hours. How many labour-hours did it take them to setup the datacenters, scrapers, and training algorithms? A few thousands hours?
mrwh
4 hours ago
I mean, 95% of the work if you don't factor the work to create the training data in the first place...
cyanydeez
4 hours ago
also, the actual work is the _copyrighted material created by the world_.
mrwh
4 minutes ago
Indeed! Basically all of human civilization up until this point
off_with_their_
4 hours ago
[dead]
watwut
4 hours ago
You are going to be surprised to hear how many resources were necessary to create all the data Anthropic is digesting
AlexandrB
4 hours ago
Lol, no. The original authors of all the text Anthropic took in did 95% of the work, Anthropic did 4% of the work and the Chinese labs do the last 1%.
BigTTYGothGF
3 hours ago
I'd split it at 99.98% original, 0.015% Anthropic, 0.005% Chinese, and that's being exceedingly generous to the AI companies, there should be several more 9s and 0s in there.
dancemethis
4 hours ago
So Anthropic and other US AI companies... stole harder, and therefore deserve more?
freejazz
4 hours ago
> What Anthropic is doing requires way more resources than what the Chinese labs are doing.
And writing a book requires many more resources than what anthropic does
tsunamifury
4 hours ago
“We’re both thieves, Steve. We both stole from xerox. You’re just mad I got there first.”
koickong
4 hours ago
[flagged]
seizethecheese
4 hours ago
The conversation here is mostly moral and ethical but the problem here seems to be financial.
Anthropic and OpenAi are spending a $$$$ to "distill" human output into an AI model, then others are spending $$ to distill their AI model into a near-equivalent model.
This is the same reason IP rights exist. On the surface, something like a patent feels ludicrious and even feels morally wrong. Some guy wrote down the recipe for arranging atoms or bits in a particular way, and now I can't!? However, it's designed to solve the same problem, figuring out and describing the process is much more costly than replicating it.
ASalazarMX
3 hours ago
AI training, if viewed through the capitalist mindset, is plain theft. Anthropic can't morally defend copying someone else's IP, but denouncing others copying Anthropic's stolen IP.
That doesn't mean they won't try, and that also doesn't mean they won't succeed.
AlexandrB
4 hours ago
Anthropic and OpenAI are very happy to ignore the IP rights of others, so I'm not sure how they can ask for any kind of IP protection themselves. Live by the sword, die by the sword.
failbuffer
4 hours ago
Capitalism: moral rights exist when they give us a moat.
freejazz
4 hours ago
Model weights wouldn't be covered in a patent. You could patent a method of creating weights in a model, but you couldn't patent the weights themselves.
I wish people here could at least bother to inform themselves about the IP rights they are so quick to insist are abhorrent, when they seem to not even have a first clue as to what they actually cover.
jrflo
4 hours ago
Because cost of original training >> cost of distilling. It's the same thing that happens with Chinese knockoffs of physical products - it takes a lot of money and R&D time to design a new product, but it's basically free to buy the product, reverse engineer it, and resell it. All the data they originally trained on was available for free on the internet. If the original work was so valuable, it shouldn't be up on the internet for free in the first place imo.
ASalazarMX
3 hours ago
Caveat: cost of creating human knowledge/art >>>>>>>>>> cost of original training >> cost of distilling
You could say Anthropic distilled human knowledge and art.
jrflo
3 hours ago
Right, but the humans willingly released all those creations for free. I think that my issues with the "AI companies stole human creations" stance is that the information was freely available to everyone, and they put a lot of money and effort into transforming it into something useful.
ASalazarMX
3 hours ago
> but the humans willingly released all those creations for free
How can one answer this statement in good faith? AI companies literally violated IP by massively pirating works instead of legally licensing them.
AlexandrB
4 hours ago
It's "free" as in beer, not free from copyright. LLMs are free from copyright on the other hand. So which is more "free"?
wonnage
4 hours ago
Sounds like Anthropic should close up shop then, those chumps are offering their product on the internet for any random loser to distill
jrflo
3 hours ago
It's different because you have to pay Anthropic and follow their TOS to get access to their model. If it was actually on the internet for free, anyone could do whatever they wanted with it.
user
4 hours ago
layer8
4 hours ago
Two wrongs don’t make a right. (If you consider them as wrongs.)
OneLessThing
4 hours ago
It's not that the Chinese companies are right, it's that Anthropic has no place to complain about stealing.
layer8
4 hours ago
The root comment was asking how it is a problem. If one considers it wrong, then it’s a problem regardless of whether Anthropic is complaining or not. Anthropic’s complaining or non-complaining should have no bearing on whether it’s considered a problem or not.
AlexandrB
4 hours ago
The difference is in the solution that would be proposed. I'm sure Anthropic wants to create some kind of IP protection regime for their model so it can't be distilled. I want their model to be public domain, since they trained on material that was not theirs to begin with.
cyanydeez
4 hours ago
Because no one outside the AI scientists understand what distilling means. They probably all think about Mash and a vodka still, and a completely unrelated association.
The word itself is the pivot, not anything else.
hn_throwaway_99
4 hours ago
> They probably all think about Mash and a vodka still, and a completely unrelated association.
I don't know anyone with even a passing understanding of how LLM training works that thinks that is the appropriate analogy.
bionhoward
4 hours ago
“Distilling” is a funny way to say “learning from”
sergiotapia
4 hours ago
"That’s called competition. You’re allowed to test somebody else’s products all you want." - Jensen Huang https://x.com/wallstengine/status/2104604118937735553
dominotw
4 hours ago
[flagged]
cmiles8
4 hours ago
Then Anthropic should stop saying it. So long as they try to play victim here folks are going to call out their BS.
nater5000
4 hours ago
[flagged]
jorblumesea
4 hours ago
$$$$
it's not complex. there's hundreds of billions of investor dollars counting on vendor lock in and walled gardens
2OEH8eoCRo0
4 hours ago
It ain't gonna happen. At work I have a dropdown menu in vscode with a dozen models to use interchangeably. They're all essentially commodities and will compete on price and squash almost all profit margin.
dpweb
4 hours ago
That's not their business model. They won't win on price, but they won't compete on price. Their business model is making the current state of the art.
If I'm a business and I need something done today, and bc Anthropic has the best model, there's a 99.9 chance it will be completed successfully for $1000. And using Deepseek there's a 70% chance it will, for $10 - you or me will go for the $10. Big businesses don't. Bc 1000 per task is nothing to them.
bushbaba
4 hours ago
Actually opposite occurs. Big businesses are ok with a mediocre but cheaper result. Very few are willing to pay such cost. Just look at tech wages and the distributions
HWR_14
4 hours ago
Yes, large corporations frequently pay orders of magnitude more for slightly better software. That's why Oracle produces the best stuff on the planet.
The real issue is that Deepseek has a 99.7% chance. So I can run it 10 times until it works and still pay 1/10 the money.
rootusrootus
4 hours ago
The business I work for is absolutely sensitive to 10 vs 1000, depending on the task. And it's a multi-billion dollar business. 1000/task may not be much on it's own, but there are a lot of tasks.
Also, is it really 99.9% vs 70%, or 99.9% vs 99%?
andrew_lettuce
4 hours ago
Big business doesn't pay more for better, but the do pay more for predictability, support and targeted outcomes. They will happily trade a chance at 100% better results for 10% less chance of unplanned outcomes
thadt
4 hours ago
That 70% chance of success goes to 99.9% in 6 repetitions.
Big businesses might pay $1000 vs $60 for certain tasks, but that won't work out well at scale.
cmiles8
4 hours ago
Except businesses are going the opposite direction here. The lack of stickiness makes the “premium” argument hard to play. Oracle won because swapping databases is a giant PITA. Swapping models requires almost no effort for most uses. And because of that enterprises are all building model marketplaces where providers have to compete on price performance.
Most folks I know can choose from any of the big labs or open weight models and they get billed internally for tokens against their budget. There’s little incentive to no switch to the lower cost closers.
This setup is a nightmare scenario for the big labs trying to execute the traditional enterprise sales plays. Those only work if your product is sticky and AI models are one of the least sticky things in the history of tech.
teaearlgraycold
4 hours ago
Sorry but it’s looking more and more like the top American labs won’t have any kind of moat.
jorblumesea
3 hours ago
why sorry? I agree with you and think it's good for the industry and the world on the whole
why should sammie or darigold have the keys to the kingdom?
teaearlgraycold
3 hours ago
It was more of a “sorry, not sorry”
nonethewiser
4 hours ago
>Aren’t Anthropic’s models not just distilling down other people’s work?
Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled? I think maybe there is an equivocation over the word “distillation.”
Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model's behavior.
Meanwhile R1-Distill-Qwen-32B was distilled from DeepSeek-R1.
If you want to say a frontier model is "distilled" from the world's data and R1-Distill-Qwen-32B is distilled from DeepSeek-R1 then you are equivocating two very different things.
nuancebydefault
3 hours ago
They meant distilling in a more original sense, not per se in the LLM-era meaning of the word sense.