podgietaru
3 days ago
"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products."
I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.
I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."
But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
II2II
3 days ago
> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.
antisthenes
3 days ago
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.
Scale matters.
joshmarinacci
3 days ago
A change in quantity can become a change in quality.
rjtavares
2 days ago
I think the toxicology proverb applies here perfectly: The dose makes the poison.
dotancohen
2 days ago
I believe that it was Stalin who phrased it most eloquently: Quantity has a quality all its own.
danlitt
2 days ago
dotancohen
2 days ago
Thank you!
TeMPOraL
2 days ago
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> In other words: it sucks when people are using your work in a manner that you find offensive
That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.
Cthulhu_
2 days ago
Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').
But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
II2II
2 days ago
> Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet.
A more pragmatic approach is to acknowledge that concerning one's self with how something is used once it has been released is an emotional drain. It is a bit much to suggest that someone agreed to something, even if that agreement is implicit, just because they released it.
> But people are free to not publish things or post things online with a more restrictive license.
Licences are meaningless unless you have the ability to enforce them (e.g. to sue). That's why so many companies are willing to ignore the terms of open source licenses. It's also why the attempts of enforcement that we do hear about are usually backed by a third party, rather than being done by the software developer themselves. Simply put, the individual developer (or even small project) trying to make a contribution to the community would be better served by not publishing (instead of using a restrictive license) if they are concerned about how their work is used.
I don't even know if there is a good way to resolve the problem. Consider something like a DMCA Takedown notice. It removes the administrative and legal overhead to copyright infringement, yet it is also easy to abuse. For example: businesses have weaponized it by using it against individuals. Perhaps my cynicism is taking over here, but I suspect any easily accessible mechanism for enforcement would be similarly abused.
TeMPOraL
2 days ago
> But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
Right. In fact, people are also free to publish things with licenses that condition access on compensating the author/publisher, and they have both social and legal backing to enforce it. This is called "proprietary", and it's not a wrong choice - in fact outside of software, it's the default choice.
The problem is when people publish "free" and "open" as a marketing tactic, where in fact they really want to control and charge for access (whether dollars or karma or credit). That is just plain dishonesty.
user
2 days ago
amiga386
2 days ago
This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them.
Search engines were absolutely fine. They were respectful of the sources they scraped. There's now a bunch of money-obsessed shitheels doing whatever they can, as blatantly and as lazily as they can, because they believe they'll become rich. Every one of these fuckers needs to be dead or in jail before the world becomes whole again.
If nobody takes action against these fuckers, everything you care about will just nope out of existence, and these fuckers will be all you have left.
TeMPOraL
2 days ago
Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers.
I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they claimed was published for everyone to benefit, ending up as training data, and thus actually benefiting humanity, at scale far beyond the original publication ever would. Which is the Dog in the Manger attitude I point out.
amiga386
2 days ago
The problem is the rat race for the pot of gold at the end of the rainbow.
A similar thing happened when crypto was in ascendence. Web pages were getting crypto-miners injected into them. Celebrities were shilling NFTs. Everyone and their dog was on the make. The thought of riches broke the minds of millions.
The same is happening with AI. Whether it's the "major AI companies" or millions of self-interested also-rans with fewer moral scruples scraping, the problem still exists. The problem will continue to exist until you can't conceiveably make money by scraping like a bastard. If everybody identified themselves up front and respected robots.txt, there would not be a problem. But they don't, and they don't, and they pummel websites for no fucking reason, and they don't care, and they won't stop.
Websites get hammered by millions of unique IP addresses from residential ISPs which happen to belong to botnets, none of them identifying themselves as a bot user agent, all just pretending to be some slightly out-of-date version of Chrome.
And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
TeMPOraL
2 days ago
You may be right about the scale of also-ran operations, even though I disagree with the core comparison. Unlike crypto coins, which require superlinearly growing energy waste just to sustain their basic guarantees, and were created to solve "problems" that aren't problems and don't need solving (hint: trust is a feature, not a bug), AI actually works. It delivers real value for cheap. The growth isn't artificial, it's a product that doesn't even need much marketing[0] - it exploded organically since ChatGPT, since it's obviously that immediately useful for approximately everyone in some aspects of their work or life.
One nit though:
> And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
They paid for it, much like everyone else.
> Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
Here you're arguing against the basic market economy. They aren't using and disposing of anything they couldn't buy for that purpose like literally everyone else. There's no theft or trickery going on here. There's a boom, because AI is that useful, but it's still all normal resource allocation.
--
[0] - Of course the competing players invest tons in marketing to gain an edge against the other players.
amiga386
2 days ago
> They paid for it, much like everyone else.
That's really not how it works. Did you miss the part where the RAM vendors, had they known their rivals were also selling all their RAM to OpenAI, would not have sold what they sold? OpenAI, who did know this, and what it would mean for every other industry and device on the planet that needs RAM, chose to go ahead with it. Because all they care about are themselves, fuck everyone else I got mine
Markets are imperfect because of information asymmetry, aka the Cheating Bastard Problem.
After announcing the deal and the truth coming out, they could also have said "hey, we see now this crimps all other industry but ours. We'll make this right, we'll reduce our RAM order if our vendors would prefer that." But they didn't, because fuck everyone else I got mine
As the AI companies are so flush with cash, they could afford to offer a Common Crawl, so everyone who wants web resources for their models or whatever can get very fresh cached copies from a common crawler, which obeys robots.txt and crawls in a fair way so as not to overwhelm the source webservers. But fuck everyone else I got mine, so here come a million dodgy scrapers to bring down everyone's site.
And in terms of natural resources, one does not simply "pay" and get whatever they want. Any government worthy of the name puts its own residents needs above the lucrative offers made to them by amoral outsiders who couldn't care less if the residents lived or died.
If you're a consumer of products and you know the product was made by taking away vital resources from people elsewhere, then you're complicit in that action by continuing to buy the product.
hack1312
2 days ago
Your rhetoric is harsh but you’re not wrong.
chrka
2 days ago
I also recently made an open-source project with 200 GitHub stars private. I never had a problem with others using it as the basis for their own projects. In fact, that happened, and I received credit for it. But LLMs just hoover up everything, process it, and then spit it back out as if it were their own.
weitendorf
3 days ago
I open source as much software as I can because I want the models to train on it and get better at it!
Scoundreller
3 days ago
I feel the same about my old blog articles.
I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.
Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.
Please do train on my pre-LLM gold!
da02
3 days ago
Do you have any active sites or social media with your recommendations?
(I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)
Scoundreller
3 days ago
Nothing active anymore, all converted to static sites.
If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.
Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.
protocolture
2 days ago
>But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I dont get this.
You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?
So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.
Meanwhile ISPs the world over have been monetizing access to your content.
I feel like what you mean is that you want control over attribution.
danlitt
2 days ago
I think the sheer magnitude of the economics have made the scales fall from a lot of people's eyes. For decades people put stuff on the internet for free on the assumption it was "not worth" anything. It turns out that as soon as that commons can be enclosed, we can marshal hundreds of dollars for every single living human, to pay for this commons to be repackaged. The money is there, and we're happy to spend it, we just won't spend it on you.
brookst
2 days ago
So basically the “information is free, encyclopedias are expensive” phenomenon from the pre-internet days? Collection, collation, and distribution have more and different qualitative value than the sum of the individual bits.
gapan
2 days ago
I think he just means that AI companies are not people.
ryukoposting
3 days ago
> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.
danlitt
2 days ago
If LLM training is found to be a fair use (looks likely), no license will help.
dotancohen
2 days ago
My understanding of the GPL is that it currently prohibits the training of LLMs, no need to explicitly mention them. MIT does allow it, however.
chrka
2 days ago
Exactly. Just like using LibGen is prohibited.
pluc
2 days ago
Why would you offer data for free if people are going to pay someone else to access it?
The whole internet is about to experience USA-level greed because people are too dumb not to use AI. Just like Google killed small sites because people were too lazy to look at page 2.
xtiansimon
a day ago
> “…gets scraped and added to a pile of training data to later be monetized…”
I feel you. It reminds me of the "Freedom for any purpose" doctrine of FOSS and the moral question of FOSS in military usage.
vachina
3 days ago
> But now I'm really reluctant to give more stuff to the free web.
I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).
bigstrat2003
3 days ago
That's only if one is stupid enough to give an LLM access to their computer. Don't do that, it's incredibly irresponsible security wise.
dotancohen
2 days ago
The vast majority of people are in fact this stupid.
I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.
So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.
CamperBob2
3 days ago
This attitude makes zero sense to me. You benefit from the "training data" just like everyone else does. If you don't, that's a problem with you, not a problem with AI models.
AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.
As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.
So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.
akudha
3 days ago
It does make sense. When podgietaru wrote their doo-dad programs, I guess they made their work available for humans to use and they didn't mind giving their work away for free. Some other person might say "I'm giving my work away free, I don't care if humans use it or AI, for any purpose". Both are legitimate choices, each according to their own.
As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...
protocolture
2 days ago
>As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point).
Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.
Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.
>If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.
CamperBob2
3 days ago
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.
Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.
podgietaru
3 days ago
I wouldn’t be happy if those companies used my vegetables to open a Michelin star restaurant, even if I got a free McDouble out of the deal.
Plant based I guess I don’t know metaphors are hard.
CamperBob2
3 days ago
But the McDouble isn't an adequate comparison. Yes, the companies who opened Michelin-rated restaurants with the food they stole from your garden are giving you McDonalds'-level food for free and charging for the rest. Meanwhile, some other companies who raided your garden are giving you the plant-based equivalent of Ruth's Chris or Fogo de Chão for free.
The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.
foco_tubi
3 days ago
The plant metaphor breaks down when you pause, take a breath, put the McDouble down and realize that intellectual property is a completely different ownership concept than physical property.
CamperBob2
2 days ago
Exactly. Plants, like flames, reproduce themselves.
foco_tubi
2 days ago
Technically most flowering plants need an intermediary for pollination
CamperBob2
2 days ago
The best kind of correct!
cycomanic
3 days ago
But the AI companies are not publishing the models? Moreover they are charging for access to the models.
CamperBob2
3 days ago
Sure they are. Surf around on HuggingFace and you'll see dozens of open-weight models up to 120B parameters in size from for-profit US companies, freely downloadable (if not freely runnable, alas). Probably hundreds of them, at this point.
And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.
They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.
podgietaru
3 days ago
I don't know how else to explain it. It doesn't feel the same to have my contributions smushed into a linear-algebra machine? For me, the incentive was the idea that my contributions might have directly helped someone. That absolutely does not feel the same when I think "My answer has modified the back propagation of a Machine Learning training run."
It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.
Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.
CamperBob2
3 days ago
Understood, there's certainly nothing invalid about your point of view here. It's just not one that I personally can come to terms with.
I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.
jtuple
3 days ago
> Now you can do something with your time besides writing pointless one-off doo-dads
Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.
Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.
AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).
Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.
dminik
3 days ago
This is such a sad worldview. Life without any challenges. Every problem immediately solved without learning anything.
CamperBob2
3 days ago
What's sad is somebody giving you a godlike tool and watching you mope around muttering about "life without any challenges."
yuye
3 days ago
It's only considered "godlike" by those satisfied with mediocrity
CamperBob2
3 days ago
The person I replied to is apparently very content with mediocrity.
foco_tubi
3 days ago
The godlike tools are paid subscriptions at best and restricted to a handful of corpos by the government at best. The free scraps thrown to the proles are mostly novelty. The other day I asked 5.6 Sol to identify a push prop airplane and it said “this is a helicopter.”
danlitt
2 days ago
> it cost hundreds of millions of dollars to generate the weights
It cost hundreds of billions of dollars to generate the training data, they just didn't get paid.
yuye
3 days ago
>Now you can do something with your time besides writing pointless one-off doo-dads.
This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.
mtVessel
3 days ago
| Now you can do something with your time besides writing pointless one-off doo-dads.
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators.
CamperBob2
3 days ago
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators
That also makes no sense, but I don't know what else I should have expected.
qotgalaxy
3 days ago
[dead]