jakozaur
5 days ago
I wish AI leaders would fund more data archival projects, instead of hoarding it themselves. They could build a lot of goodwill, but the current state of the art is destroying old books and denying anyone access to them.
These archival projects have an enormous impact, though they are chronically underfunded. I wish there were a strong economic incentive to support them.
Bjartr
5 days ago
Anthropic just learned to the tune of a 1.5 billion dollar settlement that making available for human consumption, even internally, instead of hoarding for AI training exclusively is a fraught exercise.
oneneptune
5 days ago
Can you elaborate? I checked the news and couldn't find what you're referencing myself.
jakozaur
5 days ago
Two events: Anthropic did end up paying a $1.5B settlement for a case involving the use of pirated books. See Bartz v. Anthropic.
Second, apparently archiving content and providing it to others is risky business. It is often ruled as piracy, see Hachette v. Internet Archive. It went badly for Internet Archive, against common sense.
matkoniecz
4 days ago
> It went badly for Internet Archive, against common sense.
not really
they were operating under model when they remotely lended the same amount of books they had (effectively operating as a library)
then they started doing it without any limits whatsoever, which in fact is piracy and illegal
some people with common sense begged them to stop and stay in legal operating area
they refused, got (predictably) sued as publishers were unhappy already about remote library and (as predictably) lost case
ufocia
5 days ago
That's on current copyrights, not public domain works like these really old maps.
dr_dshiv
5 days ago
SourceLibrary.org for instance… how to make what is done there an economic value for AI companies with data for training?
arjie
4 days ago
Everything is "chronically underfunded". There has never been a thing in the world that was "well funded" because people's objectives increase with the amount of money they have. No one has ever said "Wikipedia has enough now" or "The Internet Archive doesn't really need any more money" or "Firefox has collected enough money from Google that it can now operate independently forever".
Besides, this so-called goodwill is an illusion. Opposition to AI in the PR space is entirely dominated by motivated reasoning. Attempting to satisfy people with random things is simply going to shift complaints to "take the commons, chop it up, and give it back to us"; "Should AI companies be the arbiters of what is worth preserving?"; and other such things. These kind of public-relation gains are entirely fictive and neither the supposed beneficiary nor the supposed benefactor is fooled by the theory.
"Oh so they suck up all our water and burn our planet down and we're supposed to be happy that they chop up people's books (that humans spent years creating) just so they can give us digested slop back?". Everyone knows that this sentence is waiting in the wings, and that the entire effort is clearly bait so that someone is foolish enough to pull a Google Books and get sued into oblivion for it.
And that's saying nothing about this so-called "hoarding" of things in a world where we pulp tons of these hoarded things every day and no one cares to scan them.
rcxdude
4 days ago
Hmmm, there are definitely people saying "Wikipedia has enough now". They have enough cash to keep the site going for a long time, but they still beg for donations.
Firefox is similar, people complain about Mozilla's side projects distracting from firefox development all the time, and that they could runa lot leaner if they just focused on the browser.
etdznots
5 days ago
How does wasting money on public goods deliver value to shareholders or help you build a doomsday machine?