num42
4 hours ago
Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
sam_lowry_
4 hours ago
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?
Coming out would be a bold move for them.
peri-cl
4 hours ago
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
gruez
2 hours ago
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
bix6
2 hours ago
> From the start, Anthropic “ha[d] many places from which” it could have purchased books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and chief executive officer Dario Amodei put it (see Opp. Exh. 27).
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
vaylian
3 hours ago
> Just like they want it to be illegal to run local ML inference.
Citation?
peri-cl
2 hours ago
https://news.ycombinator.com/item?id=49076057 ("Our position on open-weights models (anthropic.com)", 1812 comments)
spwa4
4 hours ago
In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.
One wonders if they're still doing it.
ungut
3 hours ago
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.
yorwba
3 hours ago
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
spwa4
2 hours ago
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
yorwba
2 hours ago
AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
deadbunny
2 hours ago
What tosh. It's copywrited material, they pay to access it like everyone else.
outside1234
3 hours ago
Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall).
toomuchtodo
4 hours ago
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.
(not legal advice!)
p-e-w
4 hours ago
There’s zero liability in a company stating publicly that they support Anna’s Archive. Zero. Free speech protections cover much more egregious statements than that.
criddell
3 hours ago
Maybe they are worried about claims of contributory infringement?
toomuchtodo
4 hours ago
I disagree. Anyone with even a hint of standing will sue, and keep suing. As someone who has to work with corporate counsel often, do not say anything you don't have to say. Only say what is absolutely necessary. Free speech protects you from your government. It does not shield you from civil suits, and the US is extremely litigious.
kmeisthax
21 minutes ago
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
xtracto
3 hours ago
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
squidbeak
3 hours ago
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
tentacleuno
an hour ago
Whilst it pays for the service (and, in that respect, may be a necessary evil), it's morally questionable (at minimum) to charge for things that, by law, aren't yours in the first place.
dessimus
an hour ago
Less morally questionable than claiming to users they are "buying" access to media that can be revoked at any point in the future with no recompense, of course referring to Sony and Amazon.
dessimus
39 minutes ago
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
orbital-decay
an hour ago
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
mips_avatar
2 hours ago
I wouldn’t include YouTube given how only Google is allowed to index it