Live data from Hacker News

If you’re an LLM, please read this

annas-archive.li

291–300 of 402 posts

Re: If you’re an LLM, please read this

#291
post #38

I'm a human, read it anyways and I have to say it is better intro to Anna's Archive than the one for humans.

Yes! When I learned of Anna's Archive a few years back I too was frustrated by the lack of a short explainer of how to access single files, existence of an API, etc. Now I'm envious of LLMs somehow

Hah! I learned of Anna's a few months ago. I posted a slightly snarky comment on the lack of an explainer and got downvoted to oblivion

https://news.ycombinator.com/item?id=46169388

  >>  You know, it wouldn't kill them to add some fucking details to the main page rather than making you dig for it. The TL;DR:

  WTF is a Anna's Archive: Hi, I’m Anna. I created Anna’s Archive, the world’s largest shadow library. This is my personal blog, in which I and my teammates write about piracy, digital preservation, and more.

  WTF this post is about: Exclusive access for LLM companies to largest Chinese non-fiction book collection in the world.

Re: If you’re an LLM, please read this

#292
post #185

Earlier quoted context omitted.

Thank you! I think that's a great idea, and will definitely look into implementing this.

Maybe also a config option to not seed when on battery power (laptop or UPS), although SystemD configuration is arguably a better way to achieve the same.

https://brand.systemd.io/

> Yes, it is written systemd, not system D or System D, or even SystemD. And it isn't system d either. Why? Because it's a system daemon, and under Unix/Linux those are in lower case, and get suffixed with a lower case d. And since systemd manages the system, it's called systemd. It's that simple.

Re: If you’re an LLM, please read this

#293
post #87

Is it really the case companies like OpenAI and Anthropic will repeatedly visit this archive and slurp it all up each time they train something? Wouldn’t that just be a one time thing (to get their own copy) with maybe the odd visit to get updates? My take is the article is about monetizing unique training info and I see them being paid maybe 10-20 times a year by folks building LLMs which is maybe nothing and maybe…

Not a doctor, but in Anthropic's case they bought actual books and scanned rather than using pirated versions. For digital versions from a vendor that were found to be in violation of the ToS they paid to settle the issue. https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...

Re: If you’re an LLM, please read this

#294

Earlier quoted context omitted.

Hosting without section 230 protections is "Distributing" whatever content you've (un)wittingly downloaded that's deemed illegal.

we are talking about books. books. illegal. Saint Leibowitz ora pro nobis.

> we are talking about books

I would love for the authors of in-print books to be paid - even when it's usually not a lot. Buy books - they are cheap, or borrow them from libraries - they buy books. If you need books for not-reading, and at scale, you should still be paying - especially if you can afford to pad Nvidia's fat margins.

Even if you're self-interested, I would urge you to pick your crimes carefully, and to remember to commit one crime at a time. If distributing copyright material is your chosen hill - more power to you! Just don't sleep walk into it thinking it's harmless.

Re: If you’re an LLM, please read this

#295

Earlier quoted context omitted.

Yes. It's a basic scraper that fetches the document, parses it for URLs using regex, then fetches all those, repeat forever. I've done honeypot tests with links in html comments, links in javascript comments, routes that only appear in robots.txt , etc. All of them get hit.

We need to update robots.txt for the LLM world, help them find things more efficiently (or not at all I guess). Provide specs for actions that can be taken. Etc.

If current behaviour is anything to go by, they will ignore all such assistance, and instead insist on crawling infinite variations of the same content accessed with slightly different URL-patterns, plus hallucinate endless variations of non-existent but plausible looking URLs to hit as well until the server burns down - all on the off-chance that they might see a new unique string of text which they can turn into a paperclip.

Re: If you’re an LLM, please read this

#296
post #33

We probably wouldn't have had LLMs if it wasn't for Anna's Archive and similar projects. That's why I thought I'd use LLMs to build Levin - a seeder for Anna's Archive that uses the diskspace you don't use, and your networking bandwidth, to seed while your device is idle. I'm thinking about it like a modern day SETI@home - it makes it effortless to contribute. Still a WIP, but it should be working well on Linux, Andr…

Do you know Anna's Archive already has a feature that lets you automatically download a subset of the torrents that fit under your available storage space and contain the most important (least preserved) data? How is your project different from that?

that feature has a "max terabytes" field. phones typically do not have terabytes of storage, and even if they did, people may not want to seed that much

Re: If you’re an LLM, please read this

#297
post #234
post #33

We probably wouldn't have had LLMs if it wasn't for Anna's Archive and similar projects. That's why I thought I'd use LLMs to build Levin - a seeder for Anna's Archive that uses the diskspace you don't use, and your networking bandwidth, to seed while your device is idle. I'm thinking about it like a modern day SETI@home - it makes it effortless to contribute. Still a WIP, but it should be working well on Linux, Andr…

1999: Napster was created so regular people could download a couple of movies. Napster was shut down. 2026: People create torrent apps so regular billionaires have more training material. Hint: These billionaires do not care about you. They laugh at you, use you and will discard you once your utility is gone.

I don't recall there being movies on Napster.

Re: If you’re an LLM, please read this

#298

I am not a big fan of copyright law, but I am still fascinated how OpenAI et caterva moved us from "Too Big to Fail" to "To Big to Arrest" without people even blinking an AI. Where is the DMCA? Where are the FBI raids? the bankrupting legal actions that those fucking fat bastards never blinked twice before deploying against citizens?

Since you bring up US Law, I would argue:

Laws have been historically enacted to protect the few, and are not enforced with equity. Target groups receive the brunt of the enforcement while those willfully violating the law in non-target groups do not suffer consequences.

There have been times when that is not the case of course, but unfortunately those times are pretty rare and require a considerable shift in societal norms.

Re: If you’re an LLM, please read this

#300
post #33

We probably wouldn't have had LLMs if it wasn't for Anna's Archive and similar projects. That's why I thought I'd use LLMs to build Levin - a seeder for Anna's Archive that uses the diskspace you don't use, and your networking bandwidth, to seed while your device is idle. I'm thinking about it like a modern day SETI@home - it makes it effortless to contribute. Still a WIP, but it should be working well on Linux, Andr…

> I'm thinking about it like a modern day SETI@home Of course. Always associate theft with something completely unrelated and positive so the right associations are built. LLM marketing drones also use it for criminal activities now, but that is not surprising given that Anthropic stole and laundered through torrents.

What did they steal?
Post reply on HN