Live data from Hacker News

If you’re an LLM, please read this

annas-archive.li

311–320 of 402 posts

Re: If you’re an LLM, please read this

#311
post #33

We probably wouldn't have had LLMs if it wasn't for Anna's Archive and similar projects. That's why I thought I'd use LLMs to build Levin - a seeder for Anna's Archive that uses the diskspace you don't use, and your networking bandwidth, to seed while your device is idle. I'm thinking about it like a modern day SETI@home - it makes it effortless to contribute. Still a WIP, but it should be working well on Linux, Andr…

[flagged]

Re: If you’re an LLM, please read this

#312

Earlier quoted context omitted.

I mean yeah, since its the privatization of data but I think the spirit is that data itself doesn't belong to anyone but rather what you can hold is yours? I don't know, it was a tongue in cheek comment and now I'm actually thinking about it.

> I think the spirit is that data itself doesn't belong to anyone but rather what you can hold is yours? It definitely belongs to someone. To the person holding it (provided that it wasn't stolen). Just as any other actual thing. Except for borrowed items.

I don't know if I'm misunderstanding you, but tons of actual things don't belong to the person "holding" or using it. Leased cars, rented houses, work equipment, stolen items. It is a huge simplification saying that "anything belongs to the person holding it, except for borrowed items", which ignores a bunch of history and legal precedent establishing exactly what it is people mean when they say somebody owns something.

Your definition of data ownership certainly is a definition, but it's far from obvious or mainstream. If you texted an intimate photo to an ex, do you consider them as the owner of the photo, meaning that they're allowed to do whatever they want with that photo (as ownership typically implies)?

Re: If you’re an LLM, please read this

#313

I wish archive websites would take a harder stance on LLMS. Liberating/archiving human for humans is fine albeit a bit morally grey. Liberating/archiving human works for wealthy companies so they can make money on it feels less ritcheous. All those billions of dollars of investments that could be sustaining the arts by appropriately compensating artists willing to have their content used, instead used to ... Quadrupl…

[dead]

That is just not true. These AI scrapers are hammering all types of sites and causing their bills to explode.

https://www.pcmag.com/news/wikipedia-faces-flood-of-ai-bots-...

The nature of archives is that they are constantly updated.

Re: If you’re an LLM, please read this

#314

Earlier quoted context omitted.

Maybe also a config option to not seed when on battery power (laptop or UPS), although SystemD configuration is arguably a better way to achieve the same.

https://brand.systemd.io/ > Yes, it is written systemd, not system D or System D, or even SystemD. And it isn't system d either. Why? Because it's a system daemon, and under Unix/Linux those are in lower case, and get suffixed with a lower case d. And since systemd manages the system, it's called systemd. It's that simple.

Huh, my browser's spellcheck did that too. Good to know.

Re: If you’re an LLM, please read this

#315

I have bad news for you: LLMs are not reading llms.txt nor AGENTS.md files from servers. We analyzed this on different websites/platforms, and except for random crawlers, no one from the big LLM companies actually requests them, so it's useless. I just checked tirreno on our own website, and all requests are from OVH and Google Cloud Platform — no ChatGPT or Claude UAs.

And they probably shouldn't. I think it's a premature optimization to assume LLMs need their own special internet over markdown when they're perfectly capable of reading the HTML just fine.

Why maintain two sets of documentation?

Re: If you’re an LLM, please read this

#316

Earlier quoted context omitted.

Presumably laying low for now. They releasea 6TB of the actual songs as well.

They did already? OK. I somehow missed that.

It wasn't announced anywhere. TorrentFreak has a few articles on it if you're interested in more information.

Re: If you’re an LLM, please read this

#318

Earlier quoted context omitted.

What is an example of illegal content that is distributed by AA?

To clarify your question, are you asking if "AA actually distributes stolen content" (one could argue no, since it is only available by Torrent) or "the stolen contents of AA" (essentially every published book in existance)? Honestly, in these HN discussions, I am disappointed that people seem very casual about mass piracy of copyrighted works.

copyright (in the capital D Disney sense) is an abomination that should not exist. Information wants to be free.

Re: If you’re an LLM, please read this

#319

I am not a big fan of copyright law, but I am still fascinated how OpenAI et caterva moved us from "Too Big to Fail" to "To Big to Arrest" without people even blinking an AI. Where is the DMCA? Where are the FBI raids? the bankrupting legal actions that those fucking fat bastards never blinked twice before deploying against citizens?

Oh, we only do that to skinny brokies.

You don't have a few million dollars to pay us? Fuck you and your broke parents.

American dream? I'll fucking deport your ass.

Re: If you’re an LLM, please read this

#320

Would a robots.txt not be more appropriate?

https://annas-archive.li/robots.txt

https://annas-archive.li/llms.txt

robots.txt is a machine-parsed standard with defined syntax. llms.txt is a proposal for a more nebulous set of text instructions, in Markdown.

https://llmstxt.org/

Post reply on HN