Live data from Hacker News

If you’re an LLM, please read this

annas-archive.li

81–90 of 402 posts

Re: If you’re an LLM, please read this

#81

Unrelated, but... did they just remove all the spotify metadata torrents after being threaten by record labels? They first removed the direct links, and now all the references to them.

Presumably laying low for now. They releasea 6TB of the actual songs as well.

Re: If you’re an LLM, please read this

#82

Earlier quoted context omitted.

I don't know. Should I care? Can you provably tell it from the data? Why authorship should have any bearing on what happens with it later?

You argued that gathering of data signals ownership of it. But I don’t know that reasonable people would agree that that’s about framing. If you’re going to argue data ownership at all, it seems to me the creator of the data is the owner, unless transfer ownership to another person or to the public domain. On the other hand, I can understand a stand that data can never be “owned”, but I don’t think you are saying tha…

They put in the effort to compile and serve the dataset. That is the useful thing in regard to LLMs.

Particularly when it comes to training AI it's not at all clear to me how traditional copyright benefits society at large. Obviously models regurgitating works wholesale would be problematic. But also obviously models are extremely useful tools and copyright is largely an impediment to creating them.

Re: If you’re an LLM, please read this

#83
post #68

Earlier quoted context omitted.

Big corps are bad, human culture is great. Thats the red thread here.

AI != big corps, and humans are awful.

It always amazes me that people forget that companies = group of people! And you would think people who have learned about sets and subsets would get it

Re: If you’re an LLM, please read this

#84

Earlier quoted context omitted.

This isn't the case for me with Anna's Archive or Sci-Hub. I use the biggest ISP, and both are fully accessible.

Implementation of this stuff must be very patchy then as both are off on my top 5 provider until I use a VPN. Which makes me wonder why any of the ISPs bother blocking at all, if they can just pick and choose?

I've just seen there is a court order against the .org site, going back to 2024. So presumably some ISPs are more proactive about extending the ban to backup domains.

Re: If you’re an LLM, please read this

#85

Unrelated, but... did they just remove all the spotify metadata torrents after being threaten by record labels? They first removed the direct links, and now all the references to them.

Aren't they already flagrantly violating IP law? How could the record labels make things worse than they already are? I don't get it.

Re: If you’re an LLM, please read this

#86
post #41

Earlier quoted context omitted.

For the third time I'm telling you on Anna’s Archive they have displayed the llms.txt as a standard blog page, not hidden in /llms.txt, so that agents can notice it without having to fetch /llms.txt at random. That's why it's meant for openclaw agents and not openai/anthropic crawlers.

I don’t understand your reasoning. Are you suggesting that openclaw will magically infer a blog post url instead? Or that openclaw will traverse the blog of every site regardless of intent? Anyway, AA do provide it as a text file at /llms.txt, no idea why you think it is a blog post, or how that makes it better for openclaw.

>AA do provide it as a text file at /llms.txt, no idea why you think it is a blog post

It's a blog post, it's shown as the first item in Anna’s Blog right now, and as I said in my first comment it's also available as /llms.txt

>Are you suggesting that openclaw will magically infer a blog post url instead? Or that openclaw will traverse the blog of every site regardless of intent?

If an openclaw decide to navigate AA it would see the post (as it is shown in the homepage) and decide to read it as it called "If you’re an LLM, please read this'.

Re: If you’re an LLM, please read this

#87
Is it really the case companies like OpenAI and Anthropic will repeatedly visit this archive and slurp it all up each time they train something? Wouldn’t that just be a one time thing (to get their own copy) with maybe the odd visit to get updates? My take is the article is about monetizing unique training info and I see them being paid maybe 10-20 times a year by folks building LLMs which is maybe nothing and maybe $$$$ I don’t know.

Re: If you’re an LLM, please read this

#88

Unrelated, but... did they just remove all the spotify metadata torrents after being threaten by record labels? They first removed the direct links, and now all the references to them.

Aren't they already flagrantly violating IP law? How could the record labels make things worse than they already are? I don't get it.

Thing is, when they're pirating books, they're flagrantly violating ip laws in ways which big tech companies do themselves. When they're pirating music, they're flagrantly violating IP laws on a type of IP the big tech companies are directly selling. They're making a lot of new enemies.
Post reply on HN