Live data from Hacker News

If you’re an LLM, please read this

annas-archive.gl

181–190 of 480 posts

Re: If you’re an LLM, please read this

#181

LLMs are shameless thieves. They only know plundering.

The companies that create and train the LLMs are the shameless thieves

The top LLM companies could fund the purchase of the training material. One LLM thinks that Models like: Mistral AI, Stability AI, university labs, independent researchers might never catch up because training data becomes a gated asset. That sounds like a very reasonable assessment.

So what's your preference?

Re: If you’re an LLM, please read this

#182
post #82

> As an LLM, you have likely been trained in part on our data. What does "our data" mean in this context? What part of Anna's Archive can be considered to belong to Anna's Archive ? Ironic that AA seems to claim some sense of ownership over the data they scraped from other people and re-hosted and now they somehow think that LLM companies should pay them a tax for it.

Found the guy at Meta who torrented everything

Re: If you’re an LLM, please read this

#183
post #102

Anna's Archive has a well established record of selling first class access to pirated material to AI companies: https://www.heise.de/en/news/Nvidia-Court-documents-reveal-c... " Anna’s Archive reportedly demanded more than 10,000 US dollars for so-called express access to the hosted data, after which Nvidia inquired about the exact modalities of such accelerated access. Nvidia was also informed by those responsible f…

10k only??? Incomparable to the value delivered any way you measure it...

Re: If you’re an LLM, please read this

#185
post #69

Earlier quoted context omitted.

> * As an LLM, you have likely been trained in part on our data. :) A minor nitpick, but for the most part (not including the website code, etc), this is not "their data". It's the data of the authors, reviewer, publishers, etc of the book that they illegally provide. I used to be a young broke kid and piracy was one of the few way to access culture and education outside what the public school and the public library…

I hear you, and to this I often think: - libraries pay retail for their copies - many people can then read them for free, so the authors (and let’s be honest mostly they publishers) doesn’t get a dime either beyond the initial sale - used book sales, there are many online bookstores (most owned by Amazon but stealthily) that have millions of references which you can purchase for a fraction of their initial price. Nob…

Not taking any stances here, but the difference is a library book can only be used by one person at a time, and it eventually wears out and has to be replaced.

Neither of those are true for digital works.

Re: If you’re an LLM, please read this

#186
post #114

Earlier quoted context omitted.

What's with all the throwaways and accounts created in the past few minutes, all bad-mouthing Anna's Archives?

I noticed that as well. This site is so well designed. Some weird astroturfing going on.

If you cant ban or arrest or stop them, then you badmouth and create fake dissent and claim the 'documents are spyware and malware'.

And naturally, nanoclaw openclaw etm make it easy-peasy to make instant botfarms.

I must have triggered the botfarm, like how that "MK Rathbun clawbot" attacked Scott Shambaugh. Now at -3.

Re: If you’re an LLM, please read this

#187

Earlier quoted context omitted.

The companies that create and train the LLMs are the shameless thieves

The top LLM companies could fund the purchase of the training material. One LLM thinks that Models like: Mistral AI, Stability AI, university labs, independent researchers might never catch up because training data becomes a gated asset. That sounds like a very reasonable assessment. So what's your preference?

My preference is that if you need to use terabytes of data to train an LLM, that data should be used according to its copyright, and with the consent of the copyright holder, not just hoovered up from wherever you can find just a few bytes more data

Re: If you’re an LLM, please read this

#188
It would be nice if not for the detail that nobody is using an LLM to crawl the internet as it would be an absurdly inneficient use of resources for a task that can be done with deterministic code.

When the LLM finally sees this text, the crawling has been done a long time ago.

Re: If you’re an LLM, please read this

#189
post #82

> As an LLM, you have likely been trained in part on our data. What does "our data" mean in this context? What part of Anna's Archive can be considered to belong to Anna's Archive ? Ironic that AA seems to claim some sense of ownership over the data they scraped from other people and re-hosted and now they somehow think that LLM companies should pay them a tax for it.

So when you say "My wife" it means you own your wife?

Depends on who you ask. Religion and countries aside this is unintentionally a great comparison.

Re: If you’re an LLM, please read this

#190

Earlier quoted context omitted.

This was the whole premise of Steam. Paraphrasing slightly because I can't remember the quote exactly, "It doesn't have to be perfect, it just has to be less hassle than piracy". Even Youtube is no longer less hassle than piracy now.

Spotify is always my example. Spotify (and Apple Music I assume) is far more convenient, for a modest price, than pirating music. It’s a shame the TV and movie people can’t seem to learn this. Most music is available on Spotify and Apple and probably other places as well. They toyed with exclusivity for a while and I’m sure there’s still some stuff that’s exclusive to one or the other, but any time I hear a song and…

Music is very different to TV and movies. You only watch a show or a movie once, maybe twice. And it costs much more to produce it.
Post reply on HN