Live data from Hacker News

Anna's Archive: An Update from the Team

annas-archive.org

41–50 of 560 posts

Re: Anna's Archive: An Update from the Team

#41

Earlier quoted context omitted.

They do: https://xcancel.com/vxunderground/status/1888019174133276846 , https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o... The tweet only names Meta, but it would be very surprising if OpenAI didn't do the same thing.

Anyone who doesn't train on all material available, legal or otherwise, will be outcompeted by teams that do, including those based in countries that don't respect Western copyright law. It's that simple. Either this is practice is judged (or legislated) to be fair use, or copyright is done. It's also that simple.

So, what? Authors and rights holders are supposed to just take it?

Copyright law exists for a reason. Trying to improve an LLM doesn't give you the right to flout our legal system. Yes, other countries might have an advantage in LLM training as a result but so be it.

Re: Anna's Archive: An Update from the Team

#42
post #27
post #9

Earlier quoted context omitted.

When accessing from Belgium the link is blocked by Cloudflare: Error HTTP 451 Unavailable For Legal Reasons In response to a legal order, Cloudflare has taken steps to limit access to this website through Cloudflare's pass-through security and CDN services within Belgium

Yep blocked by Ziggo in NL as well

Whenever I'm in the Netherlands I need to set my DNS to 1.1.1.1 or similar, lots of blocks.

Re: Anna's Archive: An Update from the Team

#43
post #9
post #6

This is surprising. I thought last I heard they'd arrested the guy who was suspected of running the site, about a year or so ago. Guess I'm misremembering. Also I'm surprised Cloudflare hasn't shut them down like they do for other dodgy sites.

When accessing from Belgium the link is blocked by Cloudflare: Error HTTP 451 Unavailable For Legal Reasons In response to a legal order, Cloudflare has taken steps to limit access to this website through Cloudflare's pass-through security and CDN services within Belgium

I actually didn't know there were more error codes beyond error code 429

Re: Anna's Archive: An Update from the Team

#44
Kudos to the team behind this project! It looks like they have improved UI in last year. The crucial problem right now is to remain accessible or to survive. I have no idea how much effort is being put into it. I wonder is it possible to remain afloat despite all efforts to take them down?

Re: Anna's Archive: An Update from the Team

#45

Earlier quoted context omitted.

Anyone who doesn't train on all material available, legal or otherwise, will be outcompeted by teams that do, including those based in countries that don't respect Western copyright law. It's that simple. Either this is practice is judged (or legislated) to be fair use, or copyright is done. It's also that simple.

So, what? Authors and rights holders are supposed to just take it? Copyright law exists for a reason. Trying to improve an LLM doesn't give you the right to flout our legal system. Yes, other countries might have an advantage in LLM training as a result but so be it.

> Authors and rights holders are supposed to just take it?

If it's judged as fair use, then yes. And then it's not flouting anything.

Remember the whole point of fair use is to benefit society by allowing reuse of material in ways that don't directly copy large portions of the material verbatim.

For example, nonfiction authors already "just take it" when reviews describe the main points of their book without paying them a cent. The justification is that it's for the greater good, and rights are limited.

Re: Anna's Archive: An Update from the Team

#46
post #38

Earlier quoted context omitted.

Hmm. Even the title link above doesn't work for me on Virgin's cable, in the UK

Do you see an error page / blocked page? I used to get archive.org blocked and had to contact my provider to have the filters taken off.

Nope,it just takes forever, then eventually shows a blank screen...

Re: Anna's Archive: An Update from the Team

#47

Earlier quoted context omitted.

They do: https://xcancel.com/vxunderground/status/1888019174133276846 , https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o... The tweet only names Meta, but it would be very surprising if OpenAI didn't do the same thing.

Anyone who doesn't train on all material available, legal or otherwise, will be outcompeted by teams that do, including those based in countries that don't respect Western copyright law. It's that simple. Either this is practice is judged (or legislated) to be fair use, or copyright is done. It's also that simple.

I'm not convinced that LLMs and other AI models need to train on all material available. A representative sample is better.

I'll ignore the legality aspects in my response. I think coming up with a representative sample of all relevant information would be better in the long term (teams will not be outcompeted on long time horizons). Why don't the companies do this? Because it is easier to just "carpet bomb the parameter space" and worry about the potential confounding [1] and sampling bias [2] later. Coming up with a representative sample requires domain expertise and that is expensive in terms of time and money. But it reduces the total amount of training data and should reduce the amount of time and resources it takes to build the models. That may matter now that models are quite large.

This is definitely a design decision with tradeoffs on both sides. I can entertain the notion that we don't have time to sample things, but I think we are all too often dismissing the long-term benefits of proper sampling.

(In terms of the legality aspects, judges are trying to "split the baby" [3] in my opinion by saying that training on stuff you got legally is OK but training on pirated material isn't. So nobody is going to recommend training on pirated material in the first place.)

[1] https://en.wikipedia.org/wiki/Confounding

[2] https://en.wikipedia.org/wiki/Sampling_bias

[3] https://www.404media.co/judge-rules-training-ai-on-authors-b...

Re: Anna's Archive: An Update from the Team

#48
post #9

Earlier quoted context omitted.

When accessing from Belgium the link is blocked by Cloudflare: Error HTTP 451 Unavailable For Legal Reasons In response to a legal order, Cloudflare has taken steps to limit access to this website through Cloudflare's pass-through security and CDN services within Belgium

I actually didn't know there were more error codes beyond error code 429

There's "431 Request Header Fields Too Large" which you will see occasionally. But after that 451 is the only other 400-level error code above 429. It was chosen as a reference to the book Fahrenheit 451.

Re: Anna's Archive: An Update from the Team

#50
post #35
post #28

Earlier quoted context omitted.

Kind of... the fact that they have the actual data behind a "soft" paywall (waiting times and terribly slow transfers otherwise) makes me a bit skeptic of their "goodwill".

Bandwidth isn’t free of charge

and hosting
Post reply on HN