Live data from Hacker News

Ask HN: Has anybody built search on top of Anna's Archive?

news.ycombinator.com

101–110 of 156 posts

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#101
post #86
post #81

Earlier quoted context omitted.

1. It'd be for the scientific community (broadly-construed). Converting media that is currently completely un-indexed into plaintext and offering a suite of search features for finding content within it would be a game-changer, IMO! If you've ever done a lit review for any field other than ML, I'm guessing you know how reliant many fields are on relatively-old books and articles (read: PDFs at best, paper-only at wor…

> I really don't see how this could ever lead to any kind of legal issue. You're not hosting any of the content itself, just offering a search feature for it. You don't need to host copyrighted material. It's all about intent . The Pirate Bay is (imo correctly, even if I disagree with other aspects about copyright law and its enforcement) seen as a place where people go to find ways to not pay authors for their conte…

> When I read the OP, I imagine this would link from the search results directly to Anna's archive and sci-hub

You could just give users ISBNs or link to the book's metadata on openlibrary[0], both of which AA's native search already does.

[0] https://openlibrary.org/

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#103

Earlier quoted context omitted.

> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.

> > or some other country that doesn't respect international copyright though. > Like the US? OpenAI et al. don't give a shit. OpenAI is not a country and therefore cannot make laws that don't respect international (or domestic) copyright. Also the US is a lot bigger than OpenAI and the big tech corps, and the law is very much on the side of copyright holders in the US.

The money is definitely in the side of big tech vs book publishers. There may be a nominal settlement to end the matter, perhaps after a decade of litigation

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#104

Earlier quoted context omitted.

> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

> that blends them thoroughly and irreversibly

It's okay, you can say 'laundering'

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#105

Earlier quoted context omitted.

> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

That's Uber's Gambit. Nothing is illegal for large enough corporations with strong network effects and deep pockets.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#106
post #28

Earlier quoted context omitted.

It would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respect international copyright though.

Do you have a reason to believe this ain't already being done? I would assume that the big guys like openai are already training on basically all text in existence.

Wasn't this confirmed what Meta does?

https://www.forbes.com/sites/danpontefract/2025/03/25/author...

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#107
post #52

Earlier quoted context omitted.

LLMs already use it, dude )

I think one use would be to search for information directly from a book, rather than get a garbled/half-hallucinated version of it.

garbled/half-hallucinated is probably what you would've gotten 8-12mo ago but now adays im sure with good prompting you can pull value from any book.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#108
post #85

Earlier quoted context omitted.

Every LLM maker probably did the same. Facebook just has disgruntled employees who leaked it

Google goes around legally scanning every book they can get their hands on with books.google.com. Legally scanning every paper they can get their hands on with scholar.google.com. I doubt they'd resort to piracy for what is basically the same information as what they've already legally acquired...

That is a good reason to think they did not but it doesn't necessarily override reasons for them to do so. Perhaps it's dubious that the subset of data they could not legally get their hands on is an advantage for training but I really don't know, and maybe nobody does. Given that, Google's execs may have been in favor of similar operations as Facebook's and their lawyers may have been willing to approve them with similar justifications.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#109
post #14

Related question, has Anna's archive been thoroughly filtered for non-copyright-related illegal material? Pedo, terrorism, etc. I've considered downloading a few chunks of it but I'm worried of ending up with content I really don't want to be anywhere near from.

The team that curates it is very dedicated and wouldn't do such a thing. The least of reasons being they don't want the heat from it.

I'm not sure what other forms of information is illegal beyond CP. In the US, bomb making instructions are not illegal. In other dictatorships or zealous religious regimes, information about democracy or works that insult Islam might be illegal

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#110

https://book-finder.tiiny.site/ More: https://rentry.co/StellaOctangulaIsCool

> https://book-finder.tiiny.site/ That just redirects to https://yandex.com/search

Well yeah but with a specific query with which you can search multiple libraries
Post reply on HN