Earlier quoted context omitted.
98% sounds good enough for the usecase suggested here.
2% garbage, if some of that garbage falls out the right way, is more than enough to seriously degrade search result quality.
Ask HN: Has anybody built search on top of Anna's Archive?
71–80 of 156 posts
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#72Don't do it. Just because you can, doesn't mean you should. Do you know if they have anywhere near the legal muscle to push back the flood of legal notices if you did this? Assume it survives because it doesn't have a wide open barn door to the public.
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#73They did! They conducted a competition https://annas-archive.org/blog/all-isbns-winners.html , in which a few submissions exceeded the minimum requirements and implemented a good search tool & visualiser.
How is this a text search of the books?
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#74You must mean free text search and page level return, because it already has full metadata indexing. The thing is AA doesn't hold the texts. They're disputable IPR and even a derived work would be a legal target.
> a derived work would be a legal target. Why would it? Google isn't prosecuted for indexing the web.
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#75Related question, has Anna's archive been thoroughly filtered for non-copyright-related illegal material? Pedo, terrorism, etc. I've considered downloading a few chunks of it but I'm worried of ending up with content I really don't want to be anywhere near from.
How might you inadvertently download illegal content while searching for legal content?
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#76Honestly I don't think it would be that costly, but it would take a pretty long time to put together. I have a (few years old) copy of Library Genesis converted to plaintext and it's around 1TB. I think libgen proper was 50-100TB at the time, so we can probably assume that AA (~1PB) would be around 10-20TB when converted to plaintext. You'd probably spend several weeks torrenting a chunk of the archive, converting ev…
A rather obvious question is if someone has trained an LLM on this archive yet.
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#77Don't do it. Just because you can, doesn't mean you should. Do you know if they have anywhere near the legal muscle to push back the flood of legal notices if you did this? Assume it survives because it doesn't have a wide open barn door to the public.
It wouldn’t be called full text search of AA, It would be called full tech search of every book in the world.
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#78Earlier quoted context omitted.
> a derived work would be a legal target. Why would it? Google isn't prosecuted for indexing the web.
Oh it certainly is. https://www.reuters.com/sustainability/boards-policy-regulat...
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#79Mebbe easier to just search Amazon or Goodreads. Like site:amazon.ca as someone has mentioned below. Every book has an ISBN 10 or 13 digit ISBN number to identify them. Unless it's some self-pub/amateur-hour situation by some paranoid prepper living in a faraday-cage-protected cage in Arkansas or Florida it's likely a publication with a title, an author and an ISBN number.
Re: Ask HN: Has anybody built search on top of Anna's Archive?
#80Earlier quoted context omitted.
Oh it certainly is. https://www.reuters.com/sustainability/boards-policy-regulat...
That's not prosecution for indexing the web. Google is being treated the same way AT&T was for telephones.