Live data from Hacker News

Ask HN: Has anybody built search on top of Anna's Archive?

news.ycombinator.com

71–80 of 156 posts

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#71
post #39

Earlier quoted context omitted.

98% sounds good enough for the usecase suggested here.

2% garbage, if some of that garbage falls out the right way, is more than enough to seriously degrade search result quality.

It's better than nothing, and nothing is what we currently have.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#72

Don't do it. Just because you can, doesn't mean you should. Do you know if they have anywhere near the legal muscle to push back the flood of legal notices if you did this? Assume it survives because it doesn't have a wide open barn door to the public.

It wouldn’t be called full text search of AA, It would be called full tech search of every book in the world.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#73
post #50

They did! They conducted a competition https://annas-archive.org/blog/all-isbns-winners.html , in which a few submissions exceeded the minimum requirements and implemented a good search tool & visualiser.

How is this a text search of the books?

The original question the poster made was not clear, so this is also an answer to it. It depends on what they meant by "search"

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#74
post #2

You must mean free text search and page level return, because it already has full metadata indexing. The thing is AA doesn't hold the texts. They're disputable IPR and even a derived work would be a legal target.

> a derived work would be a legal target. Why would it? Google isn't prosecuted for indexing the web.

Oh it certainly is. https://www.reuters.com/sustainability/boards-policy-regulat...

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#75
post #18
post #14

Related question, has Anna's archive been thoroughly filtered for non-copyright-related illegal material? Pedo, terrorism, etc. I've considered downloading a few chunks of it but I'm worried of ending up with content I really don't want to be anywhere near from.

How might you inadvertently download illegal content while searching for legal content?

Seeding torrrent blocks.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#76
post #3

Honestly I don't think it would be that costly, but it would take a pretty long time to put together. I have a (few years old) copy of Library Genesis converted to plaintext and it's around 1TB. I think libgen proper was 50-100TB at the time, so we can probably assume that AA (~1PB) would be around 10-20TB when converted to plaintext. You'd probably spend several weeks torrenting a chunk of the archive, converting ev…

Decent storage is $10/TB, so for $10,000 you could just keep the entire 1PB of data.

A rather obvious question is if someone has trained an LLM on this archive yet.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#77

Don't do it. Just because you can, doesn't mean you should. Do you know if they have anywhere near the legal muscle to push back the flood of legal notices if you did this? Assume it survives because it doesn't have a wide open barn door to the public.

It wouldn’t be called full text search of AA, It would be called full tech search of every book in the world.

You are asking a judge to consider that a book is ok to scrape because it's part of a much larger collection of books, perhaps the biggest and best collection, and therefore it's all OK because at scale means good.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#78

Earlier quoted context omitted.

> a derived work would be a legal target. Why would it? Google isn't prosecuted for indexing the web.

Oh it certainly is. https://www.reuters.com/sustainability/boards-policy-regulat...

That's not prosecution for indexing the web. Google is being treated the same way AT&T was for telephones.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#79

Mebbe easier to just search Amazon or Goodreads. Like site:amazon.ca as someone has mentioned below. Every book has an ISBN 10 or 13 digit ISBN number to identify them. Unless it's some self-pub/amateur-hour situation by some paranoid prepper living in a faraday-cage-protected cage in Arkansas or Florida it's likely a publication with a title, an author and an ISBN number.

A self-pub amateur-hour book printed by a paranoid prepper living in a faraday cage is exactly the type of book I'd probably enjoy reading, but I doubt these exist anymore.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#80

Earlier quoted context omitted.

Oh it certainly is. https://www.reuters.com/sustainability/boards-policy-regulat...

That's not prosecution for indexing the web. Google is being treated the same way AT&T was for telephones.

https://harvardlawreview.org/print/vol-138/united-states-v-g...
Post reply on HN