Live data from Hacker News

150 gigs of ebooks, categorized

reddit.com

11–20 of 28 posts

Re: 150 gigs of ebooks, categorized

#12
This site rocks. Some pirated books well worth reading (buy the book if you can afford it):

Witten, Moffat, Bell. Managing gigabytes: compressing and indexing documents and images (How to build a search engine)

Aho A.V., Lam M.S., Sethi R., Ullman J.D. Compilers: Principles, techniques, and tools (Dragon book)

Anyone else have any suggestions?

Re: 150 gigs of ebooks, categorized

#13
I'd never heard of the .djvu (Déja-vu) format before trawling through that site. I grabbed a few books to test, got a DjVu viewer (MacDjView) and was seriously impressed. Small file size yet you get the "as original" experience. No horrible plain text or HTML conversions, just renderings that look like the original book (if a little cruftier after scanning).

Re: 150 gigs of ebooks, categorized

#14
It's not that hard to download ridiculous amount of ebooks from torrents. The hard stuff is to throw out all the ...for Dummies, ... in 24 hours, ...all the irrelevant thesis and articles, all the duplicates (and always throwing out the least convenient format, lowest quality) and to properly categorize or tag them.

But I have to say that this seems to be structured quite nicely. The thing is that for a collection like that you would really need domain experts in each field to look for the important books (even if they are in low quality). Here for example a quick glance at Bioinformatics tells me that Durbin's Biological sequence analysis and Xiang's Essential Bioinformatics is missing even though it can be found on torrents. They are relevant much more than the rest of the books

Re: 150 gigs of ebooks, categorized

#15

This site rocks. Some pirated books well worth reading (buy the book if you can afford it): Witten, Moffat, Bell. Managing gigabytes: compressing and indexing documents and images (How to build a search engine) Aho A.V., Lam M.S., Sethi R., Ullman J.D. Compilers: Principles, techniques, and tools (Dragon book) Anyone else have any suggestions?

Cormen T.H., Leiserson C.E., Rivest R.L., Stein C. Introduction to algorithms

Re: 150 gigs of ebooks, categorized

#17

This site rocks. Some pirated books well worth reading (buy the book if you can afford it): Witten, Moffat, Bell. Managing gigabytes: compressing and indexing documents and images (How to build a search engine) Aho A.V., Lam M.S., Sethi R., Ullman J.D. Compilers: Principles, techniques, and tools (Dragon book) Anyone else have any suggestions?

The Feynmann Lectures on Computation. It's a scan of a library book but has come out really well.

A History of Algorithms too.. looks amazing, lots of diagrams, and shows how algorithms were derived in the past.

Re: 150 gigs of ebooks, categorized

#18

There's got to be a way to organize a distributed mirroring effort. Each person has a file list in a flat text file in a shared Dropbox folder, maybe?

Commonly when directories like this are shared on Reddit, people are encouraged to use the Coral cache to get a distributed caching effect. You just add .nyud.net:8080 after the hostname.

http://lib.homelinux.org.nyud.net:8080/_djvu/_catalog/index_...

Re: 150 gigs of ebooks, categorized

#19
Use this link for a search interface on a faster site with essentially the same contents. Just try any subject/author keywords: http://gen.lib.rus.ec

The menus there link to contents, torrents and other related stuff. E.g. the torrents are at http://free-books.dontexist.com/repository_torrent/.

This library (of which the link of reddit is one of the mirrors) has been maintained by Russian volunteers, mostly mathematicians and CS people, for a number of years. It used to be informally known as kolkhoz (Russian for 'collective farm').

Re: 150 gigs of ebooks, categorized

#20

There's got to be a way to organize a distributed mirroring effort. Each person has a file list in a flat text file in a shared Dropbox folder, maybe?

Commonly when directories like this are shared on Reddit, people are encouraged to use the Coral cache to get a distributed caching effect. You just add .nyud.net:8080 after the hostname. http://lib.homelinux.org.nyud.net:8080/_djvu/_catalog/index_...

I'm not sure the Coral cache will work with a username/password. Right now, it looks like multiple people are trying to do separate comprehensive crawls to do a complete torrent, without any coordination.
Post reply on HN