Live data from Hacker News

Visualizing All ISBNs

annas-archive.org

81–90 of 146 posts

Re: Visualizing All ISBNs

#81
post #74

The thing is, ISBNs aren't hierarchical --- they are bought in blocks (or even individually at an exorbitant markup, says the guy who bought one to reprint a single book), so this doesn't show anything really interesting/useful. A visualization using LoC or even Dewey Decimal would be far more useful, esp. if it also linked to public domain and copyright-free repositories/lists, say an interactive and visual version…

ISBN's are hierarchical, what do you mean? Like Gaul, ISBNs are divided into multiple parts, where one part is for the language, another is for the publisher, and the last is for the title. The last part is a checksum. https://en.wikipedia.org/wiki/ISBN#Overview

Yes, but this internal hierarchy for an issued number doesn't tell anything beyond those facts about a specific edition of a specific text.

One can't use ISBNs alone to create a hierarchical listing of texts which is useful for anything beyond browsing by language/publisher/order in which the ISBN was generated.

A visual and interactive representation of books by LoC or some other cataloging system would actually be useful.

Re: Visualizing All ISBNs

#82

Earlier quoted context omitted.

I'm also in NL. Ziggo's DNS server blocks it: $ dig annas-archive.org @89.101.251.228 annas-archive.org. 360 IN CNAME unavailable.for.legal.reasons. unavailable.for.legal.reasons. 339 IN A 213.46.185.10 213.46.185.10 serves a generic page mentioning Russia Today and the Pirate Bay. Not sure which one applies here.

Same for KPN: http://195.121.82.125/ Would Tweak have blocked this? Most households in the Netherlands currently have the choice of Ziggo, KPN, and Odido. Long live VPNs…

Is that three broadband providers serving the same address?? You guys are so lucky you don’t even know. In America we generally have a choice of one if you aren’t including Starlink or legacy slow satellite. And perhaps a joke of a 1-6Mbps DSL option in some parts.

Re: Visualizing All ISBNs

#83
post #76

Earlier quoted context omitted.

They can't even have a tiny fraction of the world's books. Each edition of the book gets a new ISBN... if a book is released as a paperback, hardback, kindle edition, pdf, and epub then there are supposed to be five ISBNs. The vast, vast majority have only been released as dead-tree versions. They have none of those. The books they scan may have an ISBN, but the scans do not have them. Like all Project Gutenberg book…

> The books they scan may have an ISBN, but the scans do not have them. Like all Project Gutenberg books, their books have no ISBNs at all. From a strict point of view, they've released new editions of these books. Are you saying they actively remove ISBN numbers from scans? If I downloaded one of the books, it wouldn't have an ISBN? Why? That seems like a bunch of extra processing per book, makes it harder for users…

> Are you saying they actively remove ISBN numbers from scans?

No, he‘s playing the pointless „well, actually a scan of a book is a different thing from the book itself“ game.

Re: Visualizing All ISBNs

#84
post #36

Earlier quoted context omitted.

You mention the example of romance novels above. There's a schlocky Victorian pulp novel that's of no use to anyone - except that it happens to contain a fantastically detailed description of an abandoned saltings in my hometown that nobody ever thought to record in any way. For me, those two paragraphs are gold. If the novel hadn't been digitised as part of Google's Books Archive Project, I wouldn't have been able t…

Well I guess your one valuable paragraph that matters only to you justifies backing up millions (billions?) of human and soon to be AI generated books, because someone, somewhere, at some time will find a line or two valuable. Maybe. I retract my position, let's back up everything!

Let’s say there are ten billion such marginally-useful books published by the time the next few decades. Many epub books are like a couple MB. So 30 petabytes total. That’s something you could fit in one room. One rich guy could buy enough hard drives to do that today. Why not?

Re: Visualizing All ISBNs

#85
post #80

Earlier quoted context omitted.

> All action is "subjective and subject to biases that change over time". This is poppycock. Backing up all books -- the very action discussed by the person you're answering -- is by definition neither subjective nor subject to biases. > This would then imply I could never take any action, because it's just subjective and biased. And even if the first quoted claim were true, this, too, clearly isn't. Nowhere does the…

> least constructive comment I've seen on HN in quite some time But we should still archive it. Some day it might be useful to someone ;)

Exactly :-)

Re: Visualizing All ISBNs

#86
post #78

It appears that the IP of the server is blocked in the EU. I get this from my ISP (Ziggo, in the Netherlands): Deze website is geblokkeerd Europese sancties De Raad van Europa heeft besloten dat de websites van RT (voorheen Russia Today) en Sputnik News niet meer mogen worden doorgegeven. De website die je probeert te bezoeken, valt onder deze Europese sanctie. VodafoneZiggo is verplicht de sanctie uit te voeren en h…

No issue here in France.

Re: Visualizing All ISBNs

#87
post #74

Earlier quoted context omitted.

ISBN's are hierarchical, what do you mean? Like Gaul, ISBNs are divided into multiple parts, where one part is for the language, another is for the publisher, and the last is for the title. The last part is a checksum. https://en.wikipedia.org/wiki/ISBN#Overview

Yes, but this internal hierarchy for an issued number doesn't tell anything beyond those facts about a specific edition of a specific text. One can't use ISBNs alone to create a hierarchical listing of texts which is useful for anything beyond browsing by language/publisher/order in which the ISBN was generated. A visual and interactive representation of books by LoC or some other cataloging system would actually be…

totally agree, but thats not in the data. however, since blocks are assigned to agencies associated with countries and publishers, you might find some utility in showing coverage by likely language and/or country of origin and date.

Re: Visualizing All ISBNs

#88
post #74

Earlier quoted context omitted.

ISBN's are hierarchical, what do you mean? Like Gaul, ISBNs are divided into multiple parts, where one part is for the language, another is for the publisher, and the last is for the title. The last part is a checksum. https://en.wikipedia.org/wiki/ISBN#Overview

Yes, but this internal hierarchy for an issued number doesn't tell anything beyond those facts about a specific edition of a specific text. One can't use ISBNs alone to create a hierarchical listing of texts which is useful for anything beyond browsing by language/publisher/order in which the ISBN was generated. A visual and interactive representation of books by LoC or some other cataloging system would actually be…

I got into an argument with the manager of South End Press back in '94 about whether 'Futuresplash' (soon to be Macromedia Flash) had a future, he thought it did and he was right.

Years later I was working at the library and got a little bit steamed because South End Press was reusing ISBN's after books went out of print which was allowed but, I think, lame.

One of my strategies for researching a topic is looking a few up in the OPAC, finding them in the stacks, and finding more books on the topic in those areas. (In the Library of Congress system, machine vision could be under QA56 with the rest of computer science or around TA1630, thus "areas".)

From time to time I've thought about trying to replicate the feel of this with some kind of UI given that our library moved a lot of the collection into deep archives and we have a very fast 'Borrow Direct' service with other peers)

Re: Visualizing All ISBNs

#89
Anna's archive is one of the wonders of the world. If we almost destroyed our species but Anna's archive endured, there would be hope for a relatively expedient reconstruction.

Re: Visualizing All ISBNs

#90
I see that bounty at the bottom, so tossing away my chances here, but this visualization is just asking to be mapped onto a Hilbert Curve. [0] When you "stripe" the data like this, points that are sorted close together could end up pretty far apart, since a distance in the Y axis skips an entire row of data as you move down, rather than a distance in the X axis which is 1-to-1 with the source data.

If you map it onto a hilbert curve, the X and Y axis mean nothing, but visually points that are close together in the sorted list, will be visually close together in the output image.

Since the first part of an ISBN is the country, then the second part is the publisher, and the third part is the title, with a check sum at the end, I would remove the checksum and sort them each as a big number. (no hyphens)

You should end up with "islands", where you see big areas covered by big publishing countries, with these "islands" having bright spots for the publisher codes.

Bonus points for labeling these areas!

I set up something a while ago [1] for an interview that does this with weather data. It makes the seasons really obvious since they're all grouped together.

[0] https://en.wikipedia.org/wiki/Hilbert_curve

[1] https://graypegg.com/hilbert (https://github.com/graypegg/hilbertcurveplayground code if anyone wants to go for the prize using this! Please at least mention me if you decide to reuse this code, but I can't stop ya lol)

Post reply on HN