Live data from Hacker News

Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

pgdp.net

101–110 of 111 posts

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#101
post #93
post #92

Earlier quoted context omitted.

The young woman I was standing behind at the bus stop tonight had the Deathly Hallows logo tattooed on her harm. What does Harry Potter represent? You can't replace a book-length narrative with a short phrase.

Well, possibly I miss something, because I wouldn't recognize a Deathly Hallows logo if I see one. (I assume this is rather from the movies?). But there are occasional references to Harry Potter I have seen. Apart from that, I would say Harry Porter represents some things. The glass eyed bullied nerd, that steps into a magical realm to become a superhero. In general, the concept of a fantastic magical realm hidden be…

It's not the main thrust; it's the details.

I think you're right that the Deathly Hallows logo was introduced by the films.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#102

Earlier quoted context omitted.

You don't need to hash file contents (though that is often a useful thing to do). You can hash e.g. the URL that was earlier claimed to be the canonical identifier. Running it through your favorite hash function fixes your complaints about file names (choose your favorite hash function such that it is not too long and only outputs allowed characters).

Ah. The url, so I can substitute one difficult-for-human-readability with another difficult-for-human-readability, both of which are excessively long and opaque-by-design. >choose your favorite hash function such that it is not too long ISBN's 13 digits is about as long as is tolerable. Any time there is a list of authors six names long (academic titles) along with a subtitle, it's very easy to bump up against max fi…

Hashes are not excessively long unless you choose to make it so. They might be opaque/random if you want, or they might not. "Remove all special characters and keep only the first 5 characters with space padding" is a string hash function. "Keep only the first 5 vowels with space padding" is a string hash function.

Here's a friendly AI generated hash function to give you an opaque 13 digit number if you're into that:

echo -n "$URL" | sha1sum | awk '{print $1}' | xxd -r -p | od -An -t u8 | tr -d ' \n' | cut -c1-13

For example, for https://standardebooks.org/ebooks/denis-diderot/the-indiscre... you get the ID 4897562473051.

It looks like their ebook sources are all published in git repos online, so you could check out the repos, get the timestamp of the initial commits, and do a monotonic ID on that if you wanted. You could also contribute the change back to them if you think it's something others would benefit from.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#103
post #50

I'm shocked and saddened to hear this. Greg was a deep source of knowledge and support as I started and shepherded Standard Ebooks. He was generous with his time and experience, and unbelievably patient with me, some guy he had never heard of or met before who was just another cold-email in what must have been an endless stream in his inbox. We should all aspire to his high spirit of camaraderie, charity, and kindnes…

Why are there no unique numbers assigned to Standard Ebook's ebooks? I understand that there is a cost associated with ISBNs, but it's very irritating to not have something that identifies them uniquely. Most (all?) aren't even in Worldcat, so I can't use OCLC numbers for that purpose either.

> no unique numbers

This suggests a misunderstanding of the Standard Ebooks process, which allows continual incremental corrections to the authoritative source of individual books (in XHTML, on GitHub). So, a truly unique identifier would only be valid to the production output(s) from a particular state of the Git-repo sources.

https://standardebooks.org/contribute/report-errors

Recall also that final user content is made available in multiple formats, currently at least six. Example:

https://standardebooks.org/ebooks/geronimo/geronimos-story-o...

Asynchronous to the correction process, Standard Ebooks updates its own production tools. So if an individual book's content requires correction, should the "respin" be done with TOT tools, or with the versions available at time of first publication? Disclaimer: I don't actually know which is current practice -- but using the TOT tool suite is obviously vastly easier.

For most practical purposes, I'd suggest the git-commit date, along with short substrings of author name and title, would suffice.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#104
post #42
post #36

Earlier quoted context omitted.

Can you change it back? Currently it says "Greg Newby, CEO of Project Gutenberg, has died," which seems to be misleading people.

I can’t, sorry, there’s no edit button, I guess because of low karma. Maybe the mods (or who changed it) can revert it to the old title.

NB: You can email mods with such requests after the edit window has expired, or for submissions by others.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#105
post #60
post #59

Earlier quoted context omitted.

No comment on the rest of the list, but anything by H. L. Mencken is likely to be a banger.

Yeah, it's probably more enjoyable than the rest of that list put together. Still, prefaces ? Generally I try to avoid reading those.

Heh. That's why I'm intrigued by it! My bet is that it's some kind of satire based on that (almost universal) avoidance. Regardless, Mencken's wicked clever, and less-known than I think he deserves to be.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#106
post #92
post #82

Earlier quoted context omitted.

" Harry Potter, like Rambo, The Matrix, and Frankenstein, supplies metaphors and narratives through which nearly everyone today interprets the world around them, even if they haven't read it themselves" I did read the books, but I don't think I have really encountered the use of "Muggles, horcrux, mudblood" in every day life, nor do I personally feel they shaped my metaphors or narratives on how I see the world. Fran…

The young woman I was standing behind at the bus stop tonight had the Deathly Hallows logo tattooed on her harm. What does Harry Potter represent? You can't replace a book-length narrative with a short phrase.

On her arm!

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#107
post #105
post #60

Earlier quoted context omitted.

Yeah, it's probably more enjoyable than the rest of that list put together. Still, prefaces ? Generally I try to avoid reading those.

Heh. That's why I'm intrigued by it! My bet is that it's some kind of satire based on that (almost universal) avoidance. Regardless, Mencken's wicked clever, and less-known than I think he deserves to be.

Comedy is often somewhat dependent on shared cultural assumptions, although that Sumerian fart joke is still kind of funny.

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#108
post #35

Earlier quoted context omitted.

I'm at high risk for colon cancer & had my first screening last year after putting it off. For those who've not done it yet: it's really no big deal. The most challenging part is drinking the fluids. Please get screened: caught early, it's easily curable, but it definitely kills.

None of the cancer screening programs have been able to demonstrate an effect on life expectancy, neither colonoscopies nor mammographies.

From what I’ve seen, studies haven’t done a good job of actually testing for testing lifetime expectancy.

https://pmc.ncbi.nlm.nih.gov/articles/PMC11047044/

“In conclusion, the statement that cancer screenings do not save lives cannot be properly drawn from the Bretthauer's et al. meta-analysis because lifetime gains are likely underestimated and based on uncertain all-cause mortality estimates.

Lifetime gains estimated for the screened group from all-cause mortality reduction is a misleading measure and should be avoided because it implies a benefit for all persons in the screening group, including those not affected by the target cancer.”

https://bmchealthservres.biomedcentral.com/articles/10.1186/...

“Although gaps persist between the full potential benefit and benefits considering adherence, existing cancer screening technologies have offered significant value to the US population. Technologies and policy interventions that can improve adherence and/or expand the number of cancer types tested will provide significantly more value and save significantly more patient lives.”

Re: Greg Newby, CEO of Project Gutenberg Literary Archive Foundation, has died

#109

Earlier quoted context omitted.

Why are there no unique numbers assigned to Standard Ebook's ebooks? I understand that there is a cost associated with ISBNs, but it's very irritating to not have something that identifies them uniquely. Most (all?) aren't even in Worldcat, so I can't use OCLC numbers for that purpose either.

> no unique numbers This suggests a misunderstanding of the Standard Ebooks process, which allows continual incremental corrections to the authoritative source of individual books (in XHTML, on GitHub). So, a truly unique identifier would only be valid to the production output(s) from a particular state of the Git-repo sources. https://standardebooks.org/contribute/report-errors Recall also that final user content is…

>This suggests a misunderstanding of the Standard Ebooks process, which allows continual incremental corrections to the authoritative source of individual books (in XHTML, on GitHub). So, a truly unique identifier would only be valid to the production output(s) from a particular state of the Git-repo sources.

Well, one of us has a misunderstanding. Just because the printer strikes off the printing number from the colophon for each subsequent printing, they don't actually issue a new ISBN. That stays the same. If they wanted to also include a version number too, I wouldn't mind that as well, but it's not nearly as necessary as this. I use the year as a rough version number in the file names as well.

>Recall also that final user content is made available in multiple formats, currently at least six. Example:

I don't need them to issue a number per file format, but if they want to... that doesn't bother me. That's sort of self-evident which of the formats it is, after all.

>I'd suggest the git-commit date, along with short substrings of author name and title, would suffice.

It doesn't. A number of authors have at one time or another have released books with similar or identical titles that are not the same book. This is the trouble... someone who uses or would use the books is asking for something that is missing but easy to supply, and instead of a "well gee, we never considered that, let us think about it" I have a dozen assholes crawling out of the woodwork to say "no, you're doing it wrong".

I need unique identifiers that are human readable. I just do. The world discovered this need for books before you were born. They invented a global standard, even. There is an entire field of science out there about this, that you seem to be ignorant of even existing. I've been doing this for years, and I keep bumping up against it. But you think it can be solved because you used git and know about hashes or whatever, and it's just like what you deal with in your software development job!

Post reply on HN