Live data from Hacker News

Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

britannica11.org

51–60 of 138 posts

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#51

You can discover beliefs that are shocking today, such as this excerpt from the article "Adolescence": "In the case of girls, let them run, leap and climb with their brothers for the first twelve years or so of life. But as puberty approaches, with all the change, stress and strain dependent thereon, their lives should be appropriately modified. Rest should be enforced during the menstrual periods of these earlier ye…

You can nowadays paste the text from pretty much anything that's in the public domain into a near-SOTA LLM such as Kimi or GLM and it will give you a pretty nice summary of what it's about in modern language (Extremely useful: the LLM tendency to go overboard on formatting nicely balances out the wall-of-text format from historical publications, which was aimed at saving paper and minimizing manual layout effort), an…

I beg of thee, use that brain of yours and read a text that was made scarcely more than a century ago, a blink of an eye in the grand scale of the changes of the linguistic features of English, and interpret it for yourself.

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#52
post #44

Earlier quoted context omitted.

I feel exactly the same way about encyclopedias and dictionaries. And Encarta really was amazing. You'd be surprised how much modern criticism of the 11th amounts to "no entry on the Great War", except in earnest.

Thanks a lot for this incredible gem! By the way, it looks like there's a bug where I can't search for articles when already inside one. To do so, I need to go back to home > articles and then search.

If you're reading an article, just go to the top and type in the left-hand search box. That will search for articles as well as text within articles. The right-hand box searches the text of the article you're reading.

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#53
post #2

I rebuilt the 1911 Encyclopædia Britannica into a clean, structured, navigable site: https://britannica11.org/ What it does: – ~37k articles reconstructed from the original volumes – section-level structure (contents are clickable within articles) – cross-references extracted and linked – contributors indexed and searchable – original volume + page references preserved and shown while reading – links to the original…

You might want to add The Reader's Guide to the Encyclopaedia Britannica, PD text available at https://www.gutenberg.org/ebooks/74039 and scans at https://archive.org/details/readersguidetoen00londuoft - It would fit naturally with the Ancillary material that includes the topic-based index.

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#54

You can discover beliefs that are shocking today, such as this excerpt from the article "Adolescence": "In the case of girls, let them run, leap and climb with their brothers for the first twelve years or so of life. But as puberty approaches, with all the change, stress and strain dependent thereon, their lives should be appropriately modified. Rest should be enforced during the menstrual periods of these earlier ye…

You can nowadays paste the text from pretty much anything that's in the public domain into a near-SOTA LLM such as Kimi or GLM and it will give you a pretty nice summary of what it's about in modern language (Extremely useful: the LLM tendency to go overboard on formatting nicely balances out the wall-of-text format from historical publications, which was aimed at saving paper and minimizing manual layout effort), an…

How is that not "modern language"?

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#55

Small world - I'm currently cleaning up scans of the EB 9th edition to put it online as a mediawiki site; I'm including all the illustrations and plates so I'm only a third of the way through. I've been testing different OCR tools and so far I've been the most impressed with paddleOCR - it correctly split the text columns, labled the illustrations, and noted the maragin text. Still, it's not perfect, so I'm having to…

For those unfamiliar, the 1875 9th ed. was known as the scholar's edition due to how many eminent persons had contributed; it's a fascinating snapshot of the late 1800s.

Other material that would be fun to put online in a hyperlinked and indexed format include geographic and medical atlases and the Baedeker travel guides.

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#56

Very, very cool. Hats off. I've considered attempting a more limited form of this for years. For those who don't know, the 1911 Britannica is heralded for several reasons (and rightly criticized for regrettable others), but the most well-known is that it was the last encyclopedia before The Great War, and hence had a good amount of steam/optimism coming from the first and second industrial revolutions and the "Progre…

> But I would love an option (emphasis on option) to see the text side by side with the page images. ... That way, I could "confirm" or "fact check" the faithfulness of the OCR.

You can already do that on Wikisource. For example, here's p. 658 from the entry on "Molecule":

https://en.wikisource.org/wiki/Page:EB1911_-_Volume_18.djvu/...

Also OP: I noticed some fidelity issues in your version (at https://britannica11.org/article/18-0684-s2/molecule). For example parts of the math formula under the line that ends with "the molecules of other kinds" ([1]) are missing (compare [2]). Also, in your version fn. 1 of this article is attached to "as they have always done" ([3]) but it should actually be attached to "Atom" on p. 654 ([4]):

[1] https://britannica11.org/article/18-0684-s2/molecule#:~:text...

[2] https://en.wikisource.org/wiki/Page:EB1911_-_Volume_18.djvu/...

[3] https://britannica11.org/article/18-0684-s2/molecule#:~:text...

[4] https://en.wikisource.org/wiki/Page:EB1911_-_Volume_18.djvu/...

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#57

Small world - I'm currently cleaning up scans of the EB 9th edition to put it online as a mediawiki site; I'm including all the illustrations and plates so I'm only a third of the way through. I've been testing different OCR tools and so far I've been the most impressed with paddleOCR - it correctly split the text columns, labled the illustrations, and noted the maragin text. Still, it's not perfect, so I'm having to…

I'm looking forward to it. The 9th is great in its own right and a lot of it is in the 11th. Alfred Newton's nearly 200 articles on bird species and a few classic essays by Macaulay come to mind offhand.

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#58

I've been meaning to build ~exactly this experience, but for the 1952 Encyclopedia Brittanica Great Books of the World collection and its experimental index Syntopicon [0]. Would love to know more about how you OCR'd or otherwise ingested and parsed the raw material. I have a physical copy of the books, and I found some samizdat raw-image scans and started working on a custom OCR pipeline, but wondering if maybe I co…

That collection is not in the public domain, AIUI? You might be able to do it for the Harvard Classics, which has a nice collection-wide index of terms. https://en.wikisource.org/wiki/The_Harvard_Classics has links to the scans.

Oh no not in the public domain, I better not build something cool!

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#59
post #42

I've been meaning to build ~exactly this experience, but for the 1952 Encyclopedia Brittanica Great Books of the World collection and its experimental index Syntopicon [0]. Would love to know more about how you OCR'd or otherwise ingested and parsed the raw material. I have a physical copy of the books, and I found some samizdat raw-image scans and started working on a custom OCR pipeline, but wondering if maybe I co…

I'm familiar with the Synopticon, which would be fun to structure. I didn’t do OCR myself, except for the topic index and to fill in a few gaps. I started from existing Wikisource text and then built a pipeline around that: cleaning (headers, hyphenation, etc.), detecting article boundaries, reconstructing sections, and linking things back to the original page images. Most of the effort went into rendering the comple…

Ah ok thanks very much!

Re: Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica

#60

Earlier quoted context omitted.

You can nowadays paste the text from pretty much anything that's in the public domain into a near-SOTA LLM such as Kimi or GLM and it will give you a pretty nice summary of what it's about in modern language (Extremely useful: the LLM tendency to go overboard on formatting nicely balances out the wall-of-text format from historical publications, which was aimed at saving paper and minimizing manual layout effort), an…

You didn't really explain what that does for you. Why do you paste it into an LLM?

I'm not sure if you're familiar with public domain texts from around the 19th or early 20th century, but they were not intended to be skimmed or speed-read the way we'd skim a modern text prior to getting into a more attentive close-reading. Even their short magazine articles were actually the near-equivalent to our scholarly papers, and were often read aloud at length in parlor gatherings. So having a LLM split the text into manageable sections for you and provide a hint of what each lengthy wall-of-text paragraph will be about is actually a huge gain in readability.
Post reply on HN