Live data from Hacker News

Visualizing Tolkien

5013.es

41–50 of 55 posts

Re: Visualizing Tolkien

#41
post #37

Earlier quoted context omitted.

Hi author, I really liked the visualisations, especially the black hole thing. I think I may have misunderstood the purpose of your article: I first read it as a serious attempt to use textual analysis to do a comparison of the comprehensibility of Tolkein's best-known works, in which case you fell way short of the mark by going no further than word counting. I think it started off like that, but on re-reading it I s…

Isn't FK determined almost completely by sentence length? That is what I recall from messing with MS Word docs in high school.

FK combines average sentence length (total words/total sentences), average syllables per word and some fixed coefficients to come up with an equivalent school grade.

Flesch Reading Ease score, which is what I actually meant, does the same but with different coefficients to come up with a more granular difficulty score, usually in the range of 30-100.

They're both pretty arbitrary. The more I read up on this subject the more respect I have for the author's own attempts at an originality score. It's all subjective ultimately.

Re: Visualizing Tolkien

#42
post #15

>One of the reviewers argued that it was the hardest book to read because 'and' was the most used word in the book. I think the author took this way to literally. Granted, it inspired some fun (albeit odd and basic) data analysis. But the point of the reviewer isn't about word counts but about style and pace. That, narratively, the Silmarillion just felt like it continued on and on without clear sections of rising an…

> Sentences were long and atmospheric, rather than short, quick and active.

Interestingly, I think the larger frequency of 'of' actually suggests this. I suspect the Silmarillion (I haven't read it in ... a decade maybe?) has a far more complicated web of relationships–not just between characters, but between places; at the very least, it concerns itself with genealogies more.

Re: Visualizing Tolkien

#43
post #6

Why does everyone find the Silmarillion hard to read? Am I the only one, who had no problems with that style?

From what I remember the first couple of chapters contain very close to 0 dialogue, most people simply aren't used to reading books like that.

That depends. Those who like reading folklore and legends actually appreciate such style.

Re: Visualizing Tolkien

#44
post #6

Why does everyone find the Silmarillion hard to read? Am I the only one, who had no problems with that style?

I like it too. So not everyone finds it hard to read. But it's not some casual fantasy for sure.

Re: Visualizing Tolkien

#45
post #10

Earlier quoted context omitted.

Hi, author here. First, it's she, not he. Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scientific analysis but rather answer to the questions that came to my mind, by using a computer to validate hypothesis. But many thanks for the pointers & suggestions, though! I was thinking about rebuilding this to make it realtime+interactive so that more than on…

The "originality" index bothered me, because as the work of a length grows, you'd generically expect less words to be introduced -- exactly what the results show. The idea makes sense, but I'm unclear on how to actually measure it in a way that's normalized by page count.

There's the concept of a vocabulary growth curve. It shows the number of words occurring once as a function of the amount of text, e.g., how much new words in the first 1000 words? how much in first 2000 words? etc.

Re: Visualizing Tolkien

#46
post #10

Earlier quoted context omitted.

Hi, author here. First, it's she, not he. Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scientific analysis but rather answer to the questions that came to my mind, by using a computer to validate hypothesis. But many thanks for the pointers & suggestions, though! I was thinking about rebuilding this to make it realtime+interactive so that more than on…

Hi author, I really liked the visualisations, especially the black hole thing. I think I may have misunderstood the purpose of your article: I first read it as a serious attempt to use textual analysis to do a comparison of the comprehensibility of Tolkein's best-known works, in which case you fell way short of the mark by going no further than word counting. I think it started off like that, but on re-reading it I s…

> On he/she

While I consider defaulting to 'he' to be legitimate and acceptable, I actually prefer the zie/zir gender-neutral pronouns when I think about it. http://santiago.mapache.org/nonfiction/essays/zie.html If I ever begin to agonize about the gender of the person I'm talking about, that's enough to kick me over into using GNPs.

Re: Visualizing Tolkien

#47

Earlier quoted context omitted.

The "originality" index bothered me, because as the work of a length grows, you'd generically expect less words to be introduced -- exactly what the results show. The idea makes sense, but I'm unclear on how to actually measure it in a way that's normalized by page count.

There's the concept of a vocabulary growth curve. It shows the number of words occurring once as a function of the amount of text, e.g., how much new words in the first 1000 words? how much in first 2000 words? etc.

But is there a good way to boil that down to a single metric? (i.e., can we parametrize these curves with a single value?)

Re: Visualizing Tolkien

#48
post #35
post #10

Earlier quoted context omitted.

Hi, author here. First, it's she, not he. Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scientific analysis but rather answer to the questions that came to my mind, by using a computer to validate hypothesis. But many thanks for the pointers & suggestions, though! I was thinking about rebuilding this to make it realtime+interactive so that more than on…

Nice work! So my question is where did you get the text to analyze?

You can find the text for pretty much any popular work if you look in the proper places :-)

Re: Visualizing Tolkien

#49

Earlier quoted context omitted.

There's the concept of a vocabulary growth curve. It shows the number of words occurring once as a function of the amount of text, e.g., how much new words in the first 1000 words? how much in first 2000 words? etc.

But is there a good way to boil that down to a single metric? (i.e., can we parametrize these curves with a single value?)

I wouldn't know off the top of my head, but the concept was introduced by Harold Baayen. There are free PDFs online of his with lots of interesting quantitative techniques.

Re: Visualizing Tolkien

#50
"I wrote a simple program who counted how many times did each word appear in the [The Silmarillion]."

Out of curiosity, where did the author get a copy of the text to analyze?

"Turin" is the only name in The Silmarillion graph? Interesting.

Post reply on HN