Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

51–60 of 87 posts

Re: Language homogenization at Harvard

#51
post #7
post #4

I'd like to see a study of vocabulary with the number of unique words we have used over time. My sense is that older prose has a much richer and more varied vocabulary. The studies could be showing that we use a smaller dictionary today.

I had a similar intuition and graphed unique words per year while writing this. I found, surprisingly, that actually the reverse was true. Unique words per year go up, even as diversity goes down. Another finding that may explain this is that the articles get longer as time goes on - so a simple unique word count may just increase as a function of the authors using more words. There is a period in the late 80's to ea…

Maybe you should just crop the same number of words from the start of each article, and take the same number of articles from each year by random sampling. That would make things easier to compare.

Re: Language homogenization at Harvard

#52
post #48

2 counter points: 1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying! 2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934…

Quick, someone make a 1934 version of Wordle!

Re: Language homogenization at Harvard

#53

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

For instance, the addition of boilerplate to every document would increase their similarity metric. This has certainly happened over that time period with required compliance statements.

Re: Language homogenization at Harvard

#54

This article is almost criminally flawed. The author makes some horrible assumptions and presents shoddy data. Even if we assume that the latent space model is "correct" (we shouldn't), they don't present anything like variance or number of samples in a year. Then they sort of arbitrarily fit a line to data which pretty clearly looks non-linear. Suppose, for example, that The Crimson (not Harvard btw, it's a student…

For the title, I did say "At Harvard" and not "In Harvard" or "By Harvard". I think the student newspaper is indeed at Harvard. I'm also pretty clear about what text I'm looking at in the article. In the title I used "Harvard" over "The Crimson" because I figured fewer people would know "The Crimson" compared to "Harvard". Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm c…

Thank you for your article! Anecdotally, this is a phenomenon I see in high pressure service companies, like McKinsey’s of this world: there’s a very restricted, idiosyncratic vocabulary used by people working there driven by the idea that it will promote sales.

Re: Language homogenization at Harvard

#55
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

It's interesting how your histogram and candle chart shows that, while the mean has shifted, there's a sliver of samples with greater cosine distance than anything previously recorded in the dataset. So I guess while the system has become more inclusive of boilerplate language, it's also become more inclusive of far-out novelties. I'd be interested in reading those abstracts.

Re: Language homogenization at Harvard

#56
post #48

2 counter points: 1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying! 2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934…

I'm surprised you were unfamiliar with words like "mauve", "wanton" or "sallow". These are fairly common in today's UK, but maybe less so in the US vernacular?

Jocosely is the only one I had a question over, until I realised it had the root "jocose".

Re: Language homogenization at Harvard

#57

I would guess that the increased rate of information transmission across society has contributed to this trend. Copying is a key mechanism in how culture develops and propagates, and information moving more easily makes it easier for different groups to copy the dominant examples for any particular activity or discipline. A lot of information systems and cultural phenomenons have winner-take-all / power-law dynamics.…

Tangentially related, but I was thinking about this while watching the Olympics men's figure skating competition recently (given the ongoing drama on the women's side, I'll leave that aside for now). The quality of skating has drastically improved in just a decade. In 2010 the Olympic champion had no quad jumps. In 2022 the easiest jump from the top 3 competitors in their short program was a triple axel - the rest we…

They trained in 2010 just as hard as now. And learned from others for decades. I guarantee you they wanted to win as much and were trying new jumps.

What you see there is what you see in any sport - evolution of it toward better performance.

Re: Language homogenization at Harvard

#58
post #56
post #48

2 counter points: 1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying! 2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934…

I'm surprised you were unfamiliar with words like "mauve", "wanton" or "sallow". These are fairly common in today's UK, but maybe less so in the US vernacular? Jocosely is the only one I had a question over, until I realised it had the root "jocose".

It's because British people have such bad taste they still have mauve stuff. In French, the language it comes from, this color is synonymous with ugliness and the 60s hehe

Re: Language homogenization at Harvard

#59

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

> Lexical homogenization is not the same as idea homogenization

Actually it IS evidence for idea homogenization. Words represent ideas. Unless you are claiming that the same word is used to represent different ideas, which would be even more confusing than having different words representing the same idea. Therefore fewer unique words => fewer unique ideas.

Re: Language homogenization at Harvard

#60
post #21

There's no evidence from this article that the idea space is decreasing. Vocabulary is becoming somewhat more similar: the measurement is on individual words. Suppose an institution decides that the faculty is using overly obscure language (or a grant-making agency decides this) and asks that grant submitters reduce the amount of jargon, include more background, etc. Wouldn't this measurement then show a significant…

Words represent ideas. Fewer words is good evidence that there are fewer ideas.

Sometimes multiple words represent the same idea, so in that case homogenization is "good". It is also possible that multiple ideas are represented by the same word, but that is "bad" in an academic context because it leads to ambiguity. Therefore fewer unique words IS actually good evidence that the idea space is decreasing, unless you are saying academic literature is becoming more and more ambiguous.

Post reply on HN