Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

31–40 of 87 posts

Re: Language homogenization at Harvard

#31
I would interested if grant proposals which do _not_ get funded, show less of language homogenization than ones that do. Although, it still wouldn't tell you if the idea space is being constricted; that seems like a much more difficult thing to measure. It could just be the equivalent of everyone at court speaking the way the king does, because he's the king. Winning grants probably get read, and imitated, more than other grant proposals that don't get funded. Which might also happen with ideas, but I'm not convinced that is what is measured here.

Re: Language homogenization at Harvard

#32
This article is almost criminally flawed. The author makes some horrible assumptions and presents shoddy data. Even if we assume that the latent space model is "correct" (we shouldn't), they don't present anything like variance or number of samples in a year. Then they sort of arbitrarily fit a line to data which pretty clearly looks non-linear.

Suppose, for example, that The Crimson (not Harvard btw, it's a student newspaper) runs 10x as many articles this year as last. It's possible you're going to get a huge reduction in cosine distance just by virtue of a few authors producing a lot more content.

At a minimum, we need mean, variance, and number of samples. This doesn't tell you anything about "Harvard", it just tells you about the students who the Crimson choose to publish. There are lots of structural reasons within Harvard that the Crimson has probably stopped being a unified voice of the student body -- but again, that's not what the article purports to show.

Re: Language homogenization at Harvard

#33
post #7
post #4

I'd like to see a study of vocabulary with the number of unique words we have used over time. My sense is that older prose has a much richer and more varied vocabulary. The studies could be showing that we use a smaller dictionary today.

I had a similar intuition and graphed unique words per year while writing this. I found, surprisingly, that actually the reverse was true. Unique words per year go up, even as diversity goes down. Another finding that may explain this is that the articles get longer as time goes on - so a simple unique word count may just increase as a function of the authors using more words. There is a period in the late 80's to ea…

> Unique words per year go up, even as diversity goes down. Another finding that may explain this is that the articles get longer as time goes on - so a simple unique word count may just increase as a function of the authors using more words.

You'd probably see the same effect if the articles don't get any longer, but more of them get written every year. Unique word count will always go up with words produced, even if most words produced are formulaic boilerplate.

Re: Language homogenization at Harvard

#34

I would guess that the increased rate of information transmission across society has contributed to this trend. Copying is a key mechanism in how culture develops and propagates, and information moving more easily makes it easier for different groups to copy the dominant examples for any particular activity or discipline. A lot of information systems and cultural phenomenons have winner-take-all / power-law dynamics.…

Tangentially related, but I was thinking about this while watching the Olympics men's figure skating competition recently (given the ongoing drama on the women's side, I'll leave that aside for now).

The quality of skating has drastically improved in just a decade. In 2010 the Olympic champion had no quad jumps. In 2022 the easiest jump from the top 3 competitors in their short program was a triple axel - the rest were all quads.

My theory about this is that the advent of smartphones and ubiquitous video has made it much easier to constantly see what your competitors are doing. This not only pushes you to work harder and try new jumps, but also lets you see what new techniques work for other athletes. A couple decades ago people wondered if a quad jump was even possible, now people train going into it that a triple is no longer the expected limit.

Re: Language homogenization at Harvard

#35
post #7

Earlier quoted context omitted.

I had a similar intuition and graphed unique words per year while writing this. I found, surprisingly, that actually the reverse was true. Unique words per year go up, even as diversity goes down. Another finding that may explain this is that the articles get longer as time goes on - so a simple unique word count may just increase as a function of the authors using more words. There is a period in the late 80's to ea…

> Unique words per year go up, even as diversity goes down. Another finding that may explain this is that the articles get longer as time goes on - so a simple unique word count may just increase as a function of the authors using more words. You'd probably see the same effect if the articles don't get any longer, but more of them get written every year. Unique word count will always go up with words produced, even i…

One nice thing about getting feedback is learning all of the additional stuff I should have included in the blog post. I did look at number of articles per year, and it fluctuates, but there isn't a huge change across the century, and the change goes up and down. Total words, on the other hand, does trend up and goes up faster more recently.

Re: Language homogenization at Harvard

#36

Noam Chomsky on suicide watch In my point of view this is actually good news; Since I'm not a native english speaker and since this language has become the world's lingua franca it's very nice being able to understand an article just because the author's didn't waste time looking for bombastic synonyms just to sound smarter. I applaud this and projects like Simple English Wikipedia which allows us, non-native english…

> bombastic Esoteric synonyms, perhaps? :-)

I think he meant sesquipedalian synonyms.

Re: Language homogenization at Harvard

#37
As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past.

This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at random. Let's further assume in the "bad new days" people use those old ideas, change/improve upon them just slightly ("add noise"), but some old ideas are less often picked up than others. Say for example the least attractive old proposal is picked up twice, whereas the most attractive old idea is picked up and changed in 100 new proposals. Then, there is a lower average distance between proposals, all the while the total range of ideas has increased.

Because it's Friday and I am waiting for my oven to finish cooking my food, I wrote a small simulation. It's probably full of mistakes and I may have made terrible mistakes in my assumptions, but I thought it's fun:

https://imgur.com/wu232rP

Re: Language homogenization at Harvard

#38

This article is almost criminally flawed. The author makes some horrible assumptions and presents shoddy data. Even if we assume that the latent space model is "correct" (we shouldn't), they don't present anything like variance or number of samples in a year. Then they sort of arbitrarily fit a line to data which pretty clearly looks non-linear. Suppose, for example, that The Crimson (not Harvard btw, it's a student…

For the title, I did say "At Harvard" and not "In Harvard" or "By Harvard". I think the student newspaper is indeed at Harvard. I'm also pretty clear about what text I'm looking at in the article. In the title I used "Harvard" over "The Crimson" because I figured fewer people would know "The Crimson" compared to "Harvard".

Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm curious which assumptions you think are horrible or what data shoddy. I don't think, for example, that I'm assuming the latent space model is "correct" as you say. I don't think I really have any significant assumptions about the technique or the meaning behind it. I read about the technique in the linked paper and reproduced it in my blog with a different dataset and found a similar result. It's strange to me that the signal produced by this technique is as consistent as it is across the 120 years of data. Beyond that, I'm pretty explicit that I don't know what it means or why it happens.

Regarding the "arbitrarily fit" line - as I say explicitly in the post, that's a regression plot to illustrate the trend.

Regarding the possibilities that The Crimson has more articles per year - it's true that's possible. It's not reality, they run about (for a generous definition of "about") the same number of articles every year. The articles do get longer over time. Either way, it's not clear to me what impact this should have on average cosine distance.

There are a lot of things that I looked at that didn't make it into the blog post. Without including them, then perhaps it looks like I'm cutting corners. If I did include them then I think the blog post would be shooting off in many directions. For example, I considered that political violence might be related - like maybe, in times where there's lots of political violence elite institutions come together and their language becomes more similar. That didn't really pan out though. I graphed a bunch of things that ultimately I decided didn't contribute very much and did not include.

Another way of thinking about it is in the original article Rasmussen (the original author) says "Look at this elite writing in NSF grants. The cosine distance is decreasing over time." I then say "Here is some elite writing - student newspaper at an elite school. Is the cosine distance decreasing there over time too?" And, it is. That's what the blog post is trying to say.

Now, maybe the latent space is "incorrect" - although Rasmussen and I use different embeddings that find a similar trend. Maybe it's not meaningful to use cosine distance in this context. But, it does seem like something has to cause it. Whatever it is and whatever it means, it doesn't look like the kind of thing that happens entirely by chance because it is consistent in different datasets and over many years.

Re: Language homogenization at Harvard

#39

Noam Chomsky on suicide watch In my point of view this is actually good news; Since I'm not a native english speaker and since this language has become the world's lingua franca it's very nice being able to understand an article just because the author's didn't waste time looking for bombastic synonyms just to sound smarter. I applaud this and projects like Simple English Wikipedia which allows us, non-native english…

The rising proportion of non-native speakers in US academia is what I thought of when I saw the article's claim that academic vocabulary is shrinking over time. Papers written by adult language learners tend to cluster around domain terminology + the absolute most common English words + words that happen to be similar to ones from their native language, and there's not a lot of incentive for them to go read Charles Dickens until they're indistinguishable from native speakers because, really, it doesn't matter and nobody cares.

Re: Language homogenization at Harvard

#40

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

Yea exactly. This is survivorship bias - the prospects of today look at the winners of yesterday to guide their entry on how to communicate, they’re slightly more likely to get funded, etc It’s reinforcement learning over decades

Survivorship bias is when excessively high-risk behaviors are made to look good by only showing the winners. I think what you're referring to is evolution.
Post reply on HN