Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

41–50 of 87 posts

Re: Language homogenization at Harvard

#41

This article is almost criminally flawed. The author makes some horrible assumptions and presents shoddy data. Even if we assume that the latent space model is "correct" (we shouldn't), they don't present anything like variance or number of samples in a year. Then they sort of arbitrarily fit a line to data which pretty clearly looks non-linear. Suppose, for example, that The Crimson (not Harvard btw, it's a student…

For the title, I did say "At Harvard" and not "In Harvard" or "By Harvard". I think the student newspaper is indeed at Harvard. I'm also pretty clear about what text I'm looking at in the article. In the title I used "Harvard" over "The Crimson" because I figured fewer people would know "The Crimson" compared to "Harvard". Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm c…

Without commenting on the overall trend's cause, your diversity hypothesis is bunk and suggests you are looking to making things fit a diversity-related narrative:

- there's (unsurprisingly) no significant diversity-word change from 1900 to 1940 but a very significant distance drop

- there's a big diversity-word change around ~1990 with no concomitant distance change

Re: Language homogenization at Harvard

#42

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

The internet has made it very easy for people to homogenize both words and ideas very rapidly. This is the cost of making knowledge universal I guess. Maybe it’s a good thing if it allows us to progress all on the same page rather than taking the time to understand differences? If progress is good. Whatever progress is. Who knows. Although the ideas aren’t really homogenizing in a lot of cases, it is often causing dichotomies. But really it’s no more than a few factions that reinforce their homogenizations. Ok I guess it’s bad… I am rambling as much as that essay. You have increased your follow count to 6 now, good show.

Re: Language homogenization at Harvard

#43

Earlier quoted context omitted.

For the title, I did say "At Harvard" and not "In Harvard" or "By Harvard". I think the student newspaper is indeed at Harvard. I'm also pretty clear about what text I'm looking at in the article. In the title I used "Harvard" over "The Crimson" because I figured fewer people would know "The Crimson" compared to "Harvard". Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm c…

Without commenting on the overall trend's cause, your diversity hypothesis is bunk and suggests you are looking to making things fit a diversity-related narrative: - there's (unsurprisingly) no significant diversity-word change from 1900 to 1940 but a very significant distance drop - there's a big diversity-word change around ~1990 with no concomitant distance change

Your comment is a bit ironic in the sense that I can tell you didn't read the article because you reproduce conclusions from the article. That's okay! Obviously you didn't need to read it to know what I would have said. :)

Let me quote from the end:

"Another argument against connecting distance and diversity is that distance is on a long running decline from 1900 even for the first four decades while diversity words were basically flat. When diversity words pop in the 90's there isn't an immediate reaction in cosine distance, it's only about a decade later, in 2000, that cosine distance takes a steep drop."

That seems awfully similar to the two points you've raised here.

What I do find a bit distasteful is that you jump in with "your diversity hypothesis is bunk" and accuse me of trying to fit a narrative - without even reading what you're commenting on.

Re: Language homogenization at Harvard

#44

I would guess that the increased rate of information transmission across society has contributed to this trend. Copying is a key mechanism in how culture develops and propagates, and information moving more easily makes it easier for different groups to copy the dominant examples for any particular activity or discipline. A lot of information systems and cultural phenomenons have winner-take-all / power-law dynamics.…

Tangentially related, but I was thinking about this while watching the Olympics men's figure skating competition recently (given the ongoing drama on the women's side, I'll leave that aside for now). The quality of skating has drastically improved in just a decade. In 2010 the Olympic champion had no quad jumps. In 2022 the easiest jump from the top 3 competitors in their short program was a triple axel - the rest we…

I wonder if, setting quality aside, the kinds of tricks (is "tricks" right?) that skaters do becomes more similar with technology. That is, skater X was going to do some random trick, until X learned that skaters Y and Z were doing quad jumps, so X figures he must do quad jumps now.

In other words, if we could represent skating routines as vectors, would the average cosine distance between all those vectors be increasing or decreasing?

Re: Language homogenization at Harvard

#45
I'd be curious to see what the plot of average annual cosine distance would look like when using different sets of pre-trained embeddings. I suspect the corpus used is biased toward more recent documents. It wouldn't surprise me if there's more variance in the embeddings of documents that look less like those in the training set, e.g. if you were to embed documents written in German you may get some extreme outliers.

Re: Language homogenization at Harvard

#46

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

Yea exactly. This is survivorship bias - the prospects of today look at the winners of yesterday to guide their entry on how to communicate, they’re slightly more likely to get funded, etc It’s reinforcement learning over decades

Emulating the past for success in the future is a form of idea homogenization.

Re: Language homogenization at Harvard

#47

Maybe we've discovered diversity etc is worth studying, in part because it does have such a profound effect on the world. There's a reason these words and studies are considered "politicized" and that's because understanding them threatens the status quo. Just because something is politicized doesn't mean it isn't worth studying.

> Maybe we've discovered diversity etc is worth studying, in part because it does have such a profound effect on the world

Or alternatively, NSF favors grant applications that mention it and toe the party line, so everyone begins to add more and more garbage to their application to optimize for grant acceptance.

Generally if something is so profound that it disrupts the status quo we quickly see it take over. For example the smartphone. But there still isn't a scientific consensus on diversity despite the well-known liberal lean of college campuses. And companies pay lip service and it almost seems that the most successful companies implement DEI programs _after_ they become successful, implying that it provides no competitive advantage.

Re: Language homogenization at Harvard

#48
2 counter points:

1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying!

2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934 when the novel was written, but not in common usage for some decades now [1]. Using these words today would increase lexical diversity, but at the expense of communicative effectiveness!

[1] https://tinyurl.com/2p8ujk7e

Re: Language homogenization at Harvard

#49

This article is almost criminally flawed. The author makes some horrible assumptions and presents shoddy data. Even if we assume that the latent space model is "correct" (we shouldn't), they don't present anything like variance or number of samples in a year. Then they sort of arbitrarily fit a line to data which pretty clearly looks non-linear. Suppose, for example, that The Crimson (not Harvard btw, it's a student…

For the title, I did say "At Harvard" and not "In Harvard" or "By Harvard". I think the student newspaper is indeed at Harvard. I'm also pretty clear about what text I'm looking at in the article. In the title I used "Harvard" over "The Crimson" because I figured fewer people would know "The Crimson" compared to "Harvard". Regarding the flaws in the article (phew, I'm glad I didn't quite reach criminal level) - I'm c…

Isn't the pre-made model you're using trained almost entirely on recent (last decade or so) text? I didn't dig too far into it but it looks like news, web crawls, twitter, wikipedia, etc.

Re: Language homogenization at Harvard

#50

Earlier quoted context omitted.

Without commenting on the overall trend's cause, your diversity hypothesis is bunk and suggests you are looking to making things fit a diversity-related narrative: - there's (unsurprisingly) no significant diversity-word change from 1900 to 1940 but a very significant distance drop - there's a big diversity-word change around ~1990 with no concomitant distance change

Your comment is a bit ironic in the sense that I can tell you didn't read the article because you reproduce conclusions from the article. That's okay! Obviously you didn't need to read it to know what I would have said. :) Let me quote from the end: "Another argument against connecting distance and diversity is that distance is on a long running decline from 1900 even for the first four decades while diversity words…

Hey, I thought your article was nice. First it had an easy intro to word embeddings and cosine similarity. And second, you followed the investigation and even came up with the idea "against connecting distance and diversity", so it didn't seem you had the conclusion before you started the work.

If someone complains about not mentioning variance - it's still implicitly visible by the cloud of dots representing each year around the regression line.

Post reply on HN