Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

61–70 of 87 posts

Re: Language homogenization at Harvard

#61
post #45

I'd be curious to see what the plot of average annual cosine distance would look like when using different sets of pre-trained embeddings. I suspect the corpus used is biased toward more recent documents. It wouldn't surprise me if there's more variance in the embeddings of documents that look less like those in the training set, e.g. if you were to embed documents written in German you may get some extreme outliers.

This point should get more visibility — embeddings are not made in the abstract; they reflect the lexicon of their training sets, and I strongly suspect word embeddings used here reflect the lexicon and word frequencies (and their meaning/usage context) in modern literature. Using them for text in the past (nearly a century back!) warrants some skepticism.

Re: Language homogenization at Harvard

#62
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

I think that's possible, but in your model, the lower average distance has indeed decreased. Yeah, the total range (or maybe the convex hull of the idea space) is bigger, but it's not obvious that the range is what we should think about. If the top 100 proposals are small variations on the "most attractive" old idea, then a lot hangs on whether that idea is really good or not - which in turn suggests that the proposals are probably not providing enough diversity.

Re: Language homogenization at Harvard

#63
post #59

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

> Lexical homogenization is not the same as idea homogenization Actually it IS evidence for idea homogenization. Words represent ideas. Unless you are claiming that the same word is used to represent different ideas, which would be even more confusing than having different words representing the same idea. Therefore fewer unique words => fewer unique ideas.

I would not call it a consequence of idea homogenisation but it might certainly lead to this. The basic scientific idea IMHO is compatibility of research. Which is important at least in a competitive and comparitive setting. If people call a measure like 'sensitivity' different in every adjunct field it does not help reviewers. If structural elements of a proposal are similar, it helps understanding quickly the key difference that remain. And this difference should be actually the actual idea, which is often the smallest part. We actually teach students to emulate style and do the same to write successful grant applications.This sure creates a bubble and in effect hinders outsiders to enter. That we see a gradual assimilation is IMHO rather an effect of available 'training material' and interdisciplinarity. Sure there are dangers that this might be an indicator for. However, do not misinterpret it in a way that it make a grant proposal more novel just because it uses totally different language...

Re: Language homogenization at Harvard

#64
post #57

Earlier quoted context omitted.

Tangentially related, but I was thinking about this while watching the Olympics men's figure skating competition recently (given the ongoing drama on the women's side, I'll leave that aside for now). The quality of skating has drastically improved in just a decade. In 2010 the Olympic champion had no quad jumps. In 2022 the easiest jump from the top 3 competitors in their short program was a triple axel - the rest we…

They trained in 2010 just as hard as now. And learned from others for decades. I guarantee you they wanted to win as much and were trying new jumps. What you see there is what you see in any sport - evolution of it toward better performance.

Kotler's The Rise of Superman goes into how some these meta-performance skills are a big part of what's being transferred more and more effectively. There's young kids doing mind blowing things in extreme sports like skateboarding and BMX that were unimaginable to even the top echelon of the sports a few decades ago.

It's exciting to see transfer learning's effects in near real time, as a lot of that has occurred over my lifetime. I just wish meta learning was more accessible and common knowledge. As it stands, I find myself continually having to do my own research to uncover ways to improve my own improvement.

Re: Language homogenization at Harvard

#65

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

> Lexical homogenization is not the same as idea homogenization,

It is for the word thinkers.

Re: Language homogenization at Harvard

#66

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

The internet has made it very easy for people to homogenize both words and ideas very rapidly. This is the cost of making knowledge universal I guess. Maybe it’s a good thing if it allows us to progress all on the same page rather than taking the time to understand differences? If progress is good. Whatever progress is. Who knows. Although the ideas aren’t really homogenizing in a lot of cases, it is often causing di…

Just as an observation about this…I’ve increasingly begun to see the term knowledge expand to encompass information as if they are equivalent. That tracks, in my field, with changes in US K-12 education and the problems we now see in college learning behaviors.

I agree with your general ramble, but found this interesting in context.

Re: Language homogenization at Harvard

#67
post #48

2 counter points: 1. Language needs to be somewhat homogenous. At an extreme, if the languages spoken by two academics are entirely heterogenous, they (literally) won't understand a word the other is saying! 2. I recently investigated why there were 300 (!!) words I did not know the meaning of in a single George Orwell novel. Google N-Gram Viewer shows around 60-80% of these words to have been in common usage in 1934…

Quick, someone make a 1934 version of Wordle!

I can hear the steampunk aesthetic from here.

Re: Language homogenization at Harvard

#68
post #45

I'd be curious to see what the plot of average annual cosine distance would look like when using different sets of pre-trained embeddings. I suspect the corpus used is biased toward more recent documents. It wouldn't surprise me if there's more variance in the embeddings of documents that look less like those in the training set, e.g. if you were to embed documents written in German you may get some extreme outliers.

The telling part for me is the scale of the graphs from the original paper and the blog post

The paper graph shows a decrease from like .1 to .075 over 30 years

The blog post shows a decrease from .35 to .1 over 100 years

However, the crimson had a notable drop around 2000 - on the order of the entire decrease in the research study - and then looks fairly stable.

It’s almost like there are policy and effective communication reasons for narrowing your vocabulary to that appropriate for a target audience.

If you look at the graph from the underlying study, there’s a bit of a “shoulder” around 1997. That is telling. 1997 was when the NSF introduced a new, clear and explicit, set of review criteria - broader impacts and intellectual merit [0]. That change alone is likely a cause of significant (Meaningful?) linguistic narrowing. To get NSF funding, researchers now had to explain why their research matters using explicit language that aligned with specific strategic objectives of the funding agency.

Then you add a layer of Goodharts law. During this time period, two other things were also happening. First, University’s increasingly began to rely on external funding - especially public university’s. Second, the field of “faculty development” was increasingly formalizing and offering training and support on things like grant writing. Those trainings include a lot of focus on using normalized, almost shibboleth like, language in grant applications. Ensuring that there is religion of key words so that it is easy for the reviewers to establish that a particular grant application addresses the required review criteria.

So the data set used here is part of the problem - if they had used submitted rather than funded they would likely see different results. Not because ideas are narrowing but because grant applications are basically a human manifested api, and someone tried to actually standardize it.

[0] https://stem.colostate.edu/a-history-of-the-broader-impacts-...

Re: Language homogenization at Harvard

#69
post #60
post #21

There's no evidence from this article that the idea space is decreasing. Vocabulary is becoming somewhat more similar: the measurement is on individual words. Suppose an institution decides that the faculty is using overly obscure language (or a grant-making agency decides this) and asks that grant submitters reduce the amount of jargon, include more background, etc. Wouldn't this measurement then show a significant…

Words represent ideas. Fewer words is good evidence that there are fewer ideas. Sometimes multiple words represent the same idea, so in that case homogenization is "good". It is also possible that multiple ideas are represented by the same word, but that is "bad" in an academic context because it leads to ambiguity. Therefore fewer unique words IS actually good evidence that the idea space is decreasing, unless you a…

Words don’t represent JUST represent ideas.

Words transmit ideas. The difference is notable I’m this data set in particular. I posted a longer version elsewhere in the thread but to summarize.

During this time period the nsf became more strict about explicitly addressing their review criteria. Faculty became trained to use the languag explicitly and connect their ideas to the language of the review criteria. I’m effect, you started to get an ‘api’ where, irregardless of what the idea in the grant is, it’s is clearly and explicitly connected to key nsf terminology. That is a de novo narrowing of the language space no matter the underlying ideas. It’s purpose is definitively about improving communication, not narrowing ideas.

From what I can tell the paper didn’t filter anything in the data. Had they filtered “broader impacts” and “intellectual merit” I would be hard money they would get different results.

Re: Language homogenization at Harvard

#70

Noam Chomsky on suicide watch In my point of view this is actually good news; Since I'm not a native english speaker and since this language has become the world's lingua franca it's very nice being able to understand an article just because the author's didn't waste time looking for bombastic synonyms just to sound smarter. I applaud this and projects like Simple English Wikipedia which allows us, non-native english…

The rising proportion of non-native speakers in US academia is what I thought of when I saw the article's claim that academic vocabulary is shrinking over time. Papers written by adult language learners tend to cluster around domain terminology + the absolute most common English words + words that happen to be similar to ones from their native language, and there's not a lot of incentive for them to go read Charles D…

Add in “grant writing specialists” with no domain expertise but lots of “test the nsf as an api” expertise and you have even bigger sources of error in these claims.
Post reply on HN