Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

81–87 of 87 posts

Re: Language homogenization at Harvard

#81
post #62
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

I think that's possible, but in your model, the lower average distance has indeed decreased. Yeah, the total range (or maybe the convex hull of the idea space) is bigger, but it's not obvious that the range is what we should think about. If the top 100 proposals are small variations on the "most attractive" old idea, then a lot hangs on whether that idea is really good or not - which in turn suggests that the proposa…

> the lower average distance has indeed decreased

Yes, that was indeed the point I was trying to convey! In an even sillier example, assume word vectors X, then calculate "proposal by proposal" similarities (i.e. inverse distances). Then duplicated X and concatenate [X,X], recalculate "proposal by proposal" distances (now for twice as many proposals)---those distances must now be less on average because each proposal has at least one "zero distance" neighbor. HOWEVER, why would you assert that the overall "idea space" has been reduced?

Re: Language homogenization at Harvard

#82

Lexical homogenization is not the same as idea homogenization, and it does not surprise me that as time goes on the word choice in a given category of writing (esp an insular one like grant applications) would constrict. It also wouldn't surprise me a huge amount if the idea space was constricting, but just looking at the cosine distance is not enough to establish that. There probably are meaningful ways to exploit a…

The internet has made it very easy for people to homogenize both words and ideas very rapidly. This is the cost of making knowledge universal I guess. Maybe it’s a good thing if it allows us to progress all on the same page rather than taking the time to understand differences? If progress is good. Whatever progress is. Who knows. Although the ideas aren’t really homogenizing in a lot of cases, it is often causing di…

The point of the thing for me was that the Crimson analysis showed no change in rate of decline since 1900. So there is a need for more examples over larger timelines before we start modelling, perhaps.

The analysis also seems sensitive to the mapping of words to categories. Some kind of robustness analysis to this sensitivity would also be interesting to see.

Re: Language homogenization at Harvard

#83

Earlier quoted context omitted.

The internet has made it very easy for people to homogenize both words and ideas very rapidly. This is the cost of making knowledge universal I guess. Maybe it’s a good thing if it allows us to progress all on the same page rather than taking the time to understand differences? If progress is good. Whatever progress is. Who knows. Although the ideas aren’t really homogenizing in a lot of cases, it is often causing di…

The point of the thing for me was that the Crimson analysis showed no change in rate of decline since 1900. So there is a need for more examples over larger timelines before we start modelling, perhaps. The analysis also seems sensitive to the mapping of words to categories. Some kind of robustness analysis to this sensitivity would also be interesting to see.

[deleted]

Re: Language homogenization at Harvard

#84
The original paper that inspired this article is a hot mess. It’s just one huge correlation/causation confusion of a political argument posed as research. My freshman expository writing professor tore classmates to shreds over more justifiable logical constructs.

It does show lexical homogenization. It does show an increase in their shoddily assembled list so-called diversity words. The concluded implications and pretty much every other part is hand waving, assumptions and political insinuation.

Like here where the author disclaims the shoddiness of using word frequency as a measure of politicization because the nuance and subtlety of bias make quantitative measurement difficult (he should apply for a grant to research that)… but then draws the entirely unsupported conclusion that politicization is actually worse than their only data implies:

“Note that word counts is a somewhat crude way of measuring politicization. Bias, particularly in the social sciences, is often subtle, and can apply to the kinds of questions that get asked and the standard of evidence used to accept or reject a hypothesis. Thus, the fact that so many grants contain terms that are in most contexts clearly associated with left-wing political causes likely underestimates the degree of politicization in science funding.”

And the author puts that assumption to work a couple of paragraphs and charts later:

”This report presents direct evidence that scientific funding at the federal level has become more politicized and less supportive of novel ideas since 1990.”

The opening paragraphs don’t even indicate a connection beyond “Look at this trend which I assume means Y. Now look at that trend which I assume means Y. Now look at them together. Vaguely similar, eh? Awfully suspicious, eh?.”

“Taken together, the results imply that there has been a politicization of scientific funding in the US in recent years and a decrease in the diversity of ideas supported.”

How about a real “diversity word” selection criteria instead of just showing some tenuously relevant clustering? How about contextual analysis or explanation of the terms? Language and public discourse have changed a lot since the 90s. To inform and help validate their word selection analyzing ~30 years of data, they used the “DEI terms” diversity, equity and inclusion, but the DEI acronym wasn’t used until about 10 years ago— it could be a coincidence, but could indicate those terms weren’t always an accepted standard — is there evidence other terms weren’t used instead? Is there any evidence those are the best terms to use now? Did 90s political grants use euphemisms or code words for political ideas to send more neutral? Were the terms you selected used by political activists as commonly in 1993 as they are now? Would the declining prevalence of second wave feminists and original civil rights movement activists, to use two prevalent examples of very political movements ubiquitous in academia long before 1990, have been equally political but using different words? How many of those words are central to the research topic and how many are positioning unrelated research as good for society or aware of current trends or more beneficial than they are? What other words or topics showed similar/different results for comparison? So-called inclusion words removed, was the lexical similarity static? During that same period, a small number of ubiquitous spelling and grammar checkers gained sophistication and adoption; writers who disregarded or disabled them probably retired and others probably changed how they wrote; the Internet sharpened the curves of trends and memes and also probably exposed people to better writing practices; writing trends have probably come and gone; educational curricula become more standardized; the Internet facilitated access to far more grant/business writing examples; there may have been particularly influential events or pieces of writing that changed things for other reasons— how did they consider or reason about factors like that? Do Internet-focused terms which would have gained prevalence in the same time frame follow a similar trend line? Any medical terms? BPA? GMOs? Genetic testing? Any other points of comparison at all?

Anyone interested in actually advancing human understanding rather than creating political cudgels could poke holes in this garbage all day. I don’t see how anyone could look at that and think “huh — thoughtful analysis” rather than “huh — that’s some elaborate work to give a tenuous air of legitimacy to a claim they didn’t test presented in a manipulative piece which won’t get much traction beyond a politically sympathetic subreddit.”

Re: Language homogenization at Harvard

#85
post #64
post #57

Earlier quoted context omitted.

They trained in 2010 just as hard as now. And learned from others for decades. I guarantee you they wanted to win as much and were trying new jumps. What you see there is what you see in any sport - evolution of it toward better performance.

Kotler's The Rise of Superman goes into how some these meta-performance skills are a big part of what's being transferred more and more effectively. There's young kids doing mind blowing things in extreme sports like skateboarding and BMX that were unimaginable to even the top echelon of the sports a few decades ago. It's exciting to see transfer learning's effects in near real time, as a lot of that has occurred ove…

Skateboarding is fairly new sport. A few decades ago, there was nothing. Same goes for BMX. Both at current level require technology unavailable few decades ago.

The trick were really invented not that far ago.

Re: Language homogenization at Harvard

#86
post #68
post #45

I'd be curious to see what the plot of average annual cosine distance would look like when using different sets of pre-trained embeddings. I suspect the corpus used is biased toward more recent documents. It wouldn't surprise me if there's more variance in the embeddings of documents that look less like those in the training set, e.g. if you were to embed documents written in German you may get some extreme outliers.

The telling part for me is the scale of the graphs from the original paper and the blog post The paper graph shows a decrease from like .1 to .075 over 30 years The blog post shows a decrease from .35 to .1 over 100 years However, the crimson had a notable drop around 2000 - on the order of the entire decrease in the research study - and then looks fairly stable. It’s almost like there are policy and effective commun…

This. Diversity words are almost always describing the Broader Impact. Meanwhile, the original study is pretending they are for the research ideas of the grant.

Re: Language homogenization at Harvard

#87
post #81
post #62

Earlier quoted context omitted.

I think that's possible, but in your model, the lower average distance has indeed decreased. Yeah, the total range (or maybe the convex hull of the idea space) is bigger, but it's not obvious that the range is what we should think about. If the top 100 proposals are small variations on the "most attractive" old idea, then a lot hangs on whether that idea is really good or not - which in turn suggests that the proposa…

> the lower average distance has indeed decreased Yes, that was indeed the point I was trying to convey! In an even sillier example, assume word vectors X, then calculate "proposal by proposal" similarities (i.e. inverse distances). Then duplicated X and concatenate [X,X], recalculate "proposal by proposal" distances (now for twice as many proposals)---those distances must now be less on average because each proposal…

Here's one metric by which, in your first model, the overall idea space has been reduced: the distribution of models has become more concentrated. That's because 100 proposals are tiny variations around 1 basic one. The same holds in your second model: with word vectors X, if I pick (say) two ideas to fund at random, they will never be the same idea, while with (X, X), that will sometimes happen.
Post reply on HN