Live data from Hacker News

Language homogenization at Harvard

inteoryx.com

71–80 of 87 posts

Re: Language homogenization at Harvard

#71

I would interested if grant proposals which do _not_ get funded, show less of language homogenization than ones that do. Although, it still wouldn't tell you if the idea space is being constricted; that seems like a much more difficult thing to measure. It could just be the equivalent of everyone at court speaking the way the king does, because he's the king. Winning grants probably get read, and imitated, more than…

Yeah…this.

How much of the homogenization is due to “broader impacts” and “intellectual merit”. Which became the default lingo of nsf grants in about 1997

Re: Language homogenization at Harvard

#72
My first thought was that this is just the output from the loss function for a genetic algorithm. Lots of people write grant applications, some of those get accepted, the next round of people writing applications then look at the success stories from the previous round and emulate them. Rinse and repeat. New members of the grant evaluation board want to conform to the standards of their predecessors, and so look at the accepted applications from yesteryear to see what they should accept, rinse and repeat. Its a pretty classic self reenforcing feedback loop, and as other commenters have pointed out constriction in lexical space doesnt mean a constriction in idea space. That said, I really enjoyed the article! Love this kind of data driven analysis.

Re: Language homogenization at Harvard

#73
post #64
post #57

Earlier quoted context omitted.

They trained in 2010 just as hard as now. And learned from others for decades. I guarantee you they wanted to win as much and were trying new jumps. What you see there is what you see in any sport - evolution of it toward better performance.

Kotler's The Rise of Superman goes into how some these meta-performance skills are a big part of what's being transferred more and more effectively. There's young kids doing mind blowing things in extreme sports like skateboarding and BMX that were unimaginable to even the top echelon of the sports a few decades ago. It's exciting to see transfer learning's effects in near real time, as a lot of that has occurred ove…

Most of the meta performance skills and knowledge can be found in psychology textbooks.

Consult with a someone in the field.

Re: Language homogenization at Harvard

#75
Typical ML thinking, it seems.

The author should ask the question: What does my "model" actually assign here? Instead, this very core of the work is ignored and the focus rests completely on the numbers put out of a black box...

If the model simply assigns a closer similarity to more modern words (i.e., it would evaluate older words as "weird"), we would expect exactly this outcome, no?

Re: Language homogenization at Harvard

#76
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

>As others have said, the usage of similar words is no convincing evidence for homogenized ideas

I disagree. Words are used to convey ideas, so if the space of words is shrinking, one should assume that the space of ideas is shrinking. It's possible for this not to be the case, but if word-space is shrinking then the burden of proof should be on those who claim that idea-space is not shrinking.

Maybe we could use embeddings / nlp analysis to determine whether idea-space is shrinking. Or just get a bunch of people to read abstracts from different time-periods and rate how similar they are to one another in their semantic content.

Re: Language homogenization at Harvard

#77
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

>As others have said, the usage of similar words is no convincing evidence for homogenized ideas I disagree. Words are used to convey ideas, so if the space of words is shrinking, one should assume that the space of ideas is shrinking. It's possible for this not to be the case, but if word-space is shrinking then the burden of proof should be on those who claim that idea-space is not shrinking. Maybe we could use emb…

Words are like letters making up ideas, not ideas themselves. Having more than 26 letters wouldn't make us more expressive, and having fewer (like many extant languages)... wouldn't make us any less.

Re: Language homogenization at Harvard

#78
post #66

Earlier quoted context omitted.

The internet has made it very easy for people to homogenize both words and ideas very rapidly. This is the cost of making knowledge universal I guess. Maybe it’s a good thing if it allows us to progress all on the same page rather than taking the time to understand differences? If progress is good. Whatever progress is. Who knows. Although the ideas aren’t really homogenizing in a lot of cases, it is often causing di…

Just as an observation about this…I’ve increasingly begun to see the term knowledge expand to encompass information as if they are equivalent. That tracks, in my field, with changes in US K-12 education and the problems we now see in college learning behaviors. I agree with your general ramble, but found this interesting in context.

Yeah, good point, should have said information. Knowledge cannot be transferred directly to people, only information. The person needs to understand the information to convert it into knowledge. An analogy I have heard used before is that data is the primitive, the integral of data is information, and the integral of information is knowledge.

Re: Language homogenization at Harvard

#79
post #63
post #59

Earlier quoted context omitted.

> Lexical homogenization is not the same as idea homogenization Actually it IS evidence for idea homogenization. Words represent ideas. Unless you are claiming that the same word is used to represent different ideas, which would be even more confusing than having different words representing the same idea. Therefore fewer unique words => fewer unique ideas.

I would not call it a consequence of idea homogenisation but it might certainly lead to this. The basic scientific idea IMHO is compatibility of research. Which is important at least in a competitive and comparitive setting. If people call a measure like 'sensitivity' different in every adjunct field it does not help reviewers. If structural elements of a proposal are similar, it helps understanding quickly the key d…

When, for instance, a science is young, the same concept is often explained using many different terms. Over time canonical naming conventions emerge which standardize the settled parts of the field, leaving less settled future frontiers to be discussed productively with the aid of shorthand.

All named theorems do this in math, continually compressing the lexical space as a way of enabling further out ideas to be even expressed. Granted, math may be the except that proves the rule, but it is an important one.

Re: Language homogenization at Harvard

#80
post #55
post #37

As others have said, the usage of similar words is no convincing evidence for homogenized ideas. I would like to add that the publication referred to in the article notes that there are many more proposals now than there were in the past. This can naturally lead to lower average "distances" between proposals. In a simplistic example, lets assume 100 proposals existed in the "good old times" and they were different at…

It's interesting how your histogram and candle chart shows that, while the mean has shifted, there's a sliver of samples with greater cosine distance than anything previously recorded in the dataset. So I guess while the system has become more inclusive of boilerplate language, it's also become more inclusive of far-out novelties. I'd be interested in reading those abstracts.

yes, that's what i meant by range.
Post reply on HN