Live data from Hacker News

Statistical challenges and misreadings of literature create unreplicable science [pdf]

stat.columbia.edu

21–30 of 57 posts

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#21

Earlier quoted context omitted.

The goal of physics isn't to produce video games -- that a video game can generate a frame using unphysical formulae and unphysical processes (eg., rasterisation) does not invalidate physics. There's linguistics which produces descriptive theories of the practice of language. There's "computational linguistics" which aims to build discrete algorithmic models of language. Both of these are interested in explanations o…

That's like saying that it's not astronomy if I'm using a radio telescope, only optical telescopes are real astronomy. It makes zero sense. The tools are irrelevant. We don't have physics with and without ml journals. It's like saying AlphaFold isn't biology because it uses ml. It's absurd One goal of linguistics is to produce an account of how language works. If you understand how something works you can reproduce i…

My master's project was applying non-gradient-based ML methods to parameter optimisation in quantum metrology -- I'm aware that the reskin on curve-fitting algorithms called "Machine Learning" may be used as part of science. We've had taylor series approximations to functions for 300 years.

But insofar as we're talking about LLMs, and most products you can use, we're talking about an engineering use. Engineers build systems that "perform" as you say explanatory accuracy is not a kind of engineering performance.

It would equally well be said that, for a video game studio, "the more physicists I fired, the more the games perform better" -- for, of course, physicists have never studied 'making a video-game image convincing', and their techniques are best run on supercomputers, since they'd build explanatory simulations with explanatory accuracy. Video game developers do not do this. I have developed video games also, so I know.

Convincingly placing pixels on a screen as-if governed by the laws of physics does not require actually simulating them. It requires generating frames from a "point of view" that never expose the absense of real physics, or approximations thereof. The world isnt made of triangles, and you cannot clip thru solid objects.

The goal of a person making an LLM is similar to the goal of a video game developer, these are engineering goals: to give a user an experience which they will pay for. Now, a 2d 80s video-game has more actual physics in it than an LLM has models of language use, but still, this is irrelevant to the goals of these creators.

A linguist could analyse an LLM as the conditional probability structure of the english language, as captured largely recent electronic texts -- but this structure is partly what precisely linguists and others are trying to explain.

This is why all curve-fitting to historical data isn't in itself science: it is a restatement of the very target of explanation, the data. The job of scientific fields is to explain that curve.

Its explotation by engineers is explanation-free. It isnt a science, and replaces no science.

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#22

Earlier quoted context omitted.

Every single time (some 4 times in my past life) when I’m familiar with/close to the background story of something that played out in the news I come to the same conclusion: news doesn’t report the facts. I’m in Western Europe, btw. I advise to run this experiment yourself.

That's the problem when you have people responsible for news having the business model of ad-sposored entertainment and not fact reporting.

The conclusion I’m coming to is that this, too, is an exponent of an inherent weakness: we like to be told stories. As children we like to be told stories and this remains. We now have a complete spectrum of story-telling: science fiction, fiction, non-fiction, …, and news. News reporting itself is a continuum, with tabloid gossip on one side and political coverage on the other. We have so much stories to choose from.

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#23

We're increasingly aware today of how the media operates cycles of self-referential and self-justifying citations: a TV show will quote an article that reports "some people" taking an issue, which ends up being a quote from someone interviewed for another newspaper article.. and so on. This "legitimacy laundering" is rampant, and we're now getting towards media literacy levels which expose it for many people. However…

Every single time (some 4 times in my past life) when I’m familiar with/close to the background story of something that played out in the news I come to the same conclusion: news doesn’t report the facts. I’m in Western Europe, btw. I advise to run this experiment yourself.

Another name for a similar but slightly more pernicious problem is the Gell-Mann Amnesia effect: "the phenomenon of experts reading articles within their fields of expertise and finding them to be error-ridden and full of misunderstanding, but seemingly forgetting those experiences when reading articles in the same publications written on topics outside of their fields of expertise, which they believe to be credible."

https://en.wikipedia.org/wiki/Michael_Crichton#Gell-Mann_amn...

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#24

Earlier quoted context omitted.

That's like saying that it's not astronomy if I'm using a radio telescope, only optical telescopes are real astronomy. It makes zero sense. The tools are irrelevant. We don't have physics with and without ml journals. It's like saying AlphaFold isn't biology because it uses ml. It's absurd One goal of linguistics is to produce an account of how language works. If you understand how something works you can reproduce i…

My master's project was applying non-gradient-based ML methods to parameter optimisation in quantum metrology -- I'm aware that the reskin on curve-fitting algorithms called "Machine Learning" may be used as part of science. We've had taylor series approximations to functions for 300 years. But insofar as we're talking about LLMs, and most products you can use, we're talking about an engineering use. Engineers build…

And my PhD and dayjob is doing research on this.

If a student told me they had this view of what ML is, I would tell them that we've failed to educate them.

The thought that physics doesn't care about performance or approximation is silly. Just look at AlphaFold. Heck, I talk to climatologists and material scientists that want the equivalent all the time.

Prediction is the heart of all science. Whether we're talking neroscience, linguistics, physics, etc.

You think people who run things on supercomputers want to do so for some idealistic notion of what science is? No. They have to do so because they don't have good approximations. Just like with protein folding. Places like DESRES used to build supercomputers for that. This is over now.

ML models learn representations of data which can then be reused for many tasks. There's a whole field where we try to understand those representations. Those are explanations of what's going on given that they're such good predictions. The embeddings you get from an LLM are better models of language than anything linguistics ever accomplished. Their goal should be to explain them and probe their limits instead of complaining. If linguists has come up with gpt they would have been celebrating, the method by which you do science is irrelevant as long as it works.

I'll close by quoting Dawkins. "Science. It works, bitches". That's the value of science. Can we predict which molecule will cure cancer? Everything else is ideology and silly thoughts from before the paradigm shift.

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#25

Earlier quoted context omitted.

Every single time (some 4 times in my past life) when I’m familiar with/close to the background story of something that played out in the news I come to the same conclusion: news doesn’t report the facts. I’m in Western Europe, btw. I advise to run this experiment yourself.

That's the problem when you have people responsible for news having the business model of ad-sposored entertainment and not fact reporting.

In Western Europe a lot of the media isn't ad-sponsored.

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#26

Earlier quoted context omitted.

My master's project was applying non-gradient-based ML methods to parameter optimisation in quantum metrology -- I'm aware that the reskin on curve-fitting algorithms called "Machine Learning" may be used as part of science. We've had taylor series approximations to functions for 300 years. But insofar as we're talking about LLMs, and most products you can use, we're talking about an engineering use. Engineers build…

And my PhD and dayjob is doing research on this. If a student told me they had this view of what ML is, I would tell them that we've failed to educate them. The thought that physics doesn't care about performance or approximation is silly. Just look at AlphaFold. Heck, I talk to climatologists and material scientists that want the equivalent all the time. Prediction is the heart of all science. Whether we're talking…

Prediction is not the heart of science, this is early 20th C. mumbojumbo and humean nonesense that gets repeated by curve-fitters because it's all they do.

Explanation is the heart of science, not prediction. All predictions newton would have made of the orbits of the planets would have been wrong (and so on). And this goes for the vast majority of textbooks physics when its applied to very many ordinary situations: no predictive power at all.

ML models learn "representations" of the data, yes. They are models of measures. The model is just an f=sample({(all possible measures,)}). Science provides representations of the data generating process, ie., reality. It says why those are the measures, why the temperature of gas has that value, not a report on what those values were. Nor even a compressed conditional probability model of those values -- there is no atomic theory in a zip of temperatures.

The purpose of a scientific model is to explain these data-representations, not merely to predict them based on some naive regularity assumption about the measurement device: that it will always measure that way in the future.

The reality of the ML is that it offers only intra-distribution generalisation, nothing of the kind of generalisation science offers where all possible distributions induced by intervention on the (explanatory) variables of scientific models are captured. And this is often a scam that only academics can get away with, ex hyp., just assuming that the test distribution "will turn out as expected" in premise.

The reality is that this sort of repetition of historical data, requires extreme control over the data generating process which gives rise to the test distribution. How is that control delivered in practice? If its a medical lab, through untold amounts of toil delivering, say, histological slides "just right" so this dumb process almost works. If its a face tracking, well you'd be hope its not being used by the police -- because they aint orienting the camera at 3.001m at ISO 151.5 from the masses.

This is the problem with models that obtain predictive power without explanatory content: they rely on prediction time being rigged with, in-practice, extreme control mechanisms that are wholly unstated and unknown by the "modellers". Because these are no models of reality at all, but mere repetitions of historical data with unknown, unmodelled and hence unexplained similarity.

Science concerns itself with explanation, that is its goal. Prediction is instrumental. Engineering's goal is the utility of the product, and so any strategy whatsoever, even a dumb, "if its happened before, itll happen again" is permitted.

There is no textbook of physics which models reality by saying, "Well, we suppose in the future, the positions and velocities will just follow the same distribution, but we've no idea why, and how dare you ask, and get out, and doesnt my Ideal-Gas-TransformerModel look pretty? It gets the pressure right for Argon at 20.0001 C in glass jars at about 2.002L"

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#27

We're increasingly aware today of how the media operates cycles of self-referential and self-justifying citations: a TV show will quote an article that reports "some people" taking an issue, which ends up being a quote from someone interviewed for another newspaper article.. and so on. This "legitimacy laundering" is rampant, and we're now getting towards media literacy levels which expose it for many people. However…

Totally, and I think there is a fundamental, deeper, inherent problem in using statistics to determine how you want to manipulate an given object of study.

Researches find that, on average, consumers want X. Companies decide they want to maximize reach and so begin to produce X. Consumers soon have little choice except for X, reaffirming, to future researchers, that consumers want X.

This is why I think a plurality of data is so necessary especially when it comes to anything in the social domain, and why it's imperative that we begin to invest more into cross-disciplinary research. Specialization has gotten us far, but it's starting to lead to breakdowns. The only thing that might offset the statistical observation of consumer behavior is a statistical study of consumer opinion that proves to refute or contradict that behavior... but then in response some people will say "people don't know what they want" further reaffirming the conclusions that the action based on the data itself caused (behavior or observation bias). The application of statistics to social problems essentially becomes a self-fulfilling prophecy.

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#28

Earlier quoted context omitted.

Every single time (some 4 times in my past life) when I’m familiar with/close to the background story of something that played out in the news I come to the same conclusion: news doesn’t report the facts. I’m in Western Europe, btw. I advise to run this experiment yourself.

That's the problem when you have people responsible for news having the business model of ad-sposored entertainment and not fact reporting.

[deleted]

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#29

We're increasingly aware today of how the media operates cycles of self-referential and self-justifying citations: a TV show will quote an article that reports "some people" taking an issue, which ends up being a quote from someone interviewed for another newspaper article.. and so on. This "legitimacy laundering" is rampant, and we're now getting towards media literacy levels which expose it for many people. However…

Could one create a proof of pseudo-science, by injecting a faked fundamental corner stone paper, that becomes proof by inheritance that a full field is rotten?

Also why does this remind me of european politicans, claiming everyone wants to life european lifes, meanwhile whole countries goto war and atrocities without big counter-demonstrations by those western valued citizens .. narrative glider guns going ad absurdum..

Re: Statistical challenges and misreadings of literature create unreplicable science [pdf]

#30
post #25

Earlier quoted context omitted.

That's the problem when you have people responsible for news having the business model of ad-sposored entertainment and not fact reporting.

In Western Europe a lot of the media isn't ad-sponsored.

Just "bad-sponsored"
Post reply on HN