Live data from Hacker News

AAAS: Machine learning 'causing science crisis'

bbc.co.uk

81–90 of 119 posts

Re: AAAS: Machine learning 'causing science crisis'

#81
post #34

Case in point : LHC Higgs results - how many detection's vs how many events? How were the detection's determined... The answer is with a large booster [1] I postulate that out of 12 billion random events it would be remarkable if a booster didn't extract 100 or so items that looked similar to a Higgs detection. Well, let's give it 20 years and a new generation of PI's who aren't invested in this and have grad student…

90% of the work in LHC physics is estimating the amplitude of backgrounds that look almost exactly like your signal process. Coming up with the "large booster" for classification is only a small part of it. So yes, machine learning is used, but no, we don't use it blindly like you imply.

As for throwing all the data away, the article you link to actually does a good job of explaining how this is done: we look at every collision with thousands of sensors before deciding whether to keep it. At this stage there is absolutely no machine learning anyway (just physics knowledge), so be careful blaming machine learning for any missed discoveries.

Re: AAAS: Machine learning 'causing science crisis'

#83
post #47

Earlier quoted context omitted.

Some days ago there was a post here on HN about a post doc who failed to get tenure. And every other comment was like: “What did he expect, he had much less than the usual two papers a year.” So it seems that even here on HN, the mindset of quantity over quality still persists.

The problem really is: there is no alternative (which I know of). Having a number of papers in well known journals is everything we have to gauge the quality of people that look for a life time position. It’s sad but true

No, the number of papers is an okay metric, thr problem is that journals don't like to publish honest negative results

Re: AAAS: Machine learning 'causing science crisis'

#84
post #50
post #4

Is machine learning really to blame for the reproducibility crisis? I'm not in academia, but it seemed to me that the problem was entirely present without machine learning being involed. For example, Amgen reporting that of landmark cancer papers they reviewed, 47 of the 53 could not be replicated [1]. I would have assumed that most of them didn't involve 'machine learning' [1] https://www.reuters.com/article/us-scie…

The problem was there before, but there are reasons why Machine Learning is amplifying bad practices. In the past people were manually fishing for results in available datasets. Now they have algorithms to do it for them. In medicine a popular way to use ML is to improve diagnosis. Now there's already a problem in medicine that the benefits of early diagnosis are overrated and the downsides (overtreatment etc.) usual…

Can you give a citation about ML being used this way and producing overtreatment? I'm aware of experimental results with e.g. Watson that were first overstated and then rejected, bit that seemed like a situation where an institution experimented with a poor use of ML and successfully rejected it, not where they falsely accepted an ML result.

Re: AAAS: Machine learning 'causing science crisis'

#85
post #76

My impression from the article was that the doctor stating those opinions has no idea how ML works and how to apply it properly, leading to statements like that. "ML gap" is real I guess...

She has a PhD in Statistics from Stanford. The title of her thesis was "Transposable Regularized Covariance Models with Applications to High-Dimensional Data". (http://www.stat.rice.edu/~gallen/)

I think she knows what she is talking about.

Re: AAAS: Machine learning 'causing science crisis'

#86
post #6

Fails to touch on the perverse incentives in academia, "publish or perish" etc. Torturing a dataset to find a p value that a journal will like (or equivalent stat measure) is better for your career than not publishing a paper that will be discredited in time. You have no incentive at all to decide "my results are unconvincing at this point, I'm not going to submit them" and every reason to write them up as a useful c…

We need to find resources to fund "Failures in Science" journals that exclusively seek to publish interesting research that went nowhere. Personally I'd find these far more interesting to study. "Here's some background. Here's a pretty logical, plausible hypothesis we came up with and how, here's our experiment, here's our results, here's our thoughts as to why we were wildly wrong."

The problem with a failed experiment, at least in machine learning, is that it's not always clear what caused the failure. Was it not enough training, was there some small "trick" the model could have used, were the hyper parameters off etc.

There are an infinite set of configurations that could fail, and it's not sufficiently useful to know they failed without understanding why they failed. And analysing failures in a useful way is an extremely difficult and fundamental problem. On the other hand, a successful experiment is an extremely rare event and hence interesting by itself.

Re: AAAS: Machine learning 'causing science crisis'

#87
post #19

The other day someone lamented that you can't get published as an honest ML researcher, because other scientists are rendering whole professions obsolete all the time...

> you can't get published as an honest ML researcher If you research ML, you can publish in ML journals, there are several. If your research is about applying ML to domain problems, are you then an ML researcher or a domain researcher?

The point was that it is difficult to get noticed with down-to-earth work when the whole field seems to be aiming for the stars.

Re: AAAS: Machine learning 'causing science crisis'

#88
post #60

Earlier quoted context omitted.

Medicine may be better than ML but there’s not much in the difference. > COMPare: Qualitative analysis of researchers’ responses to critical correspondence on a cohort of 58 misreported trials > Background > Discrepancies between pre-specified and reported outcomes are an important and prevalent source of bias in clinical trials. COMPare (Centre for Evidence-Based Medicine Outcome Monitoring Project) monitored all tr…

I am just in the process of digging into this paper and covering it in an article, so I'm quite familiar with it. But as bad as this is: What the COMPare project is doing here is documenting the flaws of a process to counter bad scientific practice. The reality in most fields (including pretty much all of CS and ML) is that no such process exists at all, because noone even tries to fix these issues. So you have medic…

This is definitely true. A common "blueprint" for articles in applied CS is first they propose a "novel" algorithm. This algorithm may be very similar to an existing algorithm and in many cases is identical to one. Then they benchmark the algorithm and shows that it performs better on some metrics than existing solutions. These benchmarks are often quite poor, and if you vary them a little, the purported performance increases vanishes.

Re: AAAS: Machine learning 'causing science crisis'

#89
post #63

As other comments observe, the replication crisis predates the use of ML, so the causes are clearly deeper. I think there's actually a very simple explanation for this which lots and lots of people hate, so they're sort of in denial about it. Academia is entirely government funded and has little or no accountability to the outside world. Academic incentives are a closed loop in which the same sorts of people who are…

Maybe better lower corporation taxes and concurrently raise the tax on earnings, because most of the big, old corps are not really driven by people wanting to expand their knowledge but by people wanting to squeeze out a little more cash (might be different if the turf is contested)...

Re: AAAS: Machine learning 'causing science crisis'

#90
post #6

Fails to touch on the perverse incentives in academia, "publish or perish" etc. Torturing a dataset to find a p value that a journal will like (or equivalent stat measure) is better for your career than not publishing a paper that will be discredited in time. You have no incentive at all to decide "my results are unconvincing at this point, I'm not going to submit them" and every reason to write them up as a useful c…

Some days ago there was a post here on HN about a post doc who failed to get tenure. And every other comment was like: “What did he expect, he had much less than the usual two papers a year.” So it seems that even here on HN, the mindset of quantity over quality still persists.

No...that's not what it seems. I remember reading all the comments in that HN thread on Friday. The top comments were dispassionately explaining the reality of academic career trajectories. That's a basic reporting of facts, not a normative claim that quantity is better than quality. They did not themselves say you need to publish more often to be a good professor; rather they discussed the incentives that lead this to be the practical reality of the field.
Post reply on HN