Live data from Hacker News

AAAS: Machine learning 'causing science crisis'

bbc.co.uk

111–119 of 119 posts

Re: AAAS: Machine learning 'causing science crisis'

#111
post #83

Earlier quoted context omitted.

No, the number of papers is an okay metric, thr problem is that journals don't like to publish honest negative results

There are a few journals that will promise to publish your results based entirely on your pre-registered plan -- e.g. whether or not you find the correlation you were looking for. In the long term, it seems like journals that publish lots of falsified papers should be punished, and journals that don't (e.g. because of a judge-upon-pre-registration policy) should crowd them out.

Pre-registration is great, and I think it even helps researchers/scientists/post-docs to submit something (it doesn't have to be perfect, after all, it's just an experiment design, and they don't have to worry about massaging the data to have promising results to report) and then stick to the topic, and then carry out the experiment as well as they can, and gather as much high quality & fidelity data as they can, and do the analysis according to the plan, and then report it. No extra pressure to think about how to frame what you "found".

Though this will inevitably lead to the problem that grant boards face, that it'll be a lot harder to differentiate between proposed experiments. And it'll be even harder to do boring stuff. (So if we assume all submitted plans are sound, they have to publish them all. Though then we'll have journals based on how strict they are with experiment design requirements, 1 sigma, 2 sigma, 5 sigma, etc.)

Re: AAAS: Machine learning 'causing science crisis'

#112
post #5

I wonder: don't machine learning frameworks' results come with a level of confidence? Ps: I have no experience with anything regarding ML.

Yes, you can create confidence estimates for both large neural networks and gradient boosting (see for instance the thesis of Yarin Gal). This covers the majority of commercial and academic applications.

ML is actually a field with very high standards for replication, in part because emperical results are currently the focus. If certain methods don't generalize to other datasets, then all bets are off: you are dealing with data that violates the IID assumption. No statistics, bean counting, or ML is going to help you get significant results.

Re: AAAS: Machine learning 'causing science crisis'

#113
post #51

Earlier quoted context omitted.

We need to find resources to fund "Failures in Science" journals that exclusively seek to publish interesting research that went nowhere. Personally I'd find these far more interesting to study. "Here's some background. Here's a pretty logical, plausible hypothesis we came up with and how, here's our experiment, here's our results, here's our thoughts as to why we were wildly wrong."

We have that. People don't use it. PLOS e.g. specifically stated multiple times that they'll publish what meets their quality standards regardless of a positive or negative outcome. The problem is: Even if you publish failed research it won't get cited as much. And people still use citation metrics to evaluate "quality" of science. Just having journals that publish your "failed" research is a good start, but it's not…

Agreed. This isn't a one-action problem to solve. But we need to lay the foundation for alternative incentives to be possible.

Re: AAAS: Machine learning 'causing science crisis'

#114
post #12

Earlier quoted context omitted.

ML is statistics with a different name using a computer, that should always be mentioned in articles for the general public. On the other hand AI is fantasy BS hype boosting off the fact that ML sounds similar to AI to people who aren't aware Machine Learning is stats. Maybe AI one day but today it is utterly ridiculous. No really. Every single article should mention both of those things at least in passing. Downvote…

ML is different from statistics. If you want to learn more read Breiman's Statistical Modeling: Two cultures. AI is real and is a legit field of study, of which ML is currently very popular, so some articles conflate these two. But it is far from bullshit. If you want to learn more read Artificial Intelligence: A modern approach.

"statistical modelling:" AI in a pitch is exactly as much BS as interstellar travel. Definitely go ahead and study all the failed attempts. Norvig wrote the most popular text nearly a quarter of a century ago. It didn't exist then either. There is no going to alpha centuri there is no Hal9000. There is no intelligence that is artificial.

Machine Learning exists and is real and is statistical inference done by computer. There is no point where you can look at it and say "Here is the boundary between statistics and ML." Try it. Glorified curve fitting has some fantastic applications and killer demos. Along comes the hype train exactly as you'd expect. Put it on the blockchain or has the hype for that died now?

https://www.technologyreview.com/s/612437/what-is-machine-le...

Shallow BS detection, worth doing if only for one's own sanity.

Re: AAAS: Machine learning 'causing science crisis'

#115
post #24
post #12

Earlier quoted context omitted.

ML is statistics with a different name using a computer, that should always be mentioned in articles for the general public. On the other hand AI is fantasy BS hype boosting off the fact that ML sounds similar to AI to people who aren't aware Machine Learning is stats. Maybe AI one day but today it is utterly ridiculous. No really. Every single article should mention both of those things at least in passing. Downvote…

I'd say currently ML is heuristics done by a computer. The rigour of statistics isn't quite present in ML yet. At least, not in the basic courses yet.

You can equally do the thing named statistics without rigour. In fact that describes most of it. Sadly. See replication crisis, p-hacking, garden of forking data etc. etc. etc. I'd say the overwhelming majority of university stats courses have no rigour at all and that may not be a bad thing in and of itself?

Re: AAAS: Machine learning 'causing science crisis'

#116
post #76

My impression from the article was that the doctor stating those opinions has no idea how ML works and how to apply it properly, leading to statements like that. "ML gap" is real I guess...

She has a PhD in Statistics from Stanford. The title of her thesis was "Transposable Regularized Covariance Models with Applications to High-Dimensional Data". ( http://www.stat.rice.edu/~gallen/ ) I think she knows what she is talking about.

Thanks for pointing this out. The professor cited most certainly knows the best of ML. I'm guessing the author of this article simply attended the AAAS session where Allen gave a talk on recent work on addressing inferential challenges with modern ML and wrote this piece that doesn't do her work justice. See https://aaas.confex.com/aaas/2019/meetingapp.cgi/Session/215... & list of recent papers https://arxiv.org/search/?query=genevera+allen&searchtype=al...

Nearly all statisticians realize the need for more inferential thinking in modern ML. E.g. http://magazine.amstat.org/blog/2016/03/01/jordan16/ We still don't do that well in high dimension, low sample size regimes that make up the majority of life science research.

Re: AAAS: Machine learning 'causing science crisis'

#117
post #76

My impression from the article was that the doctor stating those opinions has no idea how ML works and how to apply it properly, leading to statements like that. "ML gap" is real I guess...

She has a PhD in Statistics from Stanford. The title of her thesis was "Transposable Regularized Covariance Models with Applications to High-Dimensional Data". ( http://www.stat.rice.edu/~gallen/ ) I think she knows what she is talking about.

Smart and beautiful! :)

Re: AAAS: Machine learning 'causing science crisis'

#118

How does this article manage not to mention a single actual example of ML-related misconceptions?? I'm sure they exist, but there is literally nothing here except some assertions and a plug for a vaguely remedial research line.

Here is the missing context from press release:

``` "In precision medicine, it's important to find groups of patients that have genomically similar profiles so you can develop drug therapies that are targeted to the specific genome for their disease," Allen said. "People have applied machine learning to genomic data from clinical cohorts to find groups, or clusters, of patients with similar genomic profiles.

"But there are cases where discoveries aren't reproducible; the clusters discovered in one study are completely different than the clusters found in another," she said. "Why? Because most machine-learning techniques today always say, 'I found a group.' Sometimes, it would be far more useful if they said, 'I think some of these are really grouped together, but I'm uncertain about these others.'"

Allen will discuss uncertainty and reproducibility of ML techniques for data-driven discoveries at a 10 a.m. press briefing today, and she will discuss case studies and research aimed at addressing uncertainty and reproducibility in the 3:30 p.m. general session, "Machine Learning and Statistics: Applications in Genomics and Computer Vision." Both sessions are at the Marriott Wardman Park Hotel. ``` https://eurekalert.org/pub_releases/2019-02/ru-cwt021119.php

& the context of the AAAS session https://aaas.confex.com/aaas/2019/meetingapp.cgi/Session/215...

Re: AAAS: Machine learning 'causing science crisis'

#119
post #63

As other comments observe, the replication crisis predates the use of ML, so the causes are clearly deeper. I think there's actually a very simple explanation for this which lots and lots of people hate, so they're sort of in denial about it. Academia is entirely government funded and has little or no accountability to the outside world. Academic incentives are a closed loop in which the same sorts of people who are…

Somehow, it seems like during the Cold War (50s, 60s), there were fewer scientists and the quality of the output was higher. Not sure if that is the case, or survivorship bias. But if it is the case, what was different about the system back then?

To play the devil's advocate: scientists in industry/corporations do not come out of nowhere - they come from academia. Will the academics not move to countries where academic research is better funded? The students will follow. Corporations will set up their research labs in those countries near the universities to poach the best talent. Suddenly, your country is at a disadvantage.

Post reply on HN