Live data from Hacker News

Deep learning gets the glory, deep fact checking gets ignored

rachel.fast.ai

151–160 of 174 posts

Re: Deep learning gets the glory, deep fact checking gets ignored

#151
post #6

Fantastic article by Rachel Thomas! This is basically another argument that deep learning works only as a [generative] information retrieval - i.e a stochastic parrot, due to the fact that the training data is a very lossy representation of the underlying domain. Because the data/labels of genes do not always represent the underlying domain (biology) perfectly, the output can be false/invalid/nonsensical. in cases wh…

> works only as a [generative] information retrieval

but even if we for simplicity of the argument assume that is true without question, LLM still are here to stay

Like think about how do junior devs which (in programming) average or less skill work, they "retrieve" the information about how to solve the problem from stack overflow, tutorials etc.

So giving all your devs some reasonable well done AI automation tools (not just a chat prompt!!) is like giving each a junior dev to delegate all the tedious simple tasks, too. Without having to worry about that task not allowing the junior dev to grow and learn. And to top it of if there is enough tooling (static code analysis, tests, etc.) in place the AI tooling will do the write things -> run tools -> fix issues loops just fine. And the price for that tool is like what, a 1/30th of that of a junior dev? Means more time to focus on the things which matter including teaching you actual junior devs ;)

And while I would argue AI isn't full there yet, I think the current fundation models _might_ already be good enough to get there with the right ways of wiring them up and combining them.

Re: Deep learning gets the glory, deep fact checking gets ignored

#152
post #8

Wow, great lead.... The worse science, publish or perish pulp, got more academic karma Altmetric/Citations -> $$$ AI is a perfect academic, the science and curiosity is gone and the ability to push out science looking text is supermaxxed. Tragic end solution, do the same and throw even more money at it > At a time when funding is being slashed, I believe we should be doing the opposite AI has show academia is beyond…

ah right, de-found the main source of at lest semi independent scientific advancement just because it has a lot of issues instead of fixing this issues .. always a grate idea

I would have though that the common catastrophic failure stories of fully rewriting systems all at once instead of fixing them bit by bit would help IT people to know better.

Re: Deep learning gets the glory, deep fact checking gets ignored

#153

Earlier quoted context omitted.

> Almost nobody is "anti-science". Last I checked: - 15% of Americans don't believe in Climate Change[0] - 37% believe God created man in our current form within the last ~10k years (i.e. don't believe in evolution)[1] I don't think these are just rounding errors. They're large enough numbers that you should know multiple people who hold these beliefs unless you're in a strong bubble. I'm obviously with you in news a…

I think one of the most important 'social values' for science to thrive is a culture with a freedom to disagree on essentially anything. In most of every era where there was rapid scientific progress from the Greeks to the Islamic Golden Age to the Renaissance and beyond, there was also rich, and often times rather virulent, disagreements over even the most sacred of things. Some of those disagreements were well foun…

  > Disagreeing with some consensus is not "anti-science".
Be careful of gymnastics.

Yes, science requires the ability to disagree. You can even see in my history me saying a scientist needs to be a bit anti authoritarian!

But HOW one goes about disagreeing is critical.

Sometimes I only have a hunch that what others believe is wrong. They have every right to call me stupid for that. Occasionally I'll be able to gather the evidence and prove my hunch. Then they are stupid for not believing like I do, but only after evidenced. Most of the time I'm wrong though. Trying to gather evidence I fail and just support the status quo. So I change my mind.

Most importantly, I just don't have strong opinions about most things. Opinions are unavoidable, strong ones aren't. If I care about my opinion, I must care at least as much about the evidence surrounding my opinion. That's required for science.

Look at it this way. When arguing with someone are you willing to tell them how to change your mind? I will! If you're right, I want to know! But frankly, I find most people are arguing to defend their ego. As if being wrong is something to be embarrassed about. But guess what, we're all wrong. It's all about a matter of degree though. It's less wrong to think the earth is a sphere than flat because a sphere is much closer to an oblate spheroid.

If you can't support your beliefs and if you can't change your mind, I don't care who you listen to, you're not listening to science

Re: Deep learning gets the glory, deep fact checking gets ignored

#154

It’s like fake news is taking in science now. Saying any stupid thing will attract much more view and « likes » than those debunking them. Except that we can’t compare twitter to nature journal. Science is supposed to be immune to these kind of bullshit thanks to reputed journals and pair reviewing, blocking a publication before it does any harm. Was that a failure of nature ?

no it's a long term incoming failure

partially due to legacy of science historically being rooted in "it matters more who you (or your parents) are" societies (due to them having had the money in somewhat modern history) (or like some would say the "old white man problem", except it has nothing to do with skin color, or man and only limited to do with old)

partially due to how much more "science (output)" is produced today and ways which once worked to have reasonable QA don't work that well in todays scale anymore

partially due to how many flows

partially due to human nature (as in people tend to care more about "exiting", "visible" things etc.)

People have been pushing for change in a lot of ways like:

- pushing to make full re-poducability a must have (but that is hard, especially for statistics based things only a few companies can even afford to try to run. But also hard due to it requiring a lot of transparency and open data access, and especially the alter is often very much something many owners of data sets are not okay with.

- pushing for more appreciation of null results, or failures. (To be clear I mean both appreciation in form of monetary support and in the traditional sense of the word of people (colleges) appreciating it).

- pushing for more verifying of papers by trying to reproduce it (both as in more money/time resources for it and in changing the mind set from it being a daunting unappreciated task to it being a nice thing to do)

but to little change happened in the end before modern LLM AI hit the scene and now it has made things so much harder as it's now easy to mass produce slob but reasonable looking (non) sience

Re: Deep learning gets the glory, deep fact checking gets ignored

#155
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

[deleted]

Re: Deep learning gets the glory, deep fact checking gets ignored

#156

Earlier quoted context omitted.

I don’t think that’s quite right. Disagreement for the sake of disagreement is not particularly meaningful. The basis for science is iteration on the scientific method. Which is to say: observe -> hypothesize -> falsify. Anti science means to make claims that have no basis in that process or to categorically reject the body of work that was based on that process.

People disagree because they hold a different opinion. In many eras publicly expressing differing opinions, let alone publicly challenging established ones, becomes difficult for various reasons - cultural, political, social, even economic. And I think this is, in general, the natural state of society. When people think something is right, changing their mind is often not realistically possible. And this includes eve…

  > none other than Einstein rejected a probabilistic interpretation of quantum physics
That has been communicated to you wrong and a subtle distinction makes a world of difference.

Plenty of physicists then and now still work hard on trying to figure out how to remove uncertainty in quantum mechanics. It's important to remember that randomness is a measurement of uncertainty.

We can't move forward if the current paradigm isn't challenged. But the way it is challenged is important. Einstein wasn't going around telling everyone they were wrong, but he was trying to get help in the ways he was trying to solve it. You still have to explain the rest of physics to propose something new.

Challenging ideas is fine, it's even necessary, but at the end of the day you have to pony up.

The public isn't forming opinions about things like Einstein. They just parrot authority. Most HN users don't even understand Schrödinger's cat and think there's a multiverse.

Re: Deep learning gets the glory, deep fact checking gets ignored

#157

Earlier quoted context omitted.

> Almost nobody is "anti-science". Last I checked: - 15% of Americans don't believe in Climate Change[0] - 37% believe God created man in our current form within the last ~10k years (i.e. don't believe in evolution)[1] I don't think these are just rounding errors. They're large enough numbers that you should know multiple people who hold these beliefs unless you're in a strong bubble. I'm obviously with you in news a…

According to your model, scientists who believe in God are anti-science. That's almost weirder than declaring that 15% of people not believing in anthropogenic global warming is some sort of crisis. It's a theory that seems to fit the data (with caveats), not an Axiom of Science. It's actually bizarre that 85% of people trust Science so much that they would believe in something that they have never seen any direct ev…

  > According to your model, scientists who believe in God are anti-science.
In a way, yes. But every scientist I know that also believes in God is not shy in admitting their belief is unscientific.

The reason I'm giving this a bit of a pass is because in science we need things that are falsifiable. The burden of proof should be on those believing in God. But such a belief is not falsifiable. You can't prove or disprove God. If they aren't pushy, they're okay with admitting that, and don't make a big deal out of it then I don't really care. That's just being a decent person.

But that's a very different thing than not believing in things we have strong physical evidence for, strong mathematical theories, and a long record of making counter factual predictions. The great thing about science is it makes predictions. Climate science has been making pretty good ones since the 80's. Every prediction comes with error bounds. Those are tightening but the climate today matches those predictions within error. That's falsifiable

Re: Deep learning gets the glory, deep fact checking gets ignored

#158
post #88

Earlier quoted context omitted.

> The better question is: what is the acceptable threshold? Currently we are unable to answer that question. AND THAT'S THE PROBLEM I'd be fine if we could. Well, at least far less annoyed. I'm not sure what the threshold should be, but we should always try to minimize it. At least error bounds would do a lot of good at making this happen. But right now we have no clue and that's why this is such a big question that…

There‘s also the question of „what is it failing at?“. I‘m fine with 5% failure if my soup is a bit too salty. Not fine with 0.1% failure if it contains poison.

That's a good example. I'm going to steal it ;)

(We can even get more nuanced. What kind of poison?)

Re: Deep learning gets the glory, deep fact checking gets ignored

#159
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head.

I have endash bound to ⇧⌥⌘0, and emdash bound to ⇧⌥⌘=.

Re: Deep learning gets the glory, deep fact checking gets ignored

#160

> although later investigation suggests there may have been data leakage I think this point is often forgotten. Everyone should assume data leakage until it is strongly evidenced otherwise. It is not on the reader/skeptic to prove that there is data leakage, it is the authors who have the burden of proof. It is easy to have data leakage on small datasets. Datasets where you can look at everything. Data leakage is rea…

Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…

Data leakage is an eval problem, not an accuracy problem.

That is, the problem is not that the AI is wrong X% of the time. The problem is that, in the presence of a data leak, there is no way of knowing what the value of X even is.

This problem is recursive - in the presence of a data leak, you also cannot know for sure the quantity of data that has leaked.

Post reply on HN