Live data from Hacker News

Deep learning gets the glory, deep fact checking gets ignored

rachel.fast.ai

141–150 of 174 posts

Re: Deep learning gets the glory, deep fact checking gets ignored

#141

Earlier quoted context omitted.

I think one of the most important 'social values' for science to thrive is a culture with a freedom to disagree on essentially anything. In most of every era where there was rapid scientific progress from the Greeks to the Islamic Golden Age to the Renaissance and beyond, there was also rich, and often times rather virulent, disagreements over even the most sacred of things. Some of those disagreements were well foun…

I don’t think that’s quite right. Disagreement for the sake of disagreement is not particularly meaningful. The basis for science is iteration on the scientific method. Which is to say: observe -> hypothesize -> falsify. Anti science means to make claims that have no basis in that process or to categorically reject the body of work that was based on that process.

[deleted]

Re: Deep learning gets the glory, deep fact checking gets ignored

#142

We also love deep cherry picking. Working hard to find that one awesome time some ML / AI thing worked beautifully and shouting its praises to the high heavens. Nevermind the dozens of other times we tried and failed...

Dude. I just asked my computer to write [ad lib basic utility script] and it spit out a syntactically correct C program that does it with instructions for compiling it. And then I asked it for [ad lib cocktail request] and got back thorough instructions. We did that with sand. That we got from the ground. And taught it to talk. And write C programs. Never mind what? That I had to ask twice? Or five times? What maximu…

I don't think it has anything to do with being impressed or not. It's about being careful not to put too much trust in something so fallible. Because it is so amazing, people overestimate where it can be reliably used.

Re: Deep learning gets the glory, deep fact checking gets ignored

#143
post #46

Anyone here still doing verification or reproduction work? Feels like it’s becoming rare, but I find it super valuable.

Verification work is more common than often supposed.

However, it rarely takes the form of explicit replication of the published findings. More commonly, the published work makes a claim, and such a claim leads to further hypotheses (predictions), which others may attempt to demonstrate/veriify.

During this second demonstration/study, the claims of the first study are verified.

Re: Deep learning gets the glory, deep fact checking gets ignored

#144

It’s like fake news is taking in science now. Saying any stupid thing will attract much more view and « likes » than those debunking them. Except that we can’t compare twitter to nature journal. Science is supposed to be immune to these kind of bullshit thanks to reputed journals and pair reviewing, blocking a publication before it does any harm. Was that a failure of nature ?

Have you seen the statistics about high impact journals having higher retraction/unverified rates on papers? The root causes can be argued...but keep that in mind. No single paper is proof. Bodies of work across many labs, independent verification, etc is the actual gold standard.

The purpose of peer review is to check for methodological errors, not to replicate the experiment. With a few exceptions, it can't catch many categories of serious errors.

> higher retraction/unverified

Scientific consensus doesn't advance because a single new ground-breaking claim is made in a prestigious journal. It advances when enough other scientists have built on top of that work.

The current state of science is not 'bleeding edge stuff published in a journal last week'. That bleeding edge stuff might become part of scientific consensus in a month, or year or three, or five - when enough other people build on that work.

Anybody who actually does science understands this.

Unfortunately, people with poor media literacy who only read the headlines don't understand this, and assume that the whole process is all a crock.

Re: Deep learning gets the glory, deep fact checking gets ignored

#145
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

A lot of phones do this automatically when doing double dash -- -> —

Re: Deep learning gets the glory, deep fact checking gets ignored

#146
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

option-shift-minus on a Mac (option-minus for an en dash).

Re: Deep learning gets the glory, deep fact checking gets ignored

#147
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

> Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest. If I gave a classroom of under grad students a multiple choice test where no answers were correct, I can almost guarantee almost all the tests would be filled out. Should GPT and other LLMs refuse to take a test? In my experience it will answer with the closest answer, even if none of the options are even remotely…

I think the issue is the confidence with which it lies to you.

A good analogy would be if someone claimed to be a doctor and when I asked if I should eat lead or tin for my health they said “Tin because it’s good for your complexion”.

Re: Deep learning gets the glory, deep fact checking gets ignored

#148
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

a lot of applications auto convert -- to an em dash

and a bunch of phone/tablet keyboards do so, too

I like em dashes I had considered installing a plugin to reliably turn -- into em dash in the past, if I hadn't discarded that idea you would have seen some in this post ;)

And I think I have seen at lest one spell checking browser plugin which does stuff like that.

Oh and some people use 3rd party interfaces to interact with HN, such which do auto convert consecutive dashes to em dashes.

In the places where I have been using AI from time to time it's also not supper common to use em dashes.

So IMHO "em dash" isn't a tall tell sign for something being AI written.

But then wrt. the OP comment I think you might be right anyway. It's writing style is ... strange. Like taking a writing style from a novel and not any writing style but such which over exaggerates that currently a story is told inside a story. But then fills semantics of a HN comment. Like what you might get if you ask a LLM to "tell a story" for you set of bullet points.

But this opens a question, if the story still comes from a human isn't it fine? Or is it offensive that they didn't just give us compact bullet points?

Putten that aside, there is always the option that the author is just very well read/written, maybe a book author, maybe a hobby author and picked up such a writing style.

Re: Deep learning gets the glory, deep fact checking gets ignored

#149

Earlier quoted context omitted.

I think one of the most important 'social values' for science to thrive is a culture with a freedom to disagree on essentially anything. In most of every era where there was rapid scientific progress from the Greeks to the Islamic Golden Age to the Renaissance and beyond, there was also rich, and often times rather virulent, disagreements over even the most sacred of things. Some of those disagreements were well foun…

I don’t think that’s quite right. Disagreement for the sake of disagreement is not particularly meaningful. The basis for science is iteration on the scientific method. Which is to say: observe -> hypothesize -> falsify. Anti science means to make claims that have no basis in that process or to categorically reject the body of work that was based on that process.

People disagree because they hold a different opinion. In many eras publicly expressing differing opinions, let alone publicly challenging established ones, becomes difficult for various reasons - cultural, political, social, even economic. And I think this is, in general, the natural state of society. When people think something is right, changing their mind is often not realistically possible. And this includes even the greatest of scientists.

For instance none other than Einstein rejected a probabilistic interpretation of quantum physics, the Copenhagen Interpretation, all the way to his death. Many of his most famous quotes like 'God does not play dice with the universe.' or 'Spooky action at a distance.' were essentially sardonic mocking of such an interpretation, the exact one that we hold as the standard today. It was none other than Max Planck that remarked, 'Science advances one funeral at a time' [1], precisely because of this issue.

And so freedom to express, debate, and have 'wrong ideas' in the public mindshare is quite critical, because it may very well be that those wrong ideas are simply the standard of truth tomorrow. But most societies naturally turn against this, because they believe they already know the truth, and fear the possibility of society being misled away from that truth. And so it's quite natural to try to clamp down, implicitly or explicitly, on public dissenting views, especially if they start to gain traction.

[1] - https://en.wikipedia.org/wiki/Planck's_principle

Re: Deep learning gets the glory, deep fact checking gets ignored

#150
post #119

Earlier quoted context omitted.

The difference in fields is key here: AI models are going to have a very different impact in fields where ground truth is available instantly (does the generated code have the expected output?) or takes years of manual verification. (Not a binary -- ground truth is available enough for AI to be useful to lots of programmers.)

> does the generated code have the expected output? That's many times not easy to verify at all ...

you can easily verify a lot like:

- correct syntax

- passes lints

- type checking passes

- fast test suite passes

- full test suite passes

and every time it doesn't you feed it back into the LLM, automatically, in a loop, without your involvement.

The results are often -- sadly -- too good to not slowly start using AI.

I say sadly because IMHO the IT industry has gone somewhere very wrong due to growing too fast, moving too fast and getting so much money so that the companies spear heading them could just throw more people at it instead of fixing underlying issues. There is also a huge diverge between sience about development, programming, application composition etc. (not to be confused with since about idk. data-structures and fundamental algorithms) and what the industry uses, how it advances etc.

Now I think normally the industry would auto correct at some point, but I fear with LLMs we might get even further away from any fundamental improvements, as we find even more ways to still go along and continue the mess we have.

Worse performance of LLM coding is highly dependent on how much very similar languages are represented in it's dataset, so new languages with any breakthrough/huge improvements or similar will work less good with LLMs. If that trend continues that would lock us in with very mid solutions long term.

Post reply on HN