Live data from Hacker News

Deep learning gets the glory, deep fact checking gets ignored

rachel.fast.ai

81–90 of 174 posts

Re: Deep learning gets the glory, deep fact checking gets ignored

#81
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

What kind of data did you run this on?

Re: Deep learning gets the glory, deep fact checking gets ignored

#82
post #2

Before making AI do research, perhaps we should first let it __reproduce__ research. For example, give it a paper of some deep learning technique and make it produce an implementation of that paper. Before it can do that, I have no hope that it can produce novel ideas.

Reproduciblity was never a serious issue in AI research community. I think one of the main reasons for explosive progress in AI was the open community and people could easily reproduce other people's research. If you look at top tier conferences you see that they share everything paper, latex, code, data, lecture video etc. After ChatGPT big cooperations stopped sharing their main research but it still happens at aca…

I think what I would rather like to see is the reproduction of results from experiments that the AI didn't see but are well known. Not reproducing AI papers. For example, assuming a human can build it, would an AI, not knowing anything except what was known at the time, be able to design the millikan oil drop experiment? Or would it be able to design an Taylor-Coutte setup for exploring turbulence? Would it be able to design a linear particle accelerator or a triaxial compression experiment? I think an interesting line of reasoning would be to restrict the training data to what was known before a seminal paper was produced. Like take Lorenz atmospheric circulation paper, train an AI on only data that comes from before that paper was published. Does the AI produce the same equations in the paper and the same description of chaos that Lorenz arrived at?

Re: Deep learning gets the glory, deep fact checking gets ignored

#83

Earlier quoted context omitted.

Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…

I think it's worth being highly skeptical about fraud rates that are stated to two decimal places of precision. Fraud is by design hard to accurately detect. It would be more accurate to say, Medicare decides 7.66% of its cases are fraudulent according to its own policies and procedures, which are likely conservative, and cannot take into account undetected fraud. The true rate is likely higher, perhaps much higher.…

The number seems to come from Medicare’s CERT program [0]. At a hurried glance they seem to have published data right up to present, but their most recent interpretive report I could find with error margins was from 2016. That one [1] put the CIs on those fraud rates in the +/-2% range per subtype and around +/-0.9% overall. Bearing out your point.

CERT’s annual assessments do seem to involve a large-scale, rigorous analysis of an independent sample of 50,000 cases, though. And those case audits seem, at least on paper and to a layperson, to apply rather more thorough scrutiny than Medicare’s day-to-day policies and procedures.

As @patio11 says, and to your point, “the optimal amount of fraud is non-zero”… [2]

[0] https://www.cms.gov/data-research/monitoring-programs/improp...

[1] https://www.cms.gov/research-statistics-data-and-systems/mon...

[2] https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...

Re: Deep learning gets the glory, deep fact checking gets ignored

#84
post #24

Earlier quoted context omitted.

Yes. And let's not get started on that ML Quantum Wormhole bullshit... We've taken this all too far. It is bad enough to lie to the masses in Pop-Sci articles. But we're straight up doing it in top tier journals. Some are good faith mistakes, but a lot more often they seem like due diligence just wasn't ever done. Both by researchers and reviewers. I at least have to thank the journals. I've hated them for a long tim…

this seems strange to me, shouldn’t we expect a high quality journal to retract often as we gather more information? obviously this is hyperbole of two extremes, but i certainly trust a journal far more if it actively and loudly looks to correct mistakes over one that never corrects anything or buries its retractions. a rather important piece of science is correcting mistakes by gathering and testing new information.…

  > shouldn’t we expect a high quality journal to retract often as we gather more information?
This is complicated, and kinda sad tbh. But no.

You need to carefully think about what "high quality journal" means. Typically it is based on something called Impact Factor[0]. Impact factor is judged by the number of citations a journal has received in the last 2 years. It sounds good on paper, but I think if you think about it for a second you'll notice there's a positive feedback loop. There's also no incentive that it is actually correct.

For example, a false paper can often get cited far more than a true paper. This is because when you write the academic version of "XYZ is a fucking idiot, and here's why" you cite their paper. It's good to put their bullshit down, but it can also just end up being Streisand effect-like. Journal is happy with its citations. Both people published in them. They benefit from both directions. You keep the bad paper up for the record and because as long as the authors were actually acting in good faith, you don't actually want to take it down. The problem is... how do you know?

Another weird factor used is Acceptance Rates. This again sounds nice at first. You don't want a journal publishing just anything, right?[1] The problem comes when these actually become targets (which they are). Many of the ML conferences target about 25% acceptance rate[2]. It fluctuates year to year. It should, right? Some years are just better science than other years. Good paper hits that changes things and the next year should have a boom! But that's not the level of fluctuation we're talking about. If you look at the actual number of papers accepted in that repo you'll see a disproportionate number of accepted papers ending in a 0 or 5. Then you see the 1 and 6, which is a paper being squeezed in, often for political reasons. Here, I did the first 2 tables for you. You'll see that has a very disproportionate ending of 1 and 6 and CV loves 0,1,3 These numbers should convince you that this is not a random process, though they should not convince you it is all funny business (much harder to prove). But it is at least enough to be suspicious and encourage you to dig in more.

There's a lot that's fucked up about the publishing system and academia. Lots of politics, lots of restricted research directions, lots of stupid. But also don't confuse this for people acting in bad faith or lying. Sure, that happens. But most people are trying to do good and very few people in academia are blatantly publishing bullshit. It's just that everything gets political. And by political I don't mean government politics, I mean the same bullshit office politics. We're not immune from that same bullshit and it happens for exactly the same reasons. It just gets messier because if you think it is hard to measure the output of an employee, try to measure the output of people who's entire job it is to create things that no one has ever thought of before. It's sure going to look like they're doing a whole lot of nothing.

So I'll just leave you with this (it'll explain [1])

  As a working scientist, Mervin Kelly (Director of Bell Labs (1925-1959)) understood the golden rule, 
  "How do you manage genius? You don't."
  https://1517.substack.com/p/why-bell-labs-worked
There's more complexity like how we aren't good at pushing out frauds and stuff but if you want that I'll save it for another comment.

[0] https://en.wikipedia.org/wiki/Impact_factor

[1] Actually I do. As long as it isn't obviously wrong, plagiarized, or falsified, then I want that published. You did work, you communicated it, now I want it to get out into the public so that it can be peer reviewed. I don't mean a journal's laughable version of peer review (3-4 unpaid people that don't study your niche and are more concerned with if it is "novel" or "impactful" quickly reading your paper and you're one of 4 on their desk they need to do this week. It's incredibly subjective and high impact papers (like Nobel Price winning papers) routinely get rejected). Peer review is the process of other researchers replicating your work, building on it, and/or countering it. Those are just new papers...

[2] https://github.com/lixin4ever/Conference-Acceptance-Rate

[3] https://imgur.com/a/YefOcuA

Re: Deep learning gets the glory, deep fact checking gets ignored

#85

> although later investigation suggests there may have been data leakage I think this point is often forgotten. Everyone should assume data leakage until it is strongly evidenced otherwise. It is not on the reader/skeptic to prove that there is data leakage, it is the authors who have the burden of proof. It is easy to have data leakage on small datasets. Datasets where you can look at everything. Data leakage is rea…

Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…

  > The better question is: what is the acceptable threshold?
Currently we are unable to answer that question. AND THAT'S THE PROBLEM

I'd be fine if we could. Well, at least far less annoyed. I'm not sure what the threshold should be, but we should always try to minimize it. At least error bounds would do a lot of good at making this happen. But right now we have no clue and that's why this is such a big question that people keep bringing up. We don't point out specific levels of error because they are small and we don't want you looking at them, rather we don't point them out because nobody has a fucking clue.

And until someone has a clue, you shouldn't trust that they error rate is low. The burden of proof is on the one making the claim of performance, not the one asking for evidence to that claim (i.e. skeptics).

Btw, I'd be careful with percentages. Especially when numbers are very high. e.g. LLMs are being trained on trillions of tokens. 10% of 1 trillion is 100 bn. The entire work of Shakespeare is 1.2M tokens... Our 10% error rate would be big enough to spoil any dataset. The bitter truth is that as the absolute number increases, the threshold for acceptable spoilage (in terms of percentage) needs to decrease.

Re: Deep learning gets the glory, deep fact checking gets ignored

#86
post #76

It's interesting to see this article in juxtaposition to the one shared recently[1], where AI skeptics were labeled as "nuts", and hallucinations were "(more or less) a solved problem". This seems to be exactly the kind of results we would expect from a system that hallucinates, has no semantic understanding of the content, and is little more than a probabilistic text generator. This doesn't mean that it can't be use…

Heh, reminds me of cryptocurrencies...

Or even of the Internet in general.

I guess it's a common pitfall with information or communication technologies ?

(Heck, or with technologies in general, but non-information or communication ones rarely scale as explosively...)

Re: Deep learning gets the glory, deep fact checking gets ignored

#87
post #6

Fantastic article by Rachel Thomas! This is basically another argument that deep learning works only as a [generative] information retrieval - i.e a stochastic parrot, due to the fact that the training data is a very lossy representation of the underlying domain. Because the data/labels of genes do not always represent the underlying domain (biology) perfectly, the output can be false/invalid/nonsensical. in cases wh…

I wonder to what extent the thought processes that lead to the situation described by Rachel Thomas, are active in other areas. Important article by the way, I agree!

Re: Deep learning gets the glory, deep fact checking gets ignored

#88

Earlier quoted context omitted.

Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…

> The better question is: what is the acceptable threshold? Currently we are unable to answer that question. AND THAT'S THE PROBLEM I'd be fine if we could. Well, at least far less annoyed. I'm not sure what the threshold should be, but we should always try to minimize it. At least error bounds would do a lot of good at making this happen. But right now we have no clue and that's why this is such a big question that…

There‘s also the question of „what is it failing at?“.

I‘m fine with 5% failure if my soup is a bit too salty. Not fine with 0.1% failure if it contains poison.

Re: Deep learning gets the glory, deep fact checking gets ignored

#89

We also love deep cherry picking. Working hard to find that one awesome time some ML / AI thing worked beautifully and shouting its praises to the high heavens. Nevermind the dozens of other times we tried and failed...

Dude. I just asked my computer to write [ad lib basic utility script] and it spit out a syntactically correct C program that does it with instructions for compiling it. And then I asked it for [ad lib cocktail request] and got back thorough instructions. We did that with sand. That we got from the ground. And taught it to talk. And write C programs. Never mind what? That I had to ask twice? Or five times? What maximu…

First off all, I appreciate your comment. Yes, it‘s fucking amazing. (I usually imagine it being „light“ and not „sand“ though. „Sand“ is much more poignant!)

But I think people aren‘t arguing about how amazing it is, but about specific applicability. There‘s also a lot of toxic hype and FUD going around, which can be tiring and frustrating.

Re: Deep learning gets the glory, deep fact checking gets ignored

#90
post #60

I once met a researcher who spent six months verifying the results of a published paper. In the end, all he received was a simple “thanks for pointing that out.” He said quietly, “Some work matters not because it’s seen, but because it keeps others from going wrong.” I believe that if we’re not even willing to carefully confirm whether our predictions match reality, then no matter how impressive the technology looks,…

While that will not land him a Nobel prize, its miles ahead in terms of achievement and added value to mankind compared to most corporate employees. We wish we could say something like that about our past decade of work

People typically get paid as a thank you for their corporate work, not just a lukewarm 'thanks for pointing that out.'
Post reply on HN