Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…
Deep learning gets the glory, deep fact checking gets ignored
81–90 of 174 posts
Re: Deep learning gets the glory, deep fact checking gets ignored
#82Before making AI do research, perhaps we should first let it __reproduce__ research. For example, give it a paper of some deep learning technique and make it produce an implementation of that paper. Before it can do that, I have no hope that it can produce novel ideas.
Reproduciblity was never a serious issue in AI research community. I think one of the main reasons for explosive progress in AI was the open community and people could easily reproduce other people's research. If you look at top tier conferences you see that they share everything paper, latex, code, data, lecture video etc. After ChatGPT big cooperations stopped sharing their main research but it still happens at aca…
Re: Deep learning gets the glory, deep fact checking gets ignored
#83Earlier quoted context omitted.
Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…
I think it's worth being highly skeptical about fraud rates that are stated to two decimal places of precision. Fraud is by design hard to accurately detect. It would be more accurate to say, Medicare decides 7.66% of its cases are fraudulent according to its own policies and procedures, which are likely conservative, and cannot take into account undetected fraud. The true rate is likely higher, perhaps much higher.…
CERT’s annual assessments do seem to involve a large-scale, rigorous analysis of an independent sample of 50,000 cases, though. And those case audits seem, at least on paper and to a layperson, to apply rather more thorough scrutiny than Medicare’s day-to-day policies and procedures.
As @patio11 says, and to your point, “the optimal amount of fraud is non-zero”… [2]
[0] https://www.cms.gov/data-research/monitoring-programs/improp...
[1] https://www.cms.gov/research-statistics-data-and-systems/mon...
[2] https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...
Re: Deep learning gets the glory, deep fact checking gets ignored
#84Earlier quoted context omitted.
Yes. And let's not get started on that ML Quantum Wormhole bullshit... We've taken this all too far. It is bad enough to lie to the masses in Pop-Sci articles. But we're straight up doing it in top tier journals. Some are good faith mistakes, but a lot more often they seem like due diligence just wasn't ever done. Both by researchers and reviewers. I at least have to thank the journals. I've hated them for a long tim…
this seems strange to me, shouldn’t we expect a high quality journal to retract often as we gather more information? obviously this is hyperbole of two extremes, but i certainly trust a journal far more if it actively and loudly looks to correct mistakes over one that never corrects anything or buries its retractions. a rather important piece of science is correcting mistakes by gathering and testing new information.…
> shouldn’t we expect a high quality journal to retract often as we gather more information?
This is complicated, and kinda sad tbh. But no.You need to carefully think about what "high quality journal" means. Typically it is based on something called Impact Factor[0]. Impact factor is judged by the number of citations a journal has received in the last 2 years. It sounds good on paper, but I think if you think about it for a second you'll notice there's a positive feedback loop. There's also no incentive that it is actually correct.
For example, a false paper can often get cited far more than a true paper. This is because when you write the academic version of "XYZ is a fucking idiot, and here's why" you cite their paper. It's good to put their bullshit down, but it can also just end up being Streisand effect-like. Journal is happy with its citations. Both people published in them. They benefit from both directions. You keep the bad paper up for the record and because as long as the authors were actually acting in good faith, you don't actually want to take it down. The problem is... how do you know?
Another weird factor used is Acceptance Rates. This again sounds nice at first. You don't want a journal publishing just anything, right?[1] The problem comes when these actually become targets (which they are). Many of the ML conferences target about 25% acceptance rate[2]. It fluctuates year to year. It should, right? Some years are just better science than other years. Good paper hits that changes things and the next year should have a boom! But that's not the level of fluctuation we're talking about. If you look at the actual number of papers accepted in that repo you'll see a disproportionate number of accepted papers ending in a 0 or 5. Then you see the 1 and 6, which is a paper being squeezed in, often for political reasons. Here, I did the first 2 tables for you. You'll see that has a very disproportionate ending of 1 and 6 and CV loves 0,1,3 These numbers should convince you that this is not a random process, though they should not convince you it is all funny business (much harder to prove). But it is at least enough to be suspicious and encourage you to dig in more.
There's a lot that's fucked up about the publishing system and academia. Lots of politics, lots of restricted research directions, lots of stupid. But also don't confuse this for people acting in bad faith or lying. Sure, that happens. But most people are trying to do good and very few people in academia are blatantly publishing bullshit. It's just that everything gets political. And by political I don't mean government politics, I mean the same bullshit office politics. We're not immune from that same bullshit and it happens for exactly the same reasons. It just gets messier because if you think it is hard to measure the output of an employee, try to measure the output of people who's entire job it is to create things that no one has ever thought of before. It's sure going to look like they're doing a whole lot of nothing.
So I'll just leave you with this (it'll explain [1])
As a working scientist, Mervin Kelly (Director of Bell Labs (1925-1959)) understood the golden rule,
"How do you manage genius? You don't."
https://1517.substack.com/p/why-bell-labs-worked
There's more complexity like how we aren't good at pushing out frauds and stuff but if you want that I'll save it for another comment.[0] https://en.wikipedia.org/wiki/Impact_factor
[1] Actually I do. As long as it isn't obviously wrong, plagiarized, or falsified, then I want that published. You did work, you communicated it, now I want it to get out into the public so that it can be peer reviewed. I don't mean a journal's laughable version of peer review (3-4 unpaid people that don't study your niche and are more concerned with if it is "novel" or "impactful" quickly reading your paper and you're one of 4 on their desk they need to do this week. It's incredibly subjective and high impact papers (like Nobel Price winning papers) routinely get rejected). Peer review is the process of other researchers replicating your work, building on it, and/or countering it. Those are just new papers...
[2] https://github.com/lixin4ever/Conference-Acceptance-Rate
Re: Deep learning gets the glory, deep fact checking gets ignored
#85> although later investigation suggests there may have been data leakage I think this point is often forgotten. Everyone should assume data leakage until it is strongly evidenced otherwise. It is not on the reader/skeptic to prove that there is data leakage, it is the authors who have the burden of proof. It is easy to have data leakage on small datasets. Datasets where you can look at everything. Data leakage is rea…
Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…
> The better question is: what is the acceptable threshold?
Currently we are unable to answer that question. AND THAT'S THE PROBLEMI'd be fine if we could. Well, at least far less annoyed. I'm not sure what the threshold should be, but we should always try to minimize it. At least error bounds would do a lot of good at making this happen. But right now we have no clue and that's why this is such a big question that people keep bringing up. We don't point out specific levels of error because they are small and we don't want you looking at them, rather we don't point them out because nobody has a fucking clue.
And until someone has a clue, you shouldn't trust that they error rate is low. The burden of proof is on the one making the claim of performance, not the one asking for evidence to that claim (i.e. skeptics).
Btw, I'd be careful with percentages. Especially when numbers are very high. e.g. LLMs are being trained on trillions of tokens. 10% of 1 trillion is 100 bn. The entire work of Shakespeare is 1.2M tokens... Our 10% error rate would be big enough to spoil any dataset. The bitter truth is that as the absolute number increases, the threshold for acceptable spoilage (in terms of percentage) needs to decrease.
Re: Deep learning gets the glory, deep fact checking gets ignored
#86It's interesting to see this article in juxtaposition to the one shared recently[1], where AI skeptics were labeled as "nuts", and hallucinations were "(more or less) a solved problem". This seems to be exactly the kind of results we would expect from a system that hallucinates, has no semantic understanding of the content, and is little more than a probabilistic text generator. This doesn't mean that it can't be use…
Or even of the Internet in general.
I guess it's a common pitfall with information or communication technologies ?
(Heck, or with technologies in general, but non-information or communication ones rarely scale as explosively...)
Re: Deep learning gets the glory, deep fact checking gets ignored
#87Fantastic article by Rachel Thomas! This is basically another argument that deep learning works only as a [generative] information retrieval - i.e a stochastic parrot, due to the fact that the training data is a very lossy representation of the underlying domain. Because the data/labels of genes do not always represent the underlying domain (biology) perfectly, the output can be false/invalid/nonsensical. in cases wh…
Re: Deep learning gets the glory, deep fact checking gets ignored
#88Earlier quoted context omitted.
Every system has problems. The better question is: what is the acceptable threshold? For an example Medicare and Medicade had a fraud rate of 7.66%. Yes, that is a lot of billions, and there is room for improvement, but that doesn’t mean the entire system is failing: 93% of cases are being covered as intended. The same could be said with these models. If the spoilage rate is 10%, does that mean the whole system is ba…
> The better question is: what is the acceptable threshold? Currently we are unable to answer that question. AND THAT'S THE PROBLEM I'd be fine if we could. Well, at least far less annoyed. I'm not sure what the threshold should be, but we should always try to minimize it. At least error bounds would do a lot of good at making this happen. But right now we have no clue and that's why this is such a big question that…
I‘m fine with 5% failure if my soup is a bit too salty. Not fine with 0.1% failure if it contains poison.
Re: Deep learning gets the glory, deep fact checking gets ignored
#89We also love deep cherry picking. Working hard to find that one awesome time some ML / AI thing worked beautifully and shouting its praises to the high heavens. Nevermind the dozens of other times we tried and failed...
Dude. I just asked my computer to write [ad lib basic utility script] and it spit out a syntactically correct C program that does it with instructions for compiling it. And then I asked it for [ad lib cocktail request] and got back thorough instructions. We did that with sand. That we got from the ground. And taught it to talk. And write C programs. Never mind what? That I had to ask twice? Or five times? What maximu…
But I think people aren‘t arguing about how amazing it is, but about specific applicability. There‘s also a lot of toxic hype and FUD going around, which can be tiring and frustrating.
Re: Deep learning gets the glory, deep fact checking gets ignored
#90I once met a researcher who spent six months verifying the results of a published paper. In the end, all he received was a simple “thanks for pointing that out.” He said quietly, “Some work matters not because it’s seen, but because it keeps others from going wrong.” I believe that if we’re not even willing to carefully confirm whether our predictions match reality, then no matter how impressive the technology looks,…
While that will not land him a Nobel prize, its miles ahead in terms of achievement and added value to mankind compared to most corporate employees. We wish we could say something like that about our past decade of work