Live data from Hacker News

Deep learning gets the glory, deep fact checking gets ignored

rachel.fast.ai

161–170 of 174 posts

Re: Deep learning gets the glory, deep fact checking gets ignored

#161
post #6

Fantastic article by Rachel Thomas! This is basically another argument that deep learning works only as a [generative] information retrieval - i.e a stochastic parrot, due to the fact that the training data is a very lossy representation of the underlying domain. Because the data/labels of genes do not always represent the underlying domain (biology) perfectly, the output can be false/invalid/nonsensical. in cases wh…

> works only as a [generative] information retrieval but even if we for simplicity of the argument assume that is true without question, LLM still are here to stay Like think about how do junior devs which (in programming) average or less skill work, they "retrieve" the information about how to solve the problem from stack overflow, tutorials etc. So giving all your devs some reasonable well done AI automation tools…

Programming languages are created by humans and the training dataset is complete enough to train LLMs with good results. Most importantly, natural language is the native domain of the programming code.

Whereas in biology, the natural domain is in physical/chemical/biological reactions occuring between organisms and molecules. The laws of interactions are not created by human, but by Creator(tm), and so the training dataset is barely capturing a tiny fraction and richness of the domain and its interactions. Because of this, any model will be inadequate

Re: Deep learning gets the glory, deep fact checking gets ignored

#162
post #66

Man, I’ve been there. Tried throwing BERT at enzyme data once—looked fine in eval, totally flopped in the wild. Classic overfit-on-vibes scenario. Honestly, for straight-up classification? I’d pick SVM or logistic any day. Transformers are cool, but unless your data’s super clean, they just hallucinate confidently. Like giving GPT a multiple-choice test on gibberish—it will pick something, and say it with its chest.…

I’m not sure anyone I know could make an em dash with their keyboard off the top of their head. [meta] Here’s where I wish I could personally flag HN accounts.

The Android client I use, Harmonic, has a shortcut to report a user, although it just prefills an email to hn@ycombinator.com.

Re: Deep learning gets the glory, deep fact checking gets ignored

#163
post #55

Earlier quoted context omitted.

Almost nobody is "anti-science". The source of that labeling and division came from appeals to authority. You must do or believe this because it's "the science." If you don't, or you disagree, then you are anti-science. It has nothing to do with science, but rather people not finding that a sufficient justification for unpopular actions. For instance it's 100% certain that banning sugary drinks would dramatically imp…

Yes, there are a lot of people who are anti-science. As in they do not believe the scientific method is a good way to find truth. There are people today who are rejecting very basic science that was accepted over a century ago.

Science is a great way to model reality. It’s unclear whether science accurately describes reality. It’s also unclear whether science is capable of determining metaphysical truth.

https://plato.stanford.edu/entries/scientific-realism/#WhatS...

https://en.m.wikipedia.org/wiki/Pessimistic_induction

Re: Deep learning gets the glory, deep fact checking gets ignored

#164

Earlier quoted context omitted.

Validation of biological labels easily takes years - in the OP's example it was a 'lucky' (huge!) coincidence that somebody already had spent years on one of the predicted proteins' labels. Nobody is going to stake 3-5 years of their career on validating some random model's predictions.

Just curious, could you expand on what about that process takes years?

Bioinformatically, you could compare your protein with known proteins and infer function from there, but OP's paper is specifically for the use case where we know nothing in our databases.

Time-wise it depends where in the process you start!

Do you know what your target protein even is? I've seen entire PhDs trying to purify a single protein - every protein is purified differently, there are dozens of methods and some work and some don't. If you can purify it you can run a barrage of tests on the protein - is it a kinase, how does it bind to different assays etc. That gives you a fairly broad idea in which area of activity your protein sits.

If you know what it is, you can clone it into your vector like E. coli. Then E. coli will hopefully express it. That's a few weeks/months of work, depending on how much you want to double-check.

You can then use fluorescent tags like GFP to show you where in the cell your protein is located. Is it in the cell-wall? is it in the nucleus? that might give you an indication to function. But you only have the location at this point.

If your protein is in an easily kept organism like mice, you can run knock-out experiments, where you use different approaches to either turn off or delete the gene that produces the protein. That takes a few months too - and chances are nothing in your phenotype will change once the gene is knocked out, protein-protein networks are resilient and there might be another protein jumping in to do the job.

if you have an idea of what your protein does, you can confirm using protein-protein binding studies - I think yeast two-hybrid is still very popular for this? It tests whether two specific proteins - your candidate and another protein - interact or bind.

None of those tests will tell you 'this is definitely a nicotinamide adenine dinucleotide binding protein', every test (and there are many more!) will add one piece to the puzzle.

Edit: of course it gets extra-annoying when all these puzzle pieces contradict each other. In my past life I've done some work with the ndh genes that sit on plant chloroplasts and are lost in orchids and some other groups of plants (including my babies), so it's interesting to see what they actually do and why they can be lost. It's called ndh because it was initially named NADH-dehydrogenase-like, because by sequence it kind of looks like a NADH dehydrogenase.

There's a guy in Japan (Toshiharu Shikanai) who worked on it most of his career and worked out that it certainly is NOT a NADH dehydrogenase and is instead a Fd-dependent plastoquinone reductase. https://www.sciencedirect.com/science/article/pii/S000527281...

Knockout experiments with ndh are annoying because it seems to be only really important in stress conditions - under regular conditions our ndh- plants behaved the same.

Again, this is only one protein, and since it's in chloroplasts it's ultra-common - most likely one of the more abundant proteins on earth (it's not in algae either). And we still call it ndh even though it is a Ferredoxin-plastoquinone reductase.

Re: Deep learning gets the glory, deep fact checking gets ignored

#165

Earlier quoted context omitted.

People disagree because they hold a different opinion. In many eras publicly expressing differing opinions, let alone publicly challenging established ones, becomes difficult for various reasons - cultural, political, social, even economic. And I think this is, in general, the natural state of society. When people think something is right, changing their mind is often not realistically possible. And this includes eve…

> none other than Einstein rejected a probabilistic interpretation of quantum physics That has been communicated to you wrong and a subtle distinction makes a world of difference. Plenty of physicists then and now still work hard on trying to figure out how to remove uncertainty in quantum mechanics. It's important to remember that randomness is a measurement of uncertainty. We can't move forward if the current parad…

A core component of the Copenhagen interpretation is that quantum mechanics is fundamentally indeterministic meaning you are inherently and inescapably left with probabilistic/statistical systems. And yes, Einstein was saying people were wrong while offering no viable alternative. His motivation was purely ideological - he believed in a rational deterministic universe, and the Copenhagen Interpretation didn't fit his worldview.

For instance this is the complete context of his spooky action at a distance quote: "I cannot seriously believe in [the Copenhagen Interpretation] because the theory cannot be reconciled with the idea that physics should represent a reality in time and space, free from spooky action at a distance." Framing things like entanglement as "spooky action at a distance" was obviously being intentionally antagonistic on top of it all as well.

---

And yes, if it wasn't clear by my tone - I believe the West in has gradually entered onto the exact sort of death of science phase I am speaking about. A century ago you had uneducated (formally at least) brothers working as bicycle repairmen pushing forward aerodynamics and building planes in their spare time. Today, as you observe, even people with excessive formal education, access to [relatively] endless resources, endless information, and more - seem to have little ambition in exploiting that, rather than passively consuming it. It goes some way to explaining why some think LLMs might lead to AGI.

Re: Deep learning gets the glory, deep fact checking gets ignored

#167

Earlier quoted context omitted.

A lot of the time, the work we do doesn’t get much recognition, and barely gets seen.But maybe it still helped in some small way. Thinking about that makes it feel a little less disappointing.

Why would you want to help a corporation, rather than being paid? A company's goals aren't the same as yours. The reason it isn't paying you is that its main goal is to make more money than it spends. It doesn't care whether you make more money than you spend.

> A company's goals aren't the same as yours.

That is true for any organisation or any person that's different from you. Companies ain't special here.

> The reason it isn't paying you is that its main goal is to make more money than it spends.

Making money is the main goal of many companies, but not all.

Almost any goals of any organisation or any person can be furthered by having more money rather than less. So everyone has a similar incentive to pay you less. (This includes charities. All else being equal, if you can pay your workers less, you can hand out more free malaria nets.) But as https://news.ycombinator.com/item?id=44179846 points out, they pay you, so that you work for them.

See also https://en.wikipedia.org/wiki/Instrumental_convergence

Re: Deep learning gets the glory, deep fact checking gets ignored

#168
post #83

Earlier quoted context omitted.

I think it's worth being highly skeptical about fraud rates that are stated to two decimal places of precision. Fraud is by design hard to accurately detect. It would be more accurate to say, Medicare decides 7.66% of its cases are fraudulent according to its own policies and procedures, which are likely conservative, and cannot take into account undetected fraud. The true rate is likely higher, perhaps much higher.…

The number seems to come from Medicare’s CERT program [0]. At a hurried glance they seem to have published data right up to present, but their most recent interpretive report I could find with error margins was from 2016. That one [1] put the CIs on those fraud rates in the +/-2% range per subtype and around +/-0.9% overall. Bearing out your point. CERT’s annual assessments do seem to involve a large-scale, rigorous…

That CERT doesn't seem to be looking for fraud, but more like errors in the bureaucracy. They request medical documents and assess them against the regular criteria, but no effort is made to find the sort of fraud criminals would engage in, like fraudulently produced documents for tests that never happened.

Re: Deep learning gets the glory, deep fact checking gets ignored

#169

Earlier quoted context omitted.

I don’t think that’s quite right. Disagreement for the sake of disagreement is not particularly meaningful. The basis for science is iteration on the scientific method. Which is to say: observe -> hypothesize -> falsify. Anti science means to make claims that have no basis in that process or to categorically reject the body of work that was based on that process.

People disagree because they hold a different opinion. In many eras publicly expressing differing opinions, let alone publicly challenging established ones, becomes difficult for various reasons - cultural, political, social, even economic. And I think this is, in general, the natural state of society. When people think something is right, changing their mind is often not realistically possible. And this includes eve…

> Science advances one funeral at a time

I’ve heard this sentiment expressed by several of my friends in academia

It extends to policy as well. I mean look at the average age and tenure of the US Senate.

Or stagnation and disruption of markets.

I want to call this the inertia of incumbency.

Re: Deep learning gets the glory, deep fact checking gets ignored

#170

Earlier quoted context omitted.

> none other than Einstein rejected a probabilistic interpretation of quantum physics That has been communicated to you wrong and a subtle distinction makes a world of difference. Plenty of physicists then and now still work hard on trying to figure out how to remove uncertainty in quantum mechanics. It's important to remember that randomness is a measurement of uncertainty. We can't move forward if the current parad…

A core component of the Copenhagen interpretation is that quantum mechanics is fundamentally indeterministic meaning you are inherently and inescapably left with probabilistic/statistical systems. And yes, Einstein was saying people were wrong while offering no viable alternative. His motivation was purely ideological - he believed in a rational deterministic universe, and the Copenhagen Interpretation didn't fit his…

You still gravely misunderstand what has happened and what the conversation in physics is. I'm not just making shit up or guessing, but I do have a degree in physics. There is much more nuance to this than you'd get from Pop-Sci or even basic classes. So I'll let you in on what physicists are talking about over beers and to one another. What's going on in the papers and between them.

What you seem to misunderstand is that science is not a mechanical process. It is artistic.

You have to find new ideas and you have to challenge conventions. It is about how you challenge these ideas. It is about how you prove you are right. You're right to say that claims are one thing and proofs are another, but this is why Einstein worked on this problem rather than stating it and moving on. There's a big difference.

Let's look at Schrodinger's Cat

I hold the same belief as Einstein and this is also true for most physicists. The cat is EITHER alive or dead. Regardless of our act of checking.

Here's the big problem... and is a big part everybody misses:

A photon is an observer. The cat is an observer. The detector that releases the poison is an observer. The literal particle that is being radiated from the isotope is, you guessed it, an observer. Literally everything is an observer. An interaction necessitates observation.

So when Einstein says "Do you believe the moon only exists when I look at it?" he's talking saying he isn't special. It would be silly to think that that happens. And truth is, he's right! We can be on opposite sides of the planet and you can see the moon while I can't and vise versa. But hey, maybe I don't exist and this is all in your head! So stop arguing with yourself I guess?

MOST physicists believe what Einstein believed.

Most of us don't believe there are these infinite universes spawning to "brute force search" the universe, checking literally every possible path. This infinite multiverse is the same thing as "we all live in a simulation". Instead, we believe that we simply do not have access to this information. That is a VERY different answer. But notice something critical, how do we differentiate the two? How would we differentiate the two? Unfortunately, so far, it looks like that is not possible. But we have reasons to believe one side over another. Given everything we know so far, the universe doesn't like to just needlessly use energy.

So we're presented with two (technically more) options:

  1) There are an infinite number of universes, corresponding to all possible events as would be viewed by all possible observers. Thus, the universe is doing a brute force search on whatever its solution space is.
  2) The information is unavailable
    2 a) We can't access that information
    2 b) We don't know how to access that information
Einstein believed 2b. Most physicists sit in 2, and I'm willing to bet even Max Tegmark believes #2 is right. More people are split between 2a and 2b, but no one is really ruling either option out. Certainly we hope the answer is 2b, but until someone proves that we can't access the information, we're going to be having this debate. Of course, there's actually one more answer that will change the debate: someone proves the answer is unprovable (this actually seems to be the current likely option btw).

We should also believe 2 because we have examples of both and they're quite common.

2a) If you get into actually doing physics work you will see how complex measurements actually are. You can't actually ever directly measure something, it is always through some chain of proxies. Really, the big question with the cat here is trying to come up with a clever solution so this. Maybe we're making false assumptions. Maybe we can detect the sound of the glass cracking when it releases the poison. Maybe the cat always purrs while it is alive. Those would be ways to indirectly determine if the cat is alive or dead. But it is a thought experiment for a reason.

2b) Heat is the great example of information loss. We have time, which forces a one-way computation. We can watch particles float around and progress from time t_0 to time t_n. We know that there was a unique path and a truthful answer to how these particles moved. But if you hand someone only the data for t_0 and t_n they will be unable to tell you what trajectories those particles took! They can only do this probabilistically!

That's doesn't mean all these potential universes exist, that just means we lost the information!

Similarly, the math says a blackhole is a singularity. A point where there's infinite density in an infinitely small space. But this doesn't mean this thing definitely 100% positively unquestionably exists. That's a laughable idea. There's other explanations that have yet to be ruled out. There's other explanations that have yet to be found!

So it must be the math that is "broken". Doesn't mean it is easy to fix, but it needs to be fixed. Our physics models are incomplete. That's okay! We still have work to do.

Think about what you've said. If it were true, no progress would ever happen. Like the universe didn't suddenly change when Newton and Leibniz invented calculus! Obviously our physics models back then were wrong. The question is how wrong. As best as we can tell, we are still converging. There is noise, but you can still converge with noise. Yes, there are big problems with academia today, and they should be pushed back against (I'm not shy about doing this myself if you check my comment history), but that's different than what you're suggesting.

So here's your mistake: you think your information is complete. Really, we have barely scratched the surface here. And go ahead, prove me wrong. There's a multitude on Nobels to be awarded for proving any of these points.

Post reply on HN