I don’t understand the negative reactions. Medical care as it exists requires the doctor and patient to have their brains switched on. I’ve almost never had a problem where a doctor provides me with a diagnosis and I go about my day. Most of the times that I have, I’ve been confident about the problem and known what I needed. The doctor was a barrier to accessing care. Dr. GPT is a good brainstorming tool. It helps s…
> I do think that people saying “doctors don’t know the state of the art” have a weaker case. This is kinda the case though. In Poland I met only one psychiatrist that knew about DSM-5. In this year. DSM-5 was a thing from 2013. Doctors are people just as us, not every single of them is good.
I used Claude Code to get a second opinion on my MRI
701–710 of 748 posts
Re: I used Claude Code to get a second opinion on my MRI
#702Earlier quoted context omitted.
Why should we not expect a computer vision model to outperform humans on reading medical images? The human experts are literally just a trained biological neural network. In this domain they are not capable of anything a computer can't already do.
> Why should we not expect a computer vision model to outperform humans on reading medical images? Humans can identify. A computer vision model can return a statistical value. Both can make errors, but these errors are orthogonal to how we work and what is being asked of them. I think a CV model can absolutely provide value as augmentation. Identifying possible misses or a different diagnosis worth considering, but t…
Also disagreement among human radiologists has been documented for decades, so the clean expert baseline you're defending doesn't actually exist outside this argument.
When the identifiers pass human-detection-rate percentages, it will most likely be cheaper to hire a fall-guy for the liability with a much smaller salary, I think this will be a big market in fact.
Re: I used Claude Code to get a second opinion on my MRI
#703Earlier quoted context omitted.
But that’s trivially false. There is an entire category of work where it is hard to come up with an answer and easy to verify the answer, which means that if you verified everything there would still be a large time savings.
I would question whether that holds in the practical LLM automation space. Can you think of any real life examples where an LLM is likely to be used? I think in practice what you're saying is there are problems where there exist efficient deterministic verification methods, and I'm sure that's true. But that's not the bulk of everyday work LLMs are being asked to do nowadays across industry.
But if you want to keep it in the realm of the everyday: you're asking if it is easier to write an email than to read it and check it covers what you wanted to say? Is it easier to search for something or to look at what's been found and say that it's what you were looking for?
Re: I used Claude Code to get a second opinion on my MRI
#704Earlier quoted context omitted.
That no one has actually solved the underlying problem at all, and the generation of the example LLM has no bearing on the nature of the fundamental problem.
You are totally misunderstanding my argument then. As I said, garbage in garbage out. Your article is just an example of that. It’s pretty obvious that if you train an LLM on bad data, you will get bad output. What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part. The fact that all you could come up with was research (not an actual example…
The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data.
It’s also one reason why Google search results have been getting so much worse - it’s hard to not find a SEO page with subtly (or not so subtly) wrong AI slop on almost every topic you can imagine. Most folks won’t recognize it, but that’s what is going on if you know what to look for.
One other way of putting it is the ouroborus problem - more and more internet content is AI generated, because of people trying to game the system, and they are making it is indistinguishable from real content as possible to get by the AI detection algorithms.
Anyone trying to train on it just ends up eating the shit from another LLM, which poisons it.
Another name for it is ‘model collapse’, which also doesn’t have a known solution yet.
Re: I used Claude Code to get a second opinion on my MRI
#705A few years ago (before the AI craze), I was misdiagnosed with tuberculosis. I had a chronic cough, and an outsourced radiologist at a clinic found signs of tuberculosis. The findings were sent to the city's tuberculosis hospital, as required by the country's law. The doctors there took the radiologist's conclusion at face value and required me to stay at their hospital for at least 8 months under a strict, prison-li…
How is it possible? You can't diagnose tuberculosis just based on imaging and tuberculosis hospital has to know that!
Re: I used Claude Code to get a second opinion on my MRI
#706Earlier quoted context omitted.
Why should we not expect a computer vision model to outperform humans on reading medical images? The human experts are literally just a trained biological neural network. In this domain they are not capable of anything a computer can't already do.
> Why should we not expect a computer vision model to outperform humans on reading medical images? Humans can identify. A computer vision model can return a statistical value. Both can make errors, but these errors are orthogonal to how we work and what is being asked of them. I think a CV model can absolutely provide value as augmentation. Identifying possible misses or a different diagnosis worth considering, but t…
I also didn't say anything about whatever Altman or any specific company is doing.
The simple fact is that we send humans to school for years to learn to read and classify these things. It's something computers will be able to do strictly better.
Re: I used Claude Code to get a second opinion on my MRI
#707Earlier quoted context omitted.
You are totally misunderstanding my argument then. As I said, garbage in garbage out. Your article is just an example of that. It’s pretty obvious that if you train an LLM on bad data, you will get bad output. What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part. The fact that all you could come up with was research (not an actual example…
I literally just grabbed a random link. I’ve seen dozens of real life examples of poisoning. The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data. It’s also one reason why Google search results have b…
And do you even realize how much data 0.001% of the training data for a frontier models is? They’re trained on 10s of trillions of tokens, meaning you’d need hundreds of millions of tokens of poisoned data.
Some of these problems you mention could become real barriers to models improvements, though there are plenty of countermeasures, such as by focusing on high quality data sources like I mentioned before.
We’ve already probably gotten as much as we’re ever going to get from simply scraping more and more unstructured text from the web as a way to improve model performance.
The type of training being done now is around tool use and solving specific types of problems better, which is the type of training data you simply don’t find lying around on the web.
Re: I used Claude Code to get a second opinion on my MRI
#708Earlier quoted context omitted.
> What happened to VERIFYING an answer? Does nobody do that anymore? The problem with medical advice is that you may not be competent to verify the answer, right? I agree that asking 5 LLMs to vote and trusting the answer is totally the wrong approach, of course. But LLMs (and traditional material) can help getting more informed. For instance, instead of going to your doctor with the LLM diagnosis and trying to convi…
Or more specifically in this case: the patient was obviously insisting on a diagnosis and treatment based on ... a slightly hurting shoulder, with zero visible or detectable phenomena. So the doctors gave him what he wanted: a treatment ... and Claude told him the treatment was a placebo. Correctly, I might add. Yeah, it is absolutely not what the patient wanted to hear. BAD doctors! Except ... no, not really. Does t…
Re: I used Claude Code to get a second opinion on my MRI
#709Earlier quoted context omitted.
> Why should we not expect a computer vision model to outperform humans on reading medical images? Humans can identify. A computer vision model can return a statistical value. Both can make errors, but these errors are orthogonal to how we work and what is being asked of them. I think a CV model can absolutely provide value as augmentation. Identifying possible misses or a different diagnosis worth considering, but t…
Thank you. I think my comment was careful to specify "in this domain" and "computer vision model"; I didn't say anything about generative AI. The reference to neural networks was hopefully an obvious rhetorical flair, rather than a one-line assertion that computers and brains are actually equivalent. I also didn't say anything about whatever Altman or any specific company is doing. The simple fact is that we send hum…
I don’t think there’s any evidence that’s true.
Re: I used Claude Code to get a second opinion on my MRI
#710Earlier quoted context omitted.
Radiologist who does read shoulder MRI would like to add that over half the annotations are wrong, glaring mistakes in anatomy and cardinal direction which begs the question of how is it making these findings without knowing what it’s looking at (here’s a hint, it’s hallucinated based on reports it sees).
What is "it"? Claude Opus 4.x? ChatGPT-5.x? GLM? DeepSeek? RadFM? Med-PaLM?