Live data from Hacker News

From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

fertrevino.com

11–20 of 102 posts

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#11
Feels like a mixed bag vs regression?

eg - GPT-5 beats GPT-4 on factual recall + reasoning (HeadQA, Medbullets, MedCalc).

But then slips on structured queries (EHRSQL), fairness (RaceBias), evidence QA (PubMedQA).

Hallucination resistance better but only modestly.

Latency seems uneven (maybe more testing?) faster on long tasks, slower on short ones.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#13
post #6

Did you try it with high reasoning effort?

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head:

“Did you try running it over and over until you got the results you wanted?”

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#14
post #13
post #6

Did you try it with high reasoning effort?

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”

What you describe is a person selecting the best results, but if you can get better results one shot with that option enabled, it’s worth testing and reporting results.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#15
post #13
post #6

Did you try it with high reasoning effort?

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”

This is not a good analogy because reasoning models are not choosing the best from a set of attempts based on knowledge of the correct answer. It really is more like what it sounds like: “did you think about it longer until you ruled out various doubts and became more confident?” Of course nobody knows quite why directing more computation in this way makes them better, and nobody seems to take the reasoning trace too seriously as a record of what is happening. But it is clear that it works!

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#16
post #13

Earlier quoted context omitted.

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”

What you describe is a person selecting the best results, but if you can get better results one shot with that option enabled, it’s worth testing and reporting results.

I get that. But then if that option doesn't help, what I've seen is that the next followup is inevitably "have you tried doing/prompting x instead of y"

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#17
post #11

Feels like a mixed bag vs regression? eg - GPT-5 beats GPT-4 on factual recall + reasoning (HeadQA, Medbullets, MedCalc). But then slips on structured queries (EHRSQL), fairness (RaceBias), evidence QA (PubMedQA). Hallucination resistance better but only modestly. Latency seems uneven (maybe more testing?) faster on long tasks, slower on short ones.

Definitely seems like GPT5 is a very incremental improvement. Not what you’d expect if AGI were imminent.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#18
post #15
post #13

Earlier quoted context omitted.

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”

This is not a good analogy because reasoning models are not choosing the best from a set of attempts based on knowledge of the correct answer. It really is more like what it sounds like: “did you think about it longer until you ruled out various doubts and became more confident?” Of course nobody knows quite why directing more computation in this way makes them better, and nobody seems to take the reasoning trace too…

> Of course nobody knows quite why directing more computation in this way makes them better, and nobody seems to take the reasoning trace too seriously as a record of what is happening. But it is clear that it works!

One thing it's hard to wrap my head around is that we are giving more and more trust to something we don't understand with the assumption (often unchecked) that it just works. Basically your refrain is used to justify all sorts of odd setup of AIs, agents, etc.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#19
post #16

Earlier quoted context omitted.

What you describe is a person selecting the best results, but if you can get better results one shot with that option enabled, it’s worth testing and reporting results.

I get that. But then if that option doesn't help, what I've seen is that the next followup is inevitably "have you tried doing/prompting x instead of y"

> I get that. But then if that option doesn't help, what I've seen is that the next followup is inevitably "have you tried doing/prompting x instead of y"

Maybe I’m misunderstanding, but it sounds like you’re framing a completely normal proces (try, fail, adjust) as if it’s unreasonable?

In reality, when something doesn’t work, it would seem to me that the obvious next step is to adapt and try again. This does not seem like a radical approach but instead seems to largely be how problem solving sort of works?

For example, when I was a kid trying to push start my motorcycle, it wouldn’t fire no matter what I did. Someone suggested a simple tweak, try a different gear. I did, and instantly the bike roared to life. What I was doing wasn’t wrong, it just needed a slight adjustment to get the result I was after.

Post reply on HN