Live data from Hacker News

From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

fertrevino.com

31–40 of 102 posts

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#31
post #13
post #6

Did you try it with high reasoning effort?

Sorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”

Or...

"Did you try a room full of chimpanzees with typewriters?"

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#32

Here's my experience: for some coding tasks where GPT 4.1, Claude Sonnet 4, Gemini 2.5 Pro were just spinning for hours and hours and getting nowhere, GPT 5 just did the job without a fuss. So, I switched immediately to GPT 5, and never looked back. Or at least I never looked back until I found out that my company has some Copilot limits for premium models and I blew through the limit. So now I keep my context small,…

its possible to use gpt-5-high on the plus plan with codex-cli, its a whole different beast! i dont think theres any other way for plus users to leverage gpt-5 with high reasoning.

codex -m gpt-5 model_reasoning_effort="high"

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#33

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

The latter. When "understand", "reason", "think", "feel", "believe", and any of a long list of similar words are in any title, it immediately makes me think the author already drank the kool aid.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#34

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

Do you yourself really understand, or are you just depolarizing neurons that have reached their threshold?

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#35
post #34

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

Do you yourself really understand, or are you just depolarizing neurons that have reached their threshold?

He doesn't know the answer to that and neither do you.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#36
post #34

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

Do you yourself really understand, or are you just depolarizing neurons that have reached their threshold?

[flagged]

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#37

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

What does understanding mean? Is there a sensible model for it? If not, we can only judge in the same way that we judge humans: by conducting examinations and determining whether the correct conclusions were reached.

Probabilities have nothing to do with it; by any appropriate definition, there exist statistical models that exhibit "understanding" and "reasoning".

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#38

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

The latter. When "understand", "reason", "think", "feel", "believe", and any of a long list of similar words are in any title, it immediately makes me think the author already drank the kool aid.

In the context of coding agents, they do simulate “reasoning” when you feed them the output and it is able to correct itself.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#39

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

Does it matter from a practical point of view? It's either true understanding or it's something else that's similar enough to share the same name.

Re: From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]

#40

I have an issue with the words "understanding", "reasoning", etc when talking about LLMs. Are they really understanding, or putting out a stream of probabilities?

The latter. When "understand", "reason", "think", "feel", "believe", and any of a long list of similar words are in any title, it immediately makes me think the author already drank the kool aid.

I agree with “feel” and “believe” but what words would you suggest instead of “understand” and “reason’?
Post reply on HN