Live data from Hacker News

GTP Blind Voting: GPT-5 vs. 4o

gptblindvoting.vercel.app

21–30 of 54 posts

Re: GTP Blind Voting: GPT-5 vs. 4o

#22
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

I'm back on ChatGPT today. The UI is so fast. I didn't realize how the buggy and slow Gemini UI was contributing to my stress levels. AIStudio is also quite slow compared to the ChatGPT app. Is it that hard to make it so that when you paste text into a box and press enter, your computer doesn't slow down and get noisy? Is it really that difficult of an engineering problem?

Re: GTP Blind Voting: GPT-5 vs. 4o

#23

I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".

If the positions are randomly assigned, it shouldn't matter. I mean, the results may be clear faster, but the overall shouldn't change even if you need to flip a coin from time to time.

Re: GTP Blind Voting: GPT-5 vs. 4o

#28
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

Before GPT-5, I've used almost exclusively o3 and sometimes o3-pro. Now I'm using GPT 5 Thinking and sometimes GPT 5 Pro. So I think that I have some control over quality. At least it thinks for few dozens of seconds every time.

> GPT 5 Pro

Have you noticed either of these things:

(1) If your first prompt is too long (50k+ tokens) but just below the limit (like 80k tokens or whatever), it cannot see the right-side of your prompt.

(2) By the second prompt, if the first prompt was long-ish, the context from the first prompt is no longer visible to the model.

Re: GTP Blind Voting: GPT-5 vs. 4o

#29
post #27

Does anyone ever get answers this short? What's the system prompt here? That may bias things a little. Also it's GPT not GTP

Yeah, it felt like two different styles (one very short, the other a little bit more verbose), but both very different from a plain query to GPT without additional system prompts.

So this might just test how the two models react to the (hidden) system prompt...

Re: GTP Blind Voting: GPT-5 vs. 4o

#30
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

What about Claude Opus 4 and 4.1?
Post reply on HN