GTP Blind Voting: GPT-5 vs. 4o
21–30 of 54 posts
Re: GTP Blind Voting: GPT-5 vs. 4o
#22My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…
Re: GTP Blind Voting: GPT-5 vs. 4o
#23I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".
Re: GTP Blind Voting: GPT-5 vs. 4o
#24Re: GTP Blind Voting: GPT-5 vs. 4o
#25I gravitated choosing the longer answer so my result was a preference for GPT5 responses
Re: GTP Blind Voting: GPT-5 vs. 4o
#26Re: GTP Blind Voting: GPT-5 vs. 4o
#27Also it's GPT not GTP
Re: GTP Blind Voting: GPT-5 vs. 4o
#28My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…
Before GPT-5, I've used almost exclusively o3 and sometimes o3-pro. Now I'm using GPT 5 Thinking and sometimes GPT 5 Pro. So I think that I have some control over quality. At least it thinks for few dozens of seconds every time.
Have you noticed either of these things:
(1) If your first prompt is too long (50k+ tokens) but just below the limit (like 80k tokens or whatever), it cannot see the right-side of your prompt.
(2) By the second prompt, if the first prompt was long-ish, the context from the first prompt is no longer visible to the model.
Re: GTP Blind Voting: GPT-5 vs. 4o
#29Does anyone ever get answers this short? What's the system prompt here? That may bias things a little. Also it's GPT not GTP
So this might just test how the two models react to the (hidden) system prompt...
Re: GTP Blind Voting: GPT-5 vs. 4o
#30My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…