Live data from Hacker News

GTP Blind Voting: GPT-5 vs. 4o

gptblindvoting.vercel.app

11–20 of 54 posts

Re: GTP Blind Voting: GPT-5 vs. 4o

#12
I took the test with 10 questions, and carefully picked the answer with more specificity and unique propositional content, that felt like it was communicating more logic that was worth reading, and also the answers that were just obviously more logical or effective, or framed better. I chose GPT-5 8 out of 10 times.

Re: GTP Blind Voting: GPT-5 vs. 4o

#16
I took the "Rank Models" and got GPT5 and Sonnet 4 tied at 25% each, Gemini and Grok close by and 4o in the dust.

But ... the advice (answers) was quite uniform. In more than a few cases I would personally choose a different approach to all of them.

It'd be fun to have a few Chinese models in the mix and see if the cultural biases show up.

Re: GTP Blind Voting: GPT-5 vs. 4o

#17
I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".

Re: GTP Blind Voting: GPT-5 vs. 4o

#19
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

Before GPT-5, I've used almost exclusively o3 and sometimes o3-pro. Now I'm using GPT 5 Thinking and sometimes GPT 5 Pro. So I think that I have some control over quality. At least it thinks for few dozens of seconds every time.

Re: GTP Blind Voting: GPT-5 vs. 4o

#20
This doesn’t really work as you can tell there’s an underlying prompt telling the model to reply in one or two sentences. That doesn’t seem like a good way to display the strengths and weaknesses of a model except in situations where you want a very short answer.
Post reply on HN