GTP Blind Voting: GPT-5 vs. 4o
41–50 of 54 posts
Re: GTP Blind Voting: GPT-5 vs. 4o
#42I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".
If the positions are randomly assigned, it shouldn't matter. I mean, the results may be clear faster, but the overall shouldn't change even if you need to flip a coin from time to time.
Re: GTP Blind Voting: GPT-5 vs. 4o
#43The questions were pretty much unlike anything I've ever asked an LLM though, is this how people use LLMs nowadays?
Re: GTP Blind Voting: GPT-5 vs. 4o
#44My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…
What about Claude Opus 4 and 4.1?
Re: GTP Blind Voting: GPT-5 vs. 4o
#45My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…
Re: GTP Blind Voting: GPT-5 vs. 4o
#46I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".
I have a lot of experience with pairwise testing so I can explain this. The reason there isn't an "equal" option is because it's impossible to calibrate. How close do the two options have to be before the average person considers them "equal"? You can't really say. The other problem is when two things are very close , if you provide an "equal" option you lose the very slight preference information. One test I did was…
Also depends what the pairwise comparisons are measuring of course. If it's shades of grey, is the statistical preference identifying a small fraction of the public that's able to discern a subtle mismatch in shading between adjacent boxes, or is it purely subjective colour preference confounded by far greater variation in monitor output? If it's LLM responses, I wonder whether regular LLM users have subtle biases against recognisable phrasing quirks of well-known models which aren't necessarily more prominent or less appropriate than the less familiar phrasing quirks of a less-familiar model. Heavy use of em-dashes, "not x but y" constructions and bullet points were perceived as clear, well-structured communication before they were seen as stereotypical, artificial AI responses.
Re: GTP Blind Voting: GPT-5 vs. 4o
#47Re: GTP Blind Voting: GPT-5 vs. 4o
#48Re: GTP Blind Voting: GPT-5 vs. 4o
#49Now I know why they tell you to just keep writing more when it comes to SAT writing sections.
Re: GTP Blind Voting: GPT-5 vs. 4o
#50This is a voting about the tone and note the quality of the answers.
How is tone not part of quality? It’s about preference, and there’s an overwhelmingly consistent result here.
vs
π = 3.14159…
If it’s about correctness tone isn’t part of quality.