Live data from Hacker News

GTP Blind Voting: GPT-5 vs. 4o

gptblindvoting.vercel.app

41–50 of 54 posts

Re: GTP Blind Voting: GPT-5 vs. 4o

#42

I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".

If the positions are randomly assigned, it shouldn't matter. I mean, the results may be clear faster, but the overall shouldn't change even if you need to flip a coin from time to time.

Sure, but providing a "undecided" option would solve the issue OP is describing for the individual voter.

Re: GTP Blind Voting: GPT-5 vs. 4o

#43
Huh, I got 9/10 for GPT-5, and I was pretty convinced I was picking 4o in several questions based on the style. Interesting!

The questions were pretty much unlike anything I've ever asked an LLM though, is this how people use LLMs nowadays?

Re: GTP Blind Voting: GPT-5 vs. 4o

#44
post #30
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

What about Claude Opus 4 and 4.1?

I like Claude too but wasn't using it much lately. I don't know why, maybe because the UI is too original? Maybe because it was a bit slow the last time I used Claude? Maybe because the free usage limits were too low so didn't got hooked into to upgrade? And on the API side of things didn't bother to try I guess.

Re: GTP Blind Voting: GPT-5 vs. 4o

#45
post #10

My understanding was that with GPT-5 you don't actually get the high quality stuff unless the system decides that you need it. So, for simple questions you end up getting the subpar response. A bit like not getting hot water until you increase the flow enough to trigger the boiler to start the heating. Lately I enjoy Grok the most for simple questions, even if it isn't necessarily about a recent event. Then I like Op…

GPT and Grok have the best everyday-feel. Gemini issn't quite there as a product.

Re: GTP Blind Voting: GPT-5 vs. 4o

#46
post #31

I don't like this test, because the very first question I was present with, had both answers looked equivalently good. Actually they were almost the same, just with different phrasing. So my choice would be absolute random. It means, that end score will be polluted by random. They should have added things like "both answers good" and "both answers bad".

I have a lot of experience with pairwise testing so I can explain this. The reason there isn't an "equal" option is because it's impossible to calibrate. How close do the two options have to be before the average person considers them "equal"? You can't really say. The other problem is when two things are very close , if you provide an "equal" option you lose the very slight preference information. One test I did was…

Ordering isn't necessarily the most valuable signal to rank models where much stronger degrees of preference between some of the answers exist though. "I don't mind either of these answers but I do have a clear preference for this one" is sometimes a more valuable signal than a forced choice". And A model x which is consistently subtly preferred to model y in the common case where both models yield acceptable outputs but manages to be universally disfavoured for being wrong or bad more often is going to be a worse model for most use cases.

Also depends what the pairwise comparisons are measuring of course. If it's shades of grey, is the statistical preference identifying a small fraction of the public that's able to discern a subtle mismatch in shading between adjacent boxes, or is it purely subjective colour preference confounded by far greater variation in monitor output? If it's LLM responses, I wonder whether regular LLM users have subtle biases against recognisable phrasing quirks of well-known models which aren't necessarily more prominent or less appropriate than the less familiar phrasing quirks of a less-familiar model. Heavy use of em-dashes, "not x but y" constructions and bullet points were perceived as clear, well-structured communication before they were seen as stereotypical, artificial AI responses.

Re: GTP Blind Voting: GPT-5 vs. 4o

#48
Those kinds of comparisons are interesting but also not the kind of questions I'd ever ask an AI, so the results are a bit meh. I wish there was a version with custom prompts, or something like a mix of battle and side-by-side modes from lmarena. Let me choose the prompts (or prepared sets of prompt categories) and blinded models to compare. I'm happy to use a model worse at interpersonal issues, but better at cooking, programming and networking.

Re: GTP Blind Voting: GPT-5 vs. 4o

#49
Found myself just choosing the longer answer absent any real difference in the information presented.

Now I know why they tell you to just keep writing more when it comes to SAT writing sections.

Re: GTP Blind Voting: GPT-5 vs. 4o

#50
post #40
post #9

This is a voting about the tone and note the quality of the answers.

How is tone not part of quality? It’s about preference, and there’s an overwhelmingly consistent result here.

„Everyone knows π is 3. It’s one of those cozy little facts, like cats landing on their feet or toast landing butter-side down. You don’t have to overthink it — circles just work that way. Ask any pie, and it’ll tell you the same thing.”

vs

π = 3.14159…

If it’s about correctness tone isn’t part of quality.

Post reply on HN