Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

461–469 of 469 posts

Re: “Car Wash” test with 53 models

#461

What do you know, the human results line up exactly with ChatGPT. What are the odds! Surely the human responders are highly ethical individuals and they wouldn't even dream of copy-pasting all the questions into ChatGPT without reading them. Realistically, this mostly tells me that the "human answers" service is dead. People will figure out a way to pass the work off to an AI, regardless of quality, as long as they c…

Yea funny coincidence, but this is not at all how the human answers were collected. Rapidata answered this in another comment below. They integrate micro-surveys into mobile apps (like Duolingo, games, etc) as an optional opt-in instead of watching ads. The users are vetted and there's no incentive to answer correctly.

But, there is a clear incentive to answer the question incorrectly. The wrong answer is funny and will give the human some level of pleasure thinking about it. I would certainly reply with "walk" just for fun and apparently 28.5% of people agree with me.

Re: “Car Wash” test with 53 models

#462

Earlier quoted context omitted.

Yea funny coincidence, but this is not at all how the human answers were collected. Rapidata answered this in another comment below. They integrate micro-surveys into mobile apps (like Duolingo, games, etc) as an optional opt-in instead of watching ads. The users are vetted and there's no incentive to answer correctly.

> there's no incentive to answer correctly Answering correctly is not in question here. This is essentially opinion polling anyway, there is no single correct answer. The incentive is exactly what you said: to skip ads. How are the users actually vetted? We have no information on this, just have to take rapidata on faith.

> there is no single correct answer

I think we all mostly agree that there is a single correct answer, and that is why this discussion exists in the first place.

Re: “Car Wash” test with 53 models

#463

Earlier quoted context omitted.

Testing some subset X does not mean the test is rigged unless they failed to disclose that. But also: GPT 5.2 Thinking, Standard Effort: Walk - https://chatgpt.com/share/699d38cb-e560-8012-8986-d27428de8a... I'm assuming "GPT 5.2 Thinking" is, in fact, a thinking model?

The problem is you haven't used the API, but you have used your ChatGPT subscriptions with personality, memories and possible customization. I can see for instance that your ChatGPT answers with emojis, while my ChatGPT subscription never does. If you ask GPT 5.2 with high reasoning efforts in the API, you get 10 out of 10: drive.

If it doesn't work at all using the most popular pricing plans (subscription), AND it doesn't work on the most popular way of accessing it (web), then it seems fair to say there's a problem.

And the problem is NOT that I'm using a product in the advertised, intended way.

Re: “Car Wash” test with 53 models

#464

Earlier quoted context omitted.

Oh wow, Sonnet still isn't handling it well: Opus 4.6: Drive ( https://claude.ai/share/d57fef01-df32-41f2-b1dc-07de7916bdc7 ) Opus 4.5: Drive ( https://claude.ai/chat/a590cac1-100a-490b-b0a2-df6676e1ae99 ) Opus 3.0: Walk ( https://claude.ai/chat/372c144c-d6eb-43f5-b7ea-fd4c51c681db ) Sonnet 4.6: Walk ( https://claude.ai/share/1f2a80f3-4741-40a5-8a05-7349ea1a17e5 ) Sonnet 4.5: Walk ( https://claude.ai/share/905afeb6-f…

This is because it is without thinking enabled. Of course the results are disappointing.

It seems entirely fair to evaluate a product based on the baseline that the company itself offers.

Re: “Car Wash” test with 53 models

#465
post #174

Earlier quoted context omitted.

This. The following question is likely to fool a lot of people, too. "I have a rooster named Pat. (Lots of other details so you're likely to forget Pat is a rooster, not a hen). Pat flies to the top of the roof and lays an egg right on the ridge of the roof. Which way will the egg roll?" But if you omit the details designed to confuse people, they're far less likely to get it wrong: "I have a rooster named Pat. Pat f…

Very problematic to think that something's reproductive attributes have to correspond to what gendered noun we call it by.

Tell me you've never done any farming in your life without telling me you've never done any farming in your life. The difference between male and female animals matters, a lot, to farmers (or ranchers). There's a reason the English language has the words cow and bull, sow and boar, ewe and ram, rooster and hen, nanny and billy, mare and stallion, and many more (and has had those words for centuries). And that reason is precisely because of how mammal (and avian) reproduction works. A cow can't do a bull's job, nor vice-versa, if you want to have calves next year, and grow the size of your herd (or sell the extra animals for income). And so, centuries ago, English-speaking farmers who didn't want to spend the extra syllables on words like "male cattle" and "female cattle" came up with handy, short words (one-syllable words for most species, though not goats and horses) to express those distinctions. Because as I mentioned, they matter a lot when you're raising animals.

Re: “Car Wash” test with 53 models

#466
post #459

Earlier quoted context omitted.

Not sure what you mean by this.

I encourage you to ask a questions so I can figure out what do you not understand. Let me also simplify my comment: “100 minutes” is not the correct answer to that question.

I'm not getting what you're trying to convey.

Re: “Car Wash” test with 53 models

#467
post #465

Earlier quoted context omitted.

Very problematic to think that something's reproductive attributes have to correspond to what gendered noun we call it by.

Tell me you've never done any farming in your life without telling me you've never done any farming in your life. The difference between male and female animals matters , a lot , to farmers (or ranchers). There's a reason the English language has the words cow and bull, sow and boar, ewe and ram, rooster and hen, nanny and billy, mare and stallion, and many more (and has had those words for centuries). And that reaso…

Some roosters lay eggs.

You might believe there is intrinsic sexual dimorphism among mammals and birds. You might even have overwhelming experimental and scientific evidence that proves it. But ask yourself: is it worth losing your job over?

Some roosters lay eggs.

Re: “Car Wash” test with 53 models

#468
post #249

Earlier quoted context omitted.

Hmm have not tested but a spark plug doesn't really need shop tools to be replaced; maybe trying with a way bigger repair like "I need my transmission replaced" would bring different results?

Replacing a spark plug requires a spark plug socket, which is a specialty tool that is generally only found in an automotive shop.

I'm guessing my car is old enough that is comes with a spark plug socket in the toolbag in the back along with the jack and spare wheel; you're right it probably isn't standard equipment anymore. (Car is Mazda from 2005 for reference)

Re: “Car Wash” test with 53 models

#469
post #416

Earlier quoted context omitted.

that's what the cultivators of these examples are preying on. but in practice what people care about is "can i get it to do ", not "is it a decider on every possible token sequence that humans perceive to be about ".

But what is being pitched as "AGI" hype is the latter.

none of what we are using today is even remotely being pitched as AGI. if anything, the foundation model makers go out of their way to pitch the opposite. this is a thing made up entirely in your head, and then you put it on others and then claim it was their doing.
Post reply on HN