Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

221–230 of 469 posts

Re: “Car Wash” test with 53 models

#221
post #219

This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.

For me it's interesting because no normal person I know would ever inject "because its better for the environment" in anything so small scale so not only it shows they suck, it shows how easy it is to inject side-ideology into simple exchanges.

Re: “Car Wash” test with 53 models

#222

Earlier quoted context omitted.

Source?

I don't know, but model names such as "kimi-k2-thinking" in the test set might offset a clue.

Yes, there are some exceptions where it clearly states that a thinking model has been chosen like for kimi, but there is no such indicator for the GPT family from OpenAI and other major models.

Re: “Car Wash” test with 53 models

#223
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

An intelligent mind should keep a balance.

U: The car wash is 50m away. Should I walk or should I drive?

A: Do you need to wash your car?

U: Yep.

A: Is the car yours?

U: Yep.

A: Is your car nearby?

U: Yes.

A: Do you need to walk to your car?

U: No.

A: Do you know if the car wash is currently working?

U: Yes.

A: Does it work?

U: YES!

A: Do you have a way to pay for the wash?

U: YES!

A: On you?

U: YES!

A: Is your clothing appropriate for the occasion?

U: Argh...

Re: “Car Wash” test with 53 models

#224

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

For all we know 30 percent of users on that site are now bots.

What incentive does the Rapiddata website have to even use any humans? How would customers even know?

Re: “Car Wash” test with 53 models

#227

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.

If walking is an option, then your car is already at the car wash. If your car was not at the car wash, then this wouldn't be a question

Re: “Car Wash” test with 53 models

#229

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

For all we know 30 percent of users on that site are now bots.

[flagged]
Post reply on HN