Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

441–450 of 469 posts

Re: “Car Wash” test with 53 models

#441

Earlier quoted context omitted.

The systems recognized the pattern that it looks like a generic article on the internet asking whether someone should walk or drive and answered it exactly as expected based on their training data. None of this should be surprising. We are the ones fooling ourselves into believing there's more intelligence in these systems than they really have. At the end of the day, it's just an impressive parlor trick.

In that sense the google AI summary search results are a better UX for this type experience

The better UX is that the google ai search summary is easy to ignore.

Re: “Car Wash” test with 53 models

#443
post #348
post #211

Earlier quoted context omitted.

You passed the intelligence check and failed the wisdom one. The key technique in the mathematical method to answer the machine question is "theory of mind".

It's not theory of mind, it's an understanding of how trick questions are structured and how to answer one. Pretty useless knowledge after high school - no wonder AI companies didn't bother training their models for that

It's not a trick question. It has a simple answer. It's literally impossible to specify a question about real world objects without some degree of prior knowledge about both the contents of the question and the expectation of the questioner coming into play.

The obvious answer here is 100 minutes because it's impossible to perfectly encapsulate every real life factor. What happens if a gamma ray burst destroys the machines? What happens if the machine operators go on strike? Etc, etc. The answer is 100.

Re: “Car Wash” test with 53 models

#444
post #211

Earlier quoted context omitted.

You passed the intelligence check and failed the wisdom one. The key technique in the mathematical method to answer the machine question is "theory of mind".

Theory of mind won’t help you answering this question. It is obviously an underspecified question (at least in any contexts where you are not actively designing/thinking about some specific industrial process). As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick. In the first case it is better to ask clarif…

>As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick.

Context would be key here. If this were a question on a grade school word problem test then just say 100, as it is as specified as it needs to be. If it's a Facebook post that says "We asked 1000 people this and only 1 got it right!" then it's probably some trick question.

If you think it's not specified enough for a grade school question, then I would challenge you to come up with a version that's specified rigorously enough for any sufficiently picky interviewee. (Hint: This is not possible)

>There is nothing “mathematical” about any of this though.

Finding the correct approach to solve a problem specified in English is a mathematical skill.

Re: “Car Wash” test with 53 models

#446
For ambiguous or intricate prompts, the immediate response protocol should be a clarifying question: 'Are you looking for A, B, C, or something else?' Tokens and advanced reasoning capabilities should be reserved until the user provides clarification. A benchmark score should reflect the quality of the conversation as a whole, rather than isolated responses.

Re: “Car Wash” test with 53 models

#447
post #184

Earlier quoted context omitted.

Same. I usually add a "Be curt" in front of every prompt in Gemini.

Is that more effective than simply adding it to your user instructions?

No you’re correct but I’ve experienced a bug with older Workspace business accounts where you can’t reach the screen for user instructions. It just remained blank.

Re: “Car Wash” test with 53 models

#448

Go ask 53 Americans. I’m willing to bet less than 11 get it right.

Don't bet too much, from the linked article ... They ran the exact same question with the same forced choice between "drive" and > "walk," no additional context, past 10,000 real people through their human feedback platform. 71.5% said drive.

... Still shockingly low.

Re: “Car Wash” test with 53 models

#449

Earlier quoted context omitted.

"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.

If walking is an option, then your car is already at the car wash. If your car was not at the car wash, then this wouldn't be a question

What if the car that you want to wash is already at the car wash, but you have a second car? That's still a dumb question nonetheless because you probably need to drive both cars back at some point.

Re: “Car Wash” test with 53 models

#450
post #285
post #219

This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.

I also don't trust the maxbenched results. I am thus making my own benchmarks: https://aibenchy.com

Maybe I am missing something obvious on the website, but where is the documentation? Where do you explain what each number mean, or at least a short overview of what the models are being tested on?
Post reply on HN