Earlier quoted context omitted.
The systems recognized the pattern that it looks like a generic article on the internet asking whether someone should walk or drive and answered it exactly as expected based on their training data. None of this should be surprising. We are the ones fooling ourselves into believing there's more intelligence in these systems than they really have. At the end of the day, it's just an impressive parlor trick.
In that sense the google AI summary search results are a better UX for this type experience
“Car Wash” test with 53 models
441–450 of 469 posts
Re: “Car Wash” test with 53 models
#442Re: “Car Wash” test with 53 models
#443Earlier quoted context omitted.
You passed the intelligence check and failed the wisdom one. The key technique in the mathematical method to answer the machine question is "theory of mind".
It's not theory of mind, it's an understanding of how trick questions are structured and how to answer one. Pretty useless knowledge after high school - no wonder AI companies didn't bother training their models for that
The obvious answer here is 100 minutes because it's impossible to perfectly encapsulate every real life factor. What happens if a gamma ray burst destroys the machines? What happens if the machine operators go on strike? Etc, etc. The answer is 100.
Re: “Car Wash” test with 53 models
#444Earlier quoted context omitted.
You passed the intelligence check and failed the wisdom one. The key technique in the mathematical method to answer the machine question is "theory of mind".
Theory of mind won’t help you answering this question. It is obviously an underspecified question (at least in any contexts where you are not actively designing/thinking about some specific industrial process). As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick. In the first case it is better to ask clarif…
Context would be key here. If this were a question on a grade school word problem test then just say 100, as it is as specified as it needs to be. If it's a Facebook post that says "We asked 1000 people this and only 1 got it right!" then it's probably some trick question.
If you think it's not specified enough for a grade school question, then I would challenge you to come up with a version that's specified rigorously enough for any sufficiently picky interviewee. (Hint: This is not possible)
>There is nothing “mathematical” about any of this though.
Finding the correct approach to solve a problem specified in English is a mathematical skill.
Re: “Car Wash” test with 53 models
#445Re: “Car Wash” test with 53 models
#446Re: “Car Wash” test with 53 models
#447Earlier quoted context omitted.
Same. I usually add a "Be curt" in front of every prompt in Gemini.
Is that more effective than simply adding it to your user instructions?
Re: “Car Wash” test with 53 models
#448Go ask 53 Americans. I’m willing to bet less than 11 get it right.
Don't bet too much, from the linked article ... They ran the exact same question with the same forced choice between "drive" and > "walk," no additional context, past 10,000 real people through their human feedback platform. 71.5% said drive.
Re: “Car Wash” test with 53 models
#449Earlier quoted context omitted.
"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.
If walking is an option, then your car is already at the car wash. If your car was not at the car wash, then this wouldn't be a question
Re: “Car Wash” test with 53 models
#450This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.
I also don't trust the maxbenched results. I am thus making my own benchmarks: https://aibenchy.com