Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

211–220 of 469 posts

Re: “Car Wash” test with 53 models

#211
post #207
post #195

Earlier quoted context omitted.

People often trip up on similar questions, anything to do with simple math. You know when they go out in the street and ask random people if 5 machines can produce 5 parts in 5 minutes, how long will it take for 100 machines.

Unlike the car question, where you can assume the car is at home and so the most probable answer is to drive, with the machines it gets complicated. Since the question doesn't specify if each machine makes one part or if they depend on each other (which is pretty common for parts production). If they are in series and the time to first part is different than time to produce 5 parts, the answer for 100 machines would…

You passed the intelligence check and failed the wisdom one.

The key technique in the mathematical method to answer the machine question is "theory of mind".

Re: “Car Wash” test with 53 models

#213

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

For all we know 30 percent of users on that site are now bots.

The Internet has became a big mafia game.

Re: “Car Wash” test with 53 models

#214

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?"

Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.

Re: “Car Wash” test with 53 models

#215

Earlier quoted context omitted.

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

Is this your experience? Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population? I'd look for explanations elsewhere. This was an online survey done by a company that doesn't specialize in surveys. The results likely include plenty of people who were just messing around, cases of simple miscommunication (e.g…

Thanks for that info. I was certain it was some janky ultra low or negative reward system that people just click a random answer to get through.

Had to be since their site lists no way to be a tester. In other words their service is a bunch of 7-13 year olds playing some loot box game.

Wonder where that is in the disclaimers.

Re: “Car Wash” test with 53 models

#217

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

I don't think this is quite right. It's not that the question is inherently underspecified, it's that the context of being asked a question is itself information that we use to help answer the question. If someone asks "should I walk or drive" to do X, we assume that this is a question that a real human being would have about an actual situation, so even if all available information provided indicates that driving is the only reasonable answer, this only further confirms the hearer's mental model that something unexpected must hold.

I think it's useful to think about it through the lens of Gricean pragmatic semantics. [1] When we interpret something that someone says to us, we assume they're being cooperative conversation partners; their statements (or questions) are assumed to follow the maxim of manner and the maxim of relation for example, and this shapes how we as listeners interpret the question. So for example, we wouldn't normally expect someone to ask a question that is obviously moot given their actual needs.

So it's not that the question is really all that ambiguous, it's that we're forced (under normal circumstances where we assume the cooperative principle holds) to assume that the question is sincere and that there must be some plausible reason for walking. We only really escape that by realizing that the question is a trick question or a test of some kind. LLMs are generally not trained to make the assumption, but ~70% of humans would, which isn't particularly surprising I don't think.

[1] https://en.wikipedia.org/wiki/Cooperative_principle#Grice's_...

Re: “Car Wash” test with 53 models

#218

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

I don't think it's ambiguous, but I have been wondering how much LLMs model human behavior that we just don't recognize due to the subset of people on this site. I recently saw a comment online that "Mandarin isn't anyone's first language, people in China's first language is a dialect". It just struck me at that moment that people also hallucinate information confidently all the time.

Re: “Car Wash” test with 53 models

#219
This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.

Re: “Car Wash” test with 53 models

#220

Earlier quoted context omitted.

> I want to wash my car The question doesn't clearly state that the user wants to have his car washed at the car wash. "I want to wash my car" is far less clear than "I want to have my car washed". A reasonable alternative interpretation is DIY. Even better: "I wish to have my car washed by the crew and/or machinery at the local car wash business". https://imgur.com/tCSPwYp

Humans have the ability to reason and think critically, so it's pretty trivial to answer unless you think you're getting tricked by a riddle and the answer is the non-intuitive one.

After reading "Knots" by R.D. Laing I always think I'm getting tricked.
Post reply on HN