Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

291–300 of 469 posts

Re: “Car Wash” test with 53 models

#291

Earlier quoted context omitted.

> Back in the 1970's we'd pull a spark plug and screw in a hose to use the compression phase to inflate tyres. You'd inflate your tires with a gasoline and air mix?

I mean... you don't breathe insides of your tires

No, but tyres are rubber and they heat up ...

One might reasonably wonder if the material might degrade or the tyre explode while running hot.

Can confirm, that doesn't happen.

Re: “Car Wash” test with 53 models

#292

Earlier quoted context omitted.

For all we know 30 percent of users on that site are now bots.

[flagged]

You are absolutely right! It's not just relevant, it's a much funnier take at robots mannerisms than what ended up having in the end.

Re: “Car Wash” test with 53 models

#293

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

> which I suppose makes sense if 30% of people simply lack the ability to reason

I think it would be better to say that 30% of people either lack the ability to reason (inarguably true in a few cases, though I'd suggest, and hope, an order of magnitude or two less than 30%, as that would be a life-altering mental impairment) or just can't generally be bothered to, or just didn't (because they couldn't be bothered, or because they felt some social pressure to answer quickly rather than taking more than an instant time to think) at the time of being asked this particular question.

An automated system like an LLM to not have this problem. It has no path to turn off or bypass any function that it has, so if it could reason it would.

Re: “Car Wash” test with 53 models

#294

Earlier quoted context omitted.

"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.

If walking is an option, then your car is already at the car wash. If your car was not at the car wash, then this wouldn't be a question

It feel a bit like this to me. That's not to say LLMs should not have detected this, but I still feel like this fits the "vibes" the question gives, and some LLMs fall into that trap. Is it actually what's happening in the neural nets? Maybe not! But I always find it interesting or at least entertaining to approach those questions that way nonetheless; especially given the pattern matching nature of LLMs.

Re: “Car Wash” test with 53 models

#295

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

I don't think it's ambiguous, but I have been wondering how much LLMs model human behavior that we just don't recognize due to the subset of people on this site. I recently saw a comment online that "Mandarin isn't anyone's first language, people in China's first language is a dialect". It just struck me at that moment that people also hallucinate information confidently all the time.

> It just struck me at that moment that people also hallucinate information confidently all the time.

And many will just repeat what was confidently said without question.

I know this it true, because my intelligent mate down the pub says so.

Re: “Car Wash” test with 53 models

#296

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

We should also check the specifics of the experiment. Is it possible that humans participating simply copied and pasted the question and answer to an LLM?

Re: “Car Wash” test with 53 models

#297

Earlier quoted context omitted.

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

> which I suppose makes sense if 30% of people simply lack the ability to reason I think it would be better to say that 30% of people either lack the ability to reason (inarguably true in a few cases, though I'd suggest, and hope, an order of magnitude or two less than 30%, as that would be a life-altering mental impairment) or just can't generally be bothered to, or just didn't (because they couldn't be bothered, or…

This is something I have wondered about before: whether AIs are more likely to give wrong answers when you ask a stupid question instead of a sensible one. Speaking personally, I often cannot resist the temptation to give reductio-ad-absurdum answers to particularly ridiculous questions.

If 30% of humans on the internet can't be bothered to make an effort to answer stupid questions correctly, then one would expect AIs to replicate this behaviour. And if humans on the internet sometimes provide sarcastic answers when presented with ridiculous questions, one would expect AIs to replicate this behavior as well.

So you really cannot say they have no incentive to do so. The incentive they have is that they get rewarded for replicating human behaviour.

Re: “Car Wash” test with 53 models

#298

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

"Don't move -- call the service station to have someone sent over to your place to hand wash the car" would be a valid answer. It's a little "out of the box" but it makes more sense than walking to the car wash and leaving the car behind, or walking and maybe lift the car on your shoulders.

Re: “Car Wash” test with 53 models

#299

Earlier quoted context omitted.

Very awkward and formal. Anyone would call it lavage auto, lave-auto or simply lavage if the context is clear.

Maybe I'm too old or my family was weird. We called it "le carwash" with a beautifully French "carouache" pronunciation. But yeah, "lave-auto" sounds more familiar.

Honestly, If anyone asked me "T'as fait quoi?" I'd blurt out "J'ai amené ma voiture chez le lavage". Background: I stopped speaking french when I was ten and my family isn't native, but it feels more conversational than "station de lavage".
Post reply on HN