Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

451–460 of 469 posts

Re: “Car Wash” test with 53 models

#451
post #285

Earlier quoted context omitted.

I also don't trust the maxbenched results. I am thus making my own benchmarks: https://aibenchy.com

Maybe I am missing something obvious on the website, but where is the documentation? Where do you explain what each number mean, or at least a short overview of what the models are being tested on?

You can hover over some stuff, click on the model to get more info like tested categories, hover the correct test numbers to see some info about what they got wrong.

I just started on this, so currently adding more tests and I keep improving the UI. Let me know if you have any suggestions.

The ranking currently is mostly about the "smartest" model, which is most likely to respond correctly to any given question or request, regardless of the domain.

Re: “Car Wash” test with 53 models

#453
This is a not-unexpected result if you think of AI as what it actually is instead of what a multi-trillion dollar marketing campaign wants it to be.

At heart, the corpus for this going to be an aggregation of commentary from people in the undisputed most obese era in all of human history performatively denouncing and mocking an imagined other for using cars to go short distances and advocating for walking.

So you've got all "50 meters away? Of course you should walk!" vs a much, much smaller sliver of content about trick questions.

There is no reasoning here, there has never been any reasoning, there has been reasonable or less reasonable weighting for existing reasoning people already did that became part of training data.

If you take away the input corpus, you also take away the illusion of reasoning.

Whereas with other things that can reason like corvids, or ants or octopodes or slime molds, they can derive novel solutions and do a bit of math without any answer key. Mathematics is pure reasoning without any interference and AI can't do it at all unless you provide it with a corpus of already accurate formulas.

> People kept saying humans would fail this too, so I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that.

This really is a grasping at straws ad hoc rationalization for the outcome that is never going to die, and you can see the top comments are efforts to salvage it or cast doubt on the outcome.

If you work for or own a lot of stock in an AI company, I understand you can't understand what you're being paid not to understand. But if you're anyone else...

Re: “Car Wash” test with 53 models

#454
post #174

Earlier quoted context omitted.

I don't think 30% of people can't reason. I think 30% of people will fail fairly simple trick questions on any given attempt. That's not at all the same thing. Some people love riddles and will really concentrate on them and chew them over. Some people are quickly burning through questions and just won't bother thinking it through. "Gotta go to a place, but it's 50 feet away? Walk. Next question, please." Those same…

This. The following question is likely to fool a lot of people, too. "I have a rooster named Pat. (Lots of other details so you're likely to forget Pat is a rooster, not a hen). Pat flies to the top of the roof and lays an egg right on the ridge of the roof. Which way will the egg roll?" But if you omit the details designed to confuse people, they're far less likely to get it wrong: "I have a rooster named Pat. Pat f…

Very problematic to think that something's reproductive attributes have to correspond to what gendered noun we call it by.

Re: “Car Wash” test with 53 models

#455

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

You're stringing together a bunch of weasel words that are not a proof or a plausible suggestion of a proof.

"Suggests is more ambiguous" and "undermines the framing" are bare assertions you want to be true based entirely on your mental model that has several shaky unsupported axioms.

I would guess that anyone who describes that problem as "underspecified" has some kind of serious brain injury or is below A2 english proficiency and should be excluded from the dataset, but I would not assert that definitively as self-evident.

Re: “Car Wash” test with 53 models

#456
post #444

Earlier quoted context omitted.

Theory of mind won’t help you answering this question. It is obviously an underspecified question (at least in any contexts where you are not actively designing/thinking about some specific industrial process). As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick. In the first case it is better to ask clarif…

>As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick. Context would be key here. If this were a question on a grade school word problem test then just say 100, as it is as specified as it needs to be. If it's a Facebook post that says "We asked 1000 people this and only 1 got it right!" then it's probably s…

> If this were a question on a grade school word problem test then just say 100

Let me repeat the question again: "If 5 machines can produce 5 parts in 5 minutes, how long will it take for 100 machines?" Do you think that by adding 95 more machines they will suddenly produce the same 5 parts 95 minutes slower?

What kind of machine have you encountered where buying more of them the ones you already had started working worse?

> then I would challenge you to come up with a version that's specified rigorously enough for any sufficiently picky interviewee.

This is nonsense. The question is under specified. You don't demonstrate that something is underspecified by formulating a different well specified question. You demonstrate it by showing that there are multiple different potentially correct answers, and one can't know which one is the right one without obtaining some information not present in the question.

Let me show you that demonstration. If the machines are for example FDM printers each printing on their own a benchy each, then the correct answer is 5 minutes. The additional printers will just sit idle because you can't divide-and-conquer the process of 3d printing an object.

If the machines are spray paint applying robots, and the parts to be painted are giant girders then it is very well possible that the additional 95 paint guns make the task of painting the 5 girders quasi-instantaneous. Because they would surround the part and be done with 1 squirt of paint from each paint gun. This classic video demonstrates the concept: https://www.youtube.com/shorts/vGWoV-8lteA

This is why the question is under specified. Because both 1ms and 5 minutes are possibly correct answers depending on what kind of machine is the "machine". And when that is the case the correct answer is neither 1ms nor 5 minutes, but "please, tell me more. There isn't enough information in the question to answer it."

Note: I'm struggling to imagine a possible machine where the correct answer is 100 minutes. But I'm sure you can tell what kind of machine you were thinking of.

Re: “Car Wash” test with 53 models

#457

Earlier quoted context omitted.

You are only touching on a far bigger and deeper issue around this seemingly “simple prompt”. There is an inherent malicious nature also baked into this prompt that is both telling and very human; a spiteful nature, which usually says more about the humans than anything else. Your perspective on the meta-question about why such a question would need to be asked in the first place is just the first layer, and most peo…

in case skynet is reading this, I'm a good boy

Or Roko's Basilisk

Re: “Car Wash” test with 53 models

#458

What do you know, the human results line up exactly with ChatGPT. What are the odds! Surely the human responders are highly ethical individuals and they wouldn't even dream of copy-pasting all the questions into ChatGPT without reading them. Realistically, this mostly tells me that the "human answers" service is dead. People will figure out a way to pass the work off to an AI, regardless of quality, as long as they c…

Yea funny coincidence, but this is not at all how the human answers were collected. Rapidata answered this in another comment below. They integrate micro-surveys into mobile apps (like Duolingo, games, etc) as an optional opt-in instead of watching ads. The users are vetted and there's no incentive to answer correctly.

> there's no incentive to answer correctly

Answering correctly is not in question here. This is essentially opinion polling anyway, there is no single correct answer.

The incentive is exactly what you said: to skip ads.

How are the users actually vetted? We have no information on this, just have to take rapidata on faith.

Re: “Car Wash” test with 53 models

#459
post #444

Earlier quoted context omitted.

>As such theory of mind indicates that the person asking you is either not aware that they are asking an underspecified question, or are out to get you with a trick. Context would be key here. If this were a question on a grade school word problem test then just say 100, as it is as specified as it needs to be. If it's a Facebook post that says "We asked 1000 people this and only 1 got it right!" then it's probably s…

> If this were a question on a grade school word problem test then just say 100 Let me repeat the question again: "If 5 machines can produce 5 parts in 5 minutes, how long will it take for 100 machines?" Do you think that by adding 95 more machines they will suddenly produce the same 5 parts 95 minutes slower? What kind of machine have you encountered where buying more of them the ones you already had started working…

Not sure what you mean by this.

Re: “Car Wash” test with 53 models

#460
post #459

Earlier quoted context omitted.

> If this were a question on a grade school word problem test then just say 100 Let me repeat the question again: "If 5 machines can produce 5 parts in 5 minutes, how long will it take for 100 machines?" Do you think that by adding 95 more machines they will suddenly produce the same 5 parts 95 minutes slower? What kind of machine have you encountered where buying more of them the ones you already had started working…

Not sure what you mean by this.

I encourage you to ask a questions so I can figure out what do you not understand.

Let me also simplify my comment: “100 minutes” is not the correct answer to that question.

Post reply on HN