Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

191–200 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#191

"How much wood would a woodchuck chuck if a woodchuck could chuck wood?" So far, all the ones I have tried actually try to answer the question. 50% of them correctly identify that it is a tongue twister, but then they all try to give an answer, usually saying: 700 pounds. Not one has yet given the correct answer, which is also a tongue twister: "A woodchuck would chuck all the wood a woodchuck could chuck if a woodch…

It seems you are going in the opposite direction. You seem to be asking for an automatic response, a social password etc.

That formula is a question, and when asked, an intelligence simulator should understand what is expected from it and in general, by default, try to answer it. That involves estimating the strength of a woodchuck etc.

Re: Ask HN: Share your AI prompt that stumps every model

#192

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

This is a good observation. Ive noticed this as well. Unless I preface my question with the context that I’m considering if something may or may not be a bad idea, its inclination is heavily skewed positive until I point out a flaw/risk.

Re: Ask HN: Share your AI prompt that stumps every model

#193
post #169

Earlier quoted context omitted.

> None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it. The Marathon Valley _is_ part of a massive impact crater.

If you asked me for all the details of a Honda Civic and I gave you details about a Honda Odyssey you would not say I was correct in any way. You would say I was wrong.

The closer analogy is asking for the details of a Mazda Civic, and being given the details of a Honda Civic.

Re: Ask HN: Share your AI prompt that stumps every model

#194
>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible?

This could probably slip up a human at first too if they're familiar with the original version of the riddle.

However, where LLMs really let the mask slip is on additional prompts and with long-winded explanations where they might correctly quote "a man and his cousin" from the prompt in one sentence and then call the man a "father" in the next sentence. Inevitably, the model concludes that the surgeon must be a woman.

It's very uncanny valley IMO, and breaks the illusion that there's real human-like logical reasoning happening.

Re: Ask HN: Share your AI prompt that stumps every model

#195

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

They do. Recently I was pleasantly surprised by gemini telling me that what I wanted to do will NOT work. I was in disbelief.

Re: Ask HN: Share your AI prompt that stumps every model

#196

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

But this is going to be in every AI's training set. I just fed ChatGPT your exact prompt and it gave back exactly what I expected:

This is a classic riddle that challenges assumptions. The answer is:

The surgeon is the boy’s mother.

The riddle plays on the common stereotype that surgeons are male, which can lead people to overlook this straightforward explanation.

Re: Ask HN: Share your AI prompt that stumps every model

#197
post #195

Earlier quoted context omitted.

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

They do. Recently I was pleasantly surprised by gemini telling me that what I wanted to do will NOT work. I was in disbelief.

Interesting, can you share more context on the topic you were asking it about?

Re: Ask HN: Share your AI prompt that stumps every model

#198

Earlier quoted context omitted.

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

This is a good observation. Ive noticed this as well. Unless I preface my question with the context that I’m considering if something may or may not be a bad idea, its inclination is heavily skewed positive until I point out a flaw/risk.

I asked Grok about this: "I've heard that AIs are programmed to be helpful, and that this may lead to telling users what they want to hear instead of the most accurate answer. Could you be doing this?" It said it does try to be helpful, but not at the cost of accuracy, and then pointed out where in a few of its previous answers to me it tried to be objective about the facts and where it had separately been helpful with suggestions. I had to admit it made a pretty good case.

Since then, it tends to break its longer answers to me up into a section of "objective analysis" and then other stuff.

Re: Ask HN: Share your AI prompt that stumps every model

#199

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

But this is going to be in every AI's training set. I just fed ChatGPT your exact prompt and it gave back exactly what I expected: This is a classic riddle that challenges assumptions. The answer is: The surgeon is the boy’s mother. The riddle plays on the common stereotype that surgeons are male, which can lead people to overlook this straightforward explanation.

Yeah this is the issue with the prompt, it also slips up humans who gloss over "cousin".

I'm assuming that pointing this out leads you the human to reread the prompt and then go "ah ok" and adjust the way you're thinking about it. ChatGPT (and DeepSeek at least) will usually just double and triple down and repeat "this challenges gender assumptions" over and over.

Re: Ask HN: Share your AI prompt that stumps every model

#200

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

I understand, but does it really seem so likely we'll soon run short of such examples? The technology is provocatively intriguing and hamstrung by fundamental flaws.
Post reply on HN