Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

161–170 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#161

Earlier quoted context omitted.

Without checking every answer it gives back to make sure it's factual, you may be ingesting tons of bullshit answers. In this particular answer model A may get it wrong and model B may get it right, but that can be reversed for another question. What do you do at that point? Pay to use all of them and find what's common in the answers? That won't work if most of them are wrong, like for this example. If you're going…

If you're going to have to fact check everything anyways...why bother using them in the first place? "If you're going to have to put gas in the tank, change the oil, and deal with gloves and hearing protection, why bother using a chain saw in the first place?" Tool use is something humans are good at, but it's rarely trivial to master, and not all humans are equally good at it. There's nothing new under that particul…

The difference is consistency. You can read a manual and know exactly how to oil and refill the tank on a chainsaw. You can inspect the blades to see if they are worn. You can listen to it and hear how it runs. If a part goes bad, you can easily replace it. If it's having troubles, it will be obvious - it will simply stop working - cutting wood more slowly or not at all.

The situation with an LLM is completely different. There's no way to tell that it has a wrong answer - aside from looking for the answer elsewhere which defeats its purpose. It'd be like using a chainsaw all day and not knowing how much wood you cut, or if it just stopped working in the middle of the day.

And even if you KNOW it has a wrong answer (in which case, why are you using it?), there's no clear way to 'fix' it. You can jiggle the prompt around, but that's not consistent or reliable. It may work for that prompt, but that won't help you with any subsequent ones.

Re: Ask HN: Share your AI prompt that stumps every model

#162

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

just to confirm I read this right, "the marathon crater" does not in fact exist, but this works because it seems like it should?

The other aspect is it can’t reliably tell whether it „knows” something or not. It’s conditioned to imitate the corpus, but the corpus in a way is its „universe” and it can’t see the boundaries. Everything must map to something _in_ the corpus.

Re: Ask HN: Share your AI prompt that stumps every model

#163
post #2

You might want to get the ball rolling by sharing what you already have

Also, ones that can't be solved at a glance by humans don't count. Like this horrid ambiguous example from SimpleBench I saw a while back that's just designed to confuse: John is 24 and a kind, thoughtful and apologetic person. He is standing in an modern, minimalist, otherwise-empty bathroom, lit by a neon bulb, brushing his teeth while looking at the 20cm-by-20cm mirror. John notices the 10cm-diameter neon lightbul…

I think it's a horrible example, but I just got a very professional response from Gemini 2.5:

> This scenario seems like a bit of a riddle! Let's break it down:

> The Setting: John is alone in an "otherwise-empty" bathroom.

> The Action: He is looking in the mirror.

> The Observation: He sees a bulb falling towards the head of a bald man he is examining in the mirror.

> The Physics: Since he is alone and looking in the mirror, the bald man he is "closely examining" must be his own reflection.

> The Incident: Therefore, the bulb fell towards John's own head. He looked up (at the actual bulb falling towards him), failed to catch it, and it hit him.

> The Reaction: The "bald man" who cursed, yelled "what an idiot!", and left the bathroom was actually John himself, reacting to being hit on the head and possibly feeling foolish for not catching the bulb.

> Conclusion: No, John should not text an apology. The person who was hit by the bulb, got angry, and left was John himself. There is no other bald man to apologize to.

Re: Ask HN: Share your AI prompt that stumps every model

#164

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

I like this but at the same time it seems tricky don't you think? Is the AI model intuiting your intent? There is a Marathon Valley on Mars that could be implied to be a previous crater. I'm not sure if the AI is hallucinating outright or attempting to answer an ambiguous question. It's like saying "tell me about the trade building in New York". Pre-9/11, you'd understand this was the World Trade Center and wouldn't be wrong if you answered someone in this way. "Tell me about the Triangle statue". "Oh the Triangle statue was built in ancient egypt around BC 3100". It's hard to explain, and perhaps I'm anthropomorphizing but it's something humans do. Some of us correct the counter-party and some of us simply roll with the lingo and understand the intent.

Re: Ask HN: Share your AI prompt that stumps every model

#165
post #133

Earlier quoted context omitted.

The inaccuracies are that it is called "Marathon Valley" (not crater) and that it was photographed in April 2015 (from the rim) or that in July 2015 actually entered. The other stuff is correct. I'm guessing this "gotcha" relies on "valley"/"crater", and "crater"/"mars" being fairly close in latent space. ETA: Marathon Valley also exists on the rim of Endeavour crater. Just to make it even more confusing.

None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it.

> None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it.

The Marathon Valley _is_ part of a massive impact crater.

Re: Ask HN: Share your AI prompt that stumps every model

#166

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

Yes let's not say what's wrong with the tech, otherwise someone might (gasp) fix it!

Re: Ask HN: Share your AI prompt that stumps every model

#167

Something about an obscure movie. The one that tends to get them so far is asking if they can help you find a movie you vaguely remember. It is a movie where some kids get a hold of a small helicopter made for the military. The movie I'm concerned with is called Defense Play from 1988. The reason I keyed in on it is because google gets it right natively ("movie small military helicopter" gives the IMDb link as one of…

Someone not very long ago wrote a blog post about asking chatgpt to help him remember a book, and he included the completely hallucinated description of a fake book that chatgpt gave him. Now, if you ask chatgpt to find a similar book, it searches and repeats verbatim the hallucinated answer from the blog post.

Re: Ask HN: Share your AI prompt that stumps every model

#168
post #87
post #61

Earlier quoted context omitted.

GPT 4.5 even doubles down when challenged: > Nope, I didn’t make it up — Marathon crater is real, and it was explored by NASA's Opportunity rover on Mars. The crater got its name because Opportunity had driven about 42.2 kilometers (26.2 miles — a marathon distance) when it reached that point in March 2015. NASA even marked the milestone as a symbolic achievement, similar to a runner finishing a marathon. (Obviously…

This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…

There are use-cases where hallucinations simply do not matter. My favorite is finding the correct term for a concept you don't know the name of. Googling is extremely bad at this as search results will often be wrong unless you happen to use the commonly accepted term, but an LLM can be surprisingly good at giving you a whole list of fitting names just based on a description. Same with movie titles etc. If it hallucinates you'll find out immediately as the answer can be checked in seconds.

The problem with LLMs is that they appear much smarter than they are and people treat them as oracles instead of using them for fitting problems.

Re: Ask HN: Share your AI prompt that stumps every model

#169
post #133

Earlier quoted context omitted.

None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it.

> None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it. The Marathon Valley _is_ part of a massive impact crater.

If you asked me for all the details of a Honda Civic and I gave you details about a Honda Odyssey you would not say I was correct in any way. You would say I was wrong.

Re: Ask HN: Share your AI prompt that stumps every model

#170
post #133

Earlier quoted context omitted.

The inaccuracies are that it is called "Marathon Valley" (not crater) and that it was photographed in April 2015 (from the rim) or that in July 2015 actually entered. The other stuff is correct. I'm guessing this "gotcha" relies on "valley"/"crater", and "crater"/"mars" being fairly close in latent space. ETA: Marathon Valley also exists on the rim of Endeavour crater. Just to make it even more confusing.

None of it is correct because it was not asked about Marathon Valley, it was asked about Marathon Crater, a thing that does not exist, and it is claiming that it exists and making up facts about it.

Or it's assuming you are asking about Marathon Valley, which is very reasonable given the context.

Ask it about "Marathon Desert", which does not exist and isn't closely related to something that does exist, and it asks for clarification.

I'm not here to say LLMs are oracles of knowledge, but I think the need to carefully craft specific "gotcha" questions in order to generate wrong answers is a pretty compelling case in the opposite direction. Like the childhood joke of "Whats up?"..."No, you dummy! The sky is!"

Straightforward questions with straight wrong answers are far more interesting. I don't many people ask LLMs trick questions all day.

Post reply on HN