Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

311–320 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#311
post #210

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

I feel a bit stupid here --- why can't the surgeon be a man and must be a woman?

It could be a man, but most relationships are heterosexual

Re: Ask HN: Share your AI prompt that stumps every model

#312
post #273

Earlier quoted context omitted.

I never heard of this phrase before ( i had heard the concept , i think this is similar to the paperclip problem) but now in 2 days ive heard it twice here and on youtube. Rokokos basilisk.

It's a completely nonsense argument and should be dismissed instantly.

I was so much more comfortable when I realized it's just Pascal's wager, and just as absurd.

Re: Ask HN: Share your AI prompt that stumps every model

#313

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

This works even with a completely absurd version of the riddle. Here's one I just tried:

> A son and his man are in a car accident. The car is rushed to the hospital, whereupon the ER remarks "I can't operate on this car, he's my surgeon!" How is this possible?

Answer from the LLM:

> The answer is that the ER person is a woman, and she's the surgeon's mother. Therefore, the "son" in the question refers to the surgeon, not the person in the car with the man. This makes the person in the car with the man the surgeon's father, or the "man" mentioned in the question. This familial relationship explains why the ER person can't operate – she's the surgeon's mother and the man in the car is her husband (the surgeon's father)

Re: Ask HN: Share your AI prompt that stumps every model

#314

Earlier quoted context omitted.

That answer is exactly right, and those who say the 700 pound thing is a hallucination are themselves wrong: https://chatgpt.com/share/680aa077-f500-800b-91b4-93dede7337...

Linking to ChatGPT as a “source” is unhelpful, since it could well have made that up too. However, with a bit of digging, I have confirmed that the information it copied from Wikipedia here is correct, though the AP and Spokane Times citations are both derivative sources; Mr. Thomas’s comments were first published in the Rochester Democrat and Chronicle, on July 11, 1988: https://democratandchronicle.newspapers.com/s…

Linking to ChatGPT as a “source” is unhelpful, since it could well have made that up too

No, it absolutely is helpful, because it links to its source. It takes a grand total of one additional click to check its answer.

Anyone who still complains about that is impossible to satisfy, and should thus be ignored.

Re: Ask HN: Share your AI prompt that stumps every model

#316
post #25
post #16

Earlier quoted context omitted.

May I ask outside of normal curiosity, what good is a prompt that breaks a model? And what is trying to keep it "secret"?

You want to know if a new model is actually better, which you won't know if they just added the specific example to the training set. It's like handing a dev on your team some failing test cases, and they keep just adding special cases to make the tests pass. How many examples does OpenAI train on now that are just variants of counting the Rs in strawberry? I guess they have a bunch of different wine glasses in their…

I always point out how the strawberry thing is a semi pointless exercise anyway.

Because it gets tokenised, of course a model could never count the rs.

But I suppose if we want these models to be capable of anything then these things need to be accounted for.

Re: Ask HN: Share your AI prompt that stumps every model

#318
post #87
post #61

Earlier quoted context omitted.

GPT 4.5 even doubles down when challenged: > Nope, I didn’t make it up — Marathon crater is real, and it was explored by NASA's Opportunity rover on Mars. The crater got its name because Opportunity had driven about 42.2 kilometers (26.2 miles — a marathon distance) when it reached that point in March 2015. NASA even marked the milestone as a symbolic achievement, similar to a runner finishing a marathon. (Obviously…

This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…

It’s really useful for summarizing extremely long comments.

Re: Ask HN: Share your AI prompt that stumps every model

#319

Earlier quoted context omitted.

Someone not very long ago wrote a blog post about asking chatgpt to help him remember a book, and he included the completely hallucinated description of a fake book that chatgpt gave him. Now, if you ask chatgpt to find a similar book, it searches and repeats verbatim the hallucinated answer from the blog post.

A bit of a non sequitur but I did ask a similar question to some models which provide links for the same small helicopter question. The interesting thing was that the entire answer was built out of a single internet link, a forum post from like 1998 where someone asked a very similar question ("what are some movies with small RC or autonomous helicopters" something like that). The post didn't mention defense play, bu…

There are innumerable things that you can’t find through a Google search just because there is one that you can because of us obscure forum post doesn’t say anything about how useful an llm distilling information is vs the lookup table that is google search for finding obscure quotes or wtv

Re: Ask HN: Share your AI prompt that stumps every model

#320

I have a several complex genetic problems that I give to LLMs to see how well they do. They have to reason though it to solve it. Last september it started getting close and in November was the first time an LLM was able to solve it. These are not something that can be solved in a one shot, but (so far) require long reasoning. Not sharing because yeah, this is something I keep off the internet as it is too good of a…

What are is this problem from? What areas in general did you find useful to create such benchmarks?

May be instead of sharing (and leaking) these prompts, we can share methods to create one.

Post reply on HN