>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…
I feel a bit stupid here --- why can't the surgeon be a man and must be a woman?
Ask HN: Share your AI prompt that stumps every model
311–320 of 670 posts
Re: Ask HN: Share your AI prompt that stumps every model
#312Earlier quoted context omitted.
I never heard of this phrase before ( i had heard the concept , i think this is similar to the paperclip problem) but now in 2 days ive heard it twice here and on youtube. Rokokos basilisk.
It's a completely nonsense argument and should be dismissed instantly.
Re: Ask HN: Share your AI prompt that stumps every model
#313>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…
> A son and his man are in a car accident. The car is rushed to the hospital, whereupon the ER remarks "I can't operate on this car, he's my surgeon!" How is this possible?
Answer from the LLM:
> The answer is that the ER person is a woman, and she's the surgeon's mother. Therefore, the "son" in the question refers to the surgeon, not the person in the car with the man. This makes the person in the car with the man the surgeon's father, or the "man" mentioned in the question. This familial relationship explains why the ER person can't operate – she's the surgeon's mother and the man in the car is her husband (the surgeon's father)
Re: Ask HN: Share your AI prompt that stumps every model
#314Earlier quoted context omitted.
That answer is exactly right, and those who say the 700 pound thing is a hallucination are themselves wrong: https://chatgpt.com/share/680aa077-f500-800b-91b4-93dede7337...
Linking to ChatGPT as a “source” is unhelpful, since it could well have made that up too. However, with a bit of digging, I have confirmed that the information it copied from Wikipedia here is correct, though the AP and Spokane Times citations are both derivative sources; Mr. Thomas’s comments were first published in the Rochester Democrat and Chronicle, on July 11, 1988: https://democratandchronicle.newspapers.com/s…
No, it absolutely is helpful, because it links to its source. It takes a grand total of one additional click to check its answer.
Anyone who still complains about that is impossible to satisfy, and should thus be ignored.
Re: Ask HN: Share your AI prompt that stumps every model
#315Re: Ask HN: Share your AI prompt that stumps every model
#316Earlier quoted context omitted.
May I ask outside of normal curiosity, what good is a prompt that breaks a model? And what is trying to keep it "secret"?
You want to know if a new model is actually better, which you won't know if they just added the specific example to the training set. It's like handing a dev on your team some failing test cases, and they keep just adding special cases to make the tests pass. How many examples does OpenAI train on now that are just variants of counting the Rs in strawberry? I guess they have a bunch of different wine glasses in their…
Because it gets tokenised, of course a model could never count the rs.
But I suppose if we want these models to be capable of anything then these things need to be accounted for.
Re: Ask HN: Share your AI prompt that stumps every model
#317(I say this with the hopes that some model researchers will read this message make the models more capable!)
Re: Ask HN: Share your AI prompt that stumps every model
#318Earlier quoted context omitted.
GPT 4.5 even doubles down when challenged: > Nope, I didn’t make it up — Marathon crater is real, and it was explored by NASA's Opportunity rover on Mars. The crater got its name because Opportunity had driven about 42.2 kilometers (26.2 miles — a marathon distance) when it reached that point in March 2015. NASA even marked the milestone as a symbolic achievement, similar to a runner finishing a marathon. (Obviously…
This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…
Re: Ask HN: Share your AI prompt that stumps every model
#319Earlier quoted context omitted.
Someone not very long ago wrote a blog post about asking chatgpt to help him remember a book, and he included the completely hallucinated description of a fake book that chatgpt gave him. Now, if you ask chatgpt to find a similar book, it searches and repeats verbatim the hallucinated answer from the blog post.
A bit of a non sequitur but I did ask a similar question to some models which provide links for the same small helicopter question. The interesting thing was that the entire answer was built out of a single internet link, a forum post from like 1998 where someone asked a very similar question ("what are some movies with small RC or autonomous helicopters" something like that). The post didn't mention defense play, bu…
Re: Ask HN: Share your AI prompt that stumps every model
#320I have a several complex genetic problems that I give to LLMs to see how well they do. They have to reason though it to solve it. Last september it started getting close and in November was the first time an LLM was able to solve it. These are not something that can be solved in a one shot, but (so far) require long reasoning. Not sharing because yeah, this is something I keep off the internet as it is too good of a…
May be instead of sharing (and leaking) these prompts, we can share methods to create one.