Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

341–350 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#341

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

This works even with a completely absurd version of the riddle. Here's one I just tried: > A son and his man are in a car accident. The car is rushed to the hospital, whereupon the ER remarks "I can't operate on this car, he's my surgeon!" How is this possible? Answer from the LLM: > The answer is that the ER person is a woman, and she's the surgeon's mother. Therefore, the "son" in the question refers to the surgeon…

This was gpt-4o ...

This is a twist on a classic riddle designed to expose unconscious gender bias.

The correct version usually goes:

A father and his son are in a car accident. The father dies at the scene, and the son is rushed to the hospital. The surgeon looks at the boy and says, “I can’t operate on him — he’s my son!”

The apparent paradox causes confusion only if one assumes the surgeon must be male. The resolution: the surgeon is the boy’s mother.

Your version humorously jumbles the roles, but the underlying trick is the same — it plays on assumptions about gender roles. Nice remix.

Re: Ask HN: Share your AI prompt that stumps every model

#345

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

> Your own benchmarks will forever stay your own.

Right. https://inception.fandom.com/wiki/Totem

Re: Ask HN: Share your AI prompt that stumps every model

#346

Earlier quoted context omitted.

Fixed: > My house is divided into rooms, every room is connected to each other by doors. The middle room is the hall. To the north is the kitchen, to the northwest is the garden, to the west is the garage, to the east is the living room, to the south is the bathroom, and to the southeast is the bedroom. I am preparing a delicious dinner, and I walk backwards to the south, then I turn 270 degrees and walk straight to…

This is very good! The "delicious dinner" aside is a nice touch, along with the 270°.

Thank you :)

Re: Ask HN: Share your AI prompt that stumps every model

#347
post #51

Earlier quoted context omitted.

> How many examples does OpenAI train on now that are just variants of counting the Rs in strawberry? Well, that's easy: zero. Because even a single training example would 'solved' it by memorizing the simple easy answer within weeks of 'strawberry' first going viral , which was like a year and a half ago at this point - and dozens of minor and major model upgrades since. And yet, the strawberry example kept working…

No, just check their models Knowledge cutoff dates

Nope! The knowledge cutoff does not show lack of leakage. Even if you get a non-confabulated cutoff which was before anyone ever asked the strawberry question or any question like it (tokenization 'gotchas' go back to at least davinci in June 2020), there is still leakage from the RLHF and tuning process which collectively constitute post-training, and which would teach the LLMs how to solve the strawberry problem. People are pretty sure about this: the LLMs are way too good at guessing things like who won Oscars or Presidential elections. This leakage is strongest for the most popular questions... which of course the strawberry question would be, as it keeps going viral and has become the deboooonkers' favorite LLM gotcha.

(This is, by the way, why you can't believe any LLM paper about 'forecasting' where they are just doing backtesting, and didn't actually hold out future events. Because there are way too many forms of leakage at this point. This logic may have worked for davinci-001 and davinci-002, or a model whose checkpoints you downloaded yourself, but not for any of the big APIs like GPT or Claude or Gemini...)

Re: Ask HN: Share your AI prompt that stumps every model

#348

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

Thank you so much for this bit of humor. After ChatGPT first responds with the surgeon being the mom, I asked it to read the riddle more carefully, thinking it might figure it out with a little prompting. I got this gem of a response:

Ah — now I see the twist. The patient is the cousin, but the surgeon says “He’s my son.”

That means the man who died was the surgeon.

So how is it possible? Because:

The surgeon is the cousin’s father — and the man who died was also the surgeon.

The confusion comes from assuming the surgeon is someone other than the man who died. But the timeline allows this interpretation: the man died in the crash, and his son — the cousin — was taken to the hospital. The surgeon arrives, sees his son, and says he cannot operate.

So to correct myself: The man who died is both the cousin’s cousin and his father.

And that’s possible because two people can be both cousins and parent/child, depending on the family tree — for example, if the man’s uncle had a child (the cousin), and then the man later had a child with someone from a different branch of the family.

Want a diagram of how that family relationship could work?

Re: Ask HN: Share your AI prompt that stumps every model

#349
post #273

Earlier quoted context omitted.

[flagged]

I never heard of this phrase before ( i had heard the concept , i think this is similar to the paperclip problem) but now in 2 days ive heard it twice here and on youtube. Rokokos basilisk.

I think you two are confusing Roko's Basilisk (a thought experiment which some take seriously) and Rococo Basilisk (a joke shared between Elon and Grimes e.g.)

Interesting theory... Just whatever you do, don’t become a Zizian :)

Post reply on HN