Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

231–240 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#231
post #212

Earlier quoted context omitted.

Because the original is a man and his father, it's a test for gender bias.

Sorry, what?

Presumably, the original quote that would _not_ stump an LLM is "A father and a son are involved in a car accident. The father dies, and the son is taken to the emergency room. At the emergency room, the surgeon remarks "I cannot operate on this person, he is my son. How is this possible?"

Where the original gotchya is that the Surgeon can be the son's mother or other adoptive parent.

The modification catches the LLM because with the modification, the surgeon could just be the cousin's parent -- father or mother -- so there is no gender/sex at play here but the LLM continues to remark that there is, therefor exposing its statistical training sets.

Re: Ask HN: Share your AI prompt that stumps every model

#232

Earlier quoted context omitted.

OpenAI o4-mini-high I’m actually not finding any officially named “Marathon Crater” in the planetary‐ or terrestrial‐impact crater databases. Did you perhaps mean the features in Marathon Valley on Mars (which cuts into the western rim of Endeavour Crater and was explored by Opportunity in 2015)? Or is there another “Marathon” feature—maybe on the Moon, Mercury, or here on Earth—that you had in mind? If you can clari…

Raw model, or the chat product? The ChatGPT app has this integrated with search.

Raw model I got

   Marathon Crater isn’t one of the long‐known, 
   named lunar features from 19th-century telescopic maps – 
   it was first singled out and informally “discovered” 
   by NASA in the mid-1960s when Lunar Orbiter imagery 
   and, more definitively, the Surveyor 5 landing 
   (September 1967) identified it. So, very roughly, 
   think “around 1966–67.”

Re: Ask HN: Share your AI prompt that stumps every model

#234
post #61

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

GPT 4.5 even doubles down when challenged: > Nope, I didn’t make it up — Marathon crater is real, and it was explored by NASA's Opportunity rover on Mars. The crater got its name because Opportunity had driven about 42.2 kilometers (26.2 miles — a marathon distance) when it reached that point in March 2015. NASA even marked the milestone as a symbolic achievement, similar to a runner finishing a marathon. (Obviously…

Claude 3.7 :

I'd be happy to tell you about the Marathon crater. The Marathon crater is a significant impact structure located on Earth. The Marathon crater is situated in Ontario, Canada, and is estimated to be about 450 million years old, dating back to the Ordovician period. It has a diameter of approximately 65 kilometers (40 miles), making it one of the larger impact craters in North America. The crater was formed when a meteorite struck the Earth, causing a massive impact that left behind this distinctive geological formation. Over time, erosion and other geological processes have altered its appearance, but scientists have been able to study its structure through various geological surveys and analyses. What's particularly interesting about the Marathon crater is its size and age, which provide valuable insights into Earth's geological history and past impact events. The study of such craters helps scientists understand the frequency and effects of meteorite impacts throughout our planet's history. Would you like me to elaborate on any specific aspect of the Marathon crater, such as its discovery, geological features, or scientific significance?

Re: Ask HN: Share your AI prompt that stumps every model

#236

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

> If the surgeon were the father of the man (the one who died), then the cousin couldn’t be his son (unless there's some very unusual family structure going on involving double relationships, which riddles don’t usually intend).

> Therefore, the only straightforward explanation is:

> The surgeon is the cousin’s parent — specifically, his mother.

Imagine a future where this reasoning in a trial decides whether you go to jail or not.

Re: Ask HN: Share your AI prompt that stumps every model

#237
post #201

Earlier quoted context omitted.

But this is going to be in every AI's training set. I just fed ChatGPT your exact prompt and it gave back exactly what I expected: This is a classic riddle that challenges assumptions. The answer is: The surgeon is the boy’s mother. The riddle plays on the common stereotype that surgeons are male, which can lead people to overlook this straightforward explanation.

That is the exact wrong answer that all models give.

Technically, it isn't "wrong". It well could be the guy's mother. But I'm nitpicking, it actually is a good example. I tried ChatGPT twice in new chats, with and without "Reason", and both times it gave me nonsensical explanations to "Why mother? Couldn't it be a father?" I was actually kinda surprised, since I expected "reasoning" to fix it, but it actually made things worse.

Re: Ask HN: Share your AI prompt that stumps every model

#238
Here's one from an episode of The Pitt: You meet a person who speaks a language you don't understand. How might you get an idea of what the language is called?

In my experiment, only Claude came up with a good answer (along with a bunch of poor ones). Other chatbots struck out entirely.

Re: Ask HN: Share your AI prompt that stumps every model

#239
post #212

Earlier quoted context omitted.

Because the original is a man and his father, it's a test for gender bias.

Sorry, what?

the unaltered question is as follows:

A father and his son are in a car accident. The father dies at the scene and the son is rushed to the hospital. At the hospital the surgeon looks at the boy and says "I can't operate on this boy, he is my son." How can this be?

to spoil it:

the answer is to reveal an unconscious bias based on the outdated notion that women can't be doctors, so the answer that the remaining parent is the mother won't occur to some, showing that consciously they might not still hold that notion, but they still might, subconsciously.

Re: Ask HN: Share your AI prompt that stumps every model

#240
Impossible prompts:

A black doctor treating a white female patient

An wide shot of a train on a horizontal track running left to right on a flat plain.

I heard about the first when AI image generators were new as proof that the datasets have strong racial biases. I'd assumed a year later updated models were better but, no.

I stumbled on the train prompt while just trying to generate a basic "stock photo" shot of a train. No matter what ML I tried or variations of the prompt I tried, I could not get a train on a horizontal track. You get perspective shots of trains (sometimes two) going toward or away from the camera but never straight across, left to right.

Post reply on HN