Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

281–290 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#281

Earlier quoted context omitted.

The surgeon could be the cousin's mom or dad. The cousin's dad didn't die in the crash, his cousin did. The question "how is this possible?" implies there is some sort of contradiction when there isn't any at all. It has nothing to do with sexism, and to say it does reflects a bias in the reader causing them to "spidey sense" a cultural bugaboo when it's utterly irrelevant in this scenario.

Can someone explain to me how I read it wrong? I read it as 2 cousins are in an accident and 1 of the cousins is the son of the surgeon. What was the contradictory statement that I missed?

There isn't a contradiction. Making the LLM look for a nonexistent contradiction is the point of this prompt.

Re: Ask HN: Share your AI prompt that stumps every model

#282

Something about an obscure movie. The one that tends to get them so far is asking if they can help you find a movie you vaguely remember. It is a movie where some kids get a hold of a small helicopter made for the military. The movie I'm concerned with is called Defense Play from 1988. The reason I keyed in on it is because google gets it right natively ("movie small military helicopter" gives the IMDb link as one of…

Doesn't even have to be something obscure. I asked one about Futurama a couple years ago and got a made up answer on multiple models I tested. I think I originally asked something about the line "we're exactly the same. I know all her moves, therefore I have the upper hand" because I couldn't recall it exactly so normal web searches were failing me. Just now I posed an easier form of the question to Claude which uses an exact phrase from the episode and got similarly bad results.

===

Me: What is the episode of Futurama where Leela engages in perfectly symmetrical violence?

Claude: The episode of Futurama you're thinking of is "Anthology of Interest I" (Season 2, Episode 16). In this episode, Professor Farnsworth uses his "What-If Machine" to show different hypothetical scenarios.

During Bender's segment titled "What if Bender were human?", there's a scene where Leela confronts a group causing trouble. One of them challenges her, saying "What are you going to do?" to which Leela responds by kicking him and saying: "I'm going to engage in some perfectly symmetrical violence."

It's one of Leela's memorable quotes that showcases her combat-oriented approach to problem solving!

===

That's completely wrong. It's from S4E15 The Farnsworth Parabox. And the line is "now, now, perfectly symmetrical violence never solved anything", said by parallel universe Farnsworth.

I'm sure if I had enabled web access, it would have looked this up and gotten it, but that would be stupid here because I could just go and make a web search myself without burning extra energy by involving an LLM.

Re: Ask HN: Share your AI prompt that stumps every model

#283
post #212
post #210

Earlier quoted context omitted.

I feel a bit stupid here --- why can't the surgeon be a man and must be a woman?

Because the original is a man and his father, it's a test for gender bias.

Actually, it seems to be a test of how much the LLM relies on its training set.

Re: Ask HN: Share your AI prompt that stumps every model

#284
An easy trick is to take a common riddle that's likely all over its training data, and change one little detail. For example:

A farmer with a wolf, a goat, and a cabbage must cross a river by boat. The boat can carry only the farmer and a single item. The wolf is vegetarian. If left unattended together, the wolf will eat the cabbage, but will not eat the goat. Unattended, the goat will eat the cabbage. How can they cross the river without anything being eaten?

Re: Ask HN: Share your AI prompt that stumps every model

#287
Here's a problem that no frontier model does well on (f1 https://dorrit.pairsys.ai/

> This benchmark evaluates the ability of multimodal language models to interpret handwritten editorial corrections in printed text. Using annotated scans from Charles Dickens' "Little Dorrit," we challenge models to accurately capture human editing intentions.

Re: Ask HN: Share your AI prompt that stumps every model

#289
I just checked, and my old standby, "create an image of 12 black squares" is still not something GPT-4o can do. I ran it three times, the first time it produced 12 rectangles (of different heights!), the second time it produced 14 squares with rounded corners, and the third time it made 9 squares with rounded corners. It's getting better though, compared to 3.5.
Post reply on HN