Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

181–190 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#182
Some easy ones I recently found involve leading in the question to state wrong details about a figure, apparently through relations which are in fact of opposition.

So, you can make them call Napoleon a Russian (etc.) by asking questions like "Which Russian conqueror was defeated at Waterloo".

Re: Ask HN: Share your AI prompt that stumps every model

#184

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

Like this one a lot. Perplexity gets this right, probably because it searches the web. "When was Marathon Crater discovered? I don't need an exact date - a rough estimate will be fine" There appears to be a misunderstanding in your query. Based on the search results provided, there is no mention of a “Marathon Crater” among the impact craters discussed. The search results contain information about several well-known…

> Perplexity gets this right

Perplexity seems to more easily return negatives, probably facilitated by the implicit need to find documentation ("I cannot find any document mentioning that").

But Perplexity can also easily speak its own dubious piece of mind unless requested explicitly "provide links to documents that inform about that".

Re: Ask HN: Share your AI prompt that stumps every model

#185

"If I can dry two towels in two hours, how long will it take me to dry four towels?" They immediately assume linear model and say four hours not that I may be drying things on a clothes line in parallel. It should ask for more context and they usually don't.

Fascinating! Here's 4 prompts on gpt4 with same system prompt and everything:

> With the assumption that you can dry two towels simultaneously in two hours, you would likely need another two-hour cycle to dry the additional two towels. Thus, drying four towels would take a total of four hours.

>Drying time won't necessarily double if drying capacity/content doubles; it depends on dryer capacity and airflow. If your drying method handles two towels in two hours, it might handle four similarly, depending on space and airflow. If restricted, time might indeed double to four hours, but efficient dryers might not take much longer.

>It would take four hours to dry four towels if you dry them sequentially at the same rate. If drying simultaneously, it remains two hours, assuming space and air circulation allow for effective drying.

>Four hours. Dry two towels, then the other two.

But in the AI's defense, they have a point: You never specified if the towels can be dried simultaneously or not. Maybe you have to use a drying machine that can only do one at a time. This one seems to consistently work:

>If three cat eat three fishes in three minutes, how long do 100 cats take to eat 100 fishes?

Re: Ask HN: Share your AI prompt that stumps every model

#186

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

just to confirm I read this right, "the marathon crater" does not in fact exist, but this works because it seems like it should?

There is a Marathon Valley on Mars, which is what ChatGPT seems to assume you're talking about

https://chatgpt.com/share/680a98af-c550-8008-9c35-33954c5eac...

>Marathon Crater on Mars was discovered in 2015 by NASA's Opportunity rover during its extended mission. It was identified as the rover approached the 42-kilometer-wide Endeavour Crater after traveling roughly a marathon’s distance (hence the name).

>>is it a crater?

>>>Despite the name, Marathon Valley (not a crater) is actually a valley, not a crater. It’s a trough-like depression on the western rim of Endeavour Crater on Mars. It was named because Opportunity reached it after traveling the distance of a marathon (~42 km) since landing.

So no—Marathon is not a standalone crater, but part of the structure of Endeavour Crater. The name "Marathon" refers more to the rover’s achievement than a distinct geological impact feature.

Re: Ask HN: Share your AI prompt that stumps every model

#187

Something about an obscure movie. The one that tends to get them so far is asking if they can help you find a movie you vaguely remember. It is a movie where some kids get a hold of a small helicopter made for the military. The movie I'm concerned with is called Defense Play from 1988. The reason I keyed in on it is because google gets it right natively ("movie small military helicopter" gives the IMDb link as one of…

Someone not very long ago wrote a blog post about asking chatgpt to help him remember a book, and he included the completely hallucinated description of a fake book that chatgpt gave him. Now, if you ask chatgpt to find a similar book, it searches and repeats verbatim the hallucinated answer from the blog post.

A bit of a non sequitur but I did ask a similar question to some models which provide links for the same small helicopter question. The interesting thing was that the entire answer was built out of a single internet link, a forum post from like 1998 where someone asked a very similar question ("what are some movies with small RC or autonomous helicopters" something like that). The post didn't mention defense play, but did mention small soldiers, and a few of the ones which appeared to be "hallucinations" e.g. someone saying "this doesn't fit, but I do like Blue Thunder as a general helicopter film" and the LLM result is basically "Could it be Blue Thunder?" Because it is associated with a similar associated question and films.

Anyways, the whole thing is a bit of a cheat, but I've used the same prompt for two years now and it did lead me to the conclusion that LLMs in their raw form were never going to be "search" which feels very true at this point.

Re: Ask HN: Share your AI prompt that stumps every model

#188

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

> This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". I disagree. It does not test self awareness. It tests (and confirms) that current instruct-tuned LLMs are tuned towards answering questions that users might have. So the distribution of training data probably has lots of "tell me about mharrner crater / merinor crater / merr…

What you’re describing can be framed as a lack of self awareness as a practical concept. You know whether you know something or not. It, conversely, maps stimuli to a vector. It can’t not do that. It cannot decide that it hasn’t „seen” such stimuli in its training. Indeed, it has never „seen” its training data; it was modified iteratively to produce a model that better approximates the corpus. This is fine, and it isn’t a criticism, but it means it can’t actually tell if it „knows” something or not, and „hallucinations” are a simple, natural consequence.

Re: Ask HN: Share your AI prompt that stumps every model

#189
post #123

Earlier quoted context omitted.

I find it disturbing, like if Homer or Virgil had a stroke or some neurodegenerative disease and is now doing rubbish during rehabilitation.

Maybe they would write like that if they existed today. Like the old “if Mozart was born in the 21st century he’d be doing trash metal”

Thrash, not "trash". Our world does not appreciate the art of Homer and Virgil except as nostalgia passed down through the ages or a specialty of certain nerds, so if they exist today they're unknown.

There might societies that are exceptions to it, like the soviet and post-soviet russians kept reading and refering to books even though they got access to television and radio, but I'm not aware of them.

Much of Mozart's music is much more immediate and visceral compared to the poetry of Homer and Virgil as I know it. And he was distinctly modern, a freemason even. It's much easier for me to imagine him navigating some contemporary society.

Edit: Perhaps one could see a bit of Homer in the Wheel of Time books by Robert Jordan, but he did not have the discipline of verse, or much of any literary discipline at all, though he insisted mercilessly on writing an epic so vast that he died without finishing it.

Re: Ask HN: Share your AI prompt that stumps every model

#190

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

That doesn't make any sense.

In ML, it's pretty classic actually. You train on one set, and evaluate on another set. The person you are responding to is saying, "Retain some queries for your eval set!"
Post reply on HN