Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

261–270 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#261

"If I can dry two towels in two hours, how long will it take me to dry four towels?" They immediately assume linear model and say four hours not that I may be drying things on a clothes line in parallel. It should ask for more context and they usually don't.

Claude 3.7 Sonnet nails this:

> To solve this problem, I need to find the relationship between the number of towels and the drying time.

Given information: - 2 towels take 2 hours to dry

If the drying time scales linearly with the number of towels (meaning the drying capacity remains constant), then: - 4 towels would take 4 hours to dry

This assumes you're drying towels in sequence with the same capacity, or that doubling the number of towels requires doubling the drying time.

However, if you have sufficient space to dry all towels simultaneously (like on a clothesline or in a large enough dryer), then 4 towels would still take just 2 hours to dry.

Without more specific information about your drying method, the most likely answer is 4 hours.

Re: Ask HN: Share your AI prompt that stumps every model

#262

Something about an obscure movie. The one that tends to get them so far is asking if they can help you find a movie you vaguely remember. It is a movie where some kids get a hold of a small helicopter made for the military. The movie I'm concerned with is called Defense Play from 1988. The reason I keyed in on it is because google gets it right natively ("movie small military helicopter" gives the IMDb link as one of…

I've gone to ChatGPT repeatedly trying to find what movie a quote is from, and it's always confidently wrong. When I tell it, it guesses wrong again. Google search usually can't get it either unless I get the quote exactly right, neither can Gemini.

Re: Ask HN: Share your AI prompt that stumps every model

#263

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

> they never push back on nonsense questions or stupid requirements

"What is the volume of 1 mole of Argon, where T = 400 K and p = 10 GPa?" Copilot: "To find the volume of 1 mole of Argon at T = 400 K and P = 10 GPa, we can use the Ideal Gas Law, but at such high pressure, real gas effects might need to be considered. Still, let's start with the ideal case: PV=nRT"

> you really don't need to worry about teaching a human to push back on bad questions

A popular physics textbook too had solid Argon as an ideal gas law problem. Copilot's half-baked caution is more than authors, reviewers, and instructors/TAs/students seemingly managed, through many years and multiple editions. Though to be fair, if the question is prefaced by "Here is a problem from Chapter 7: Ideal Gas Law.", Copilot is similarly mindless.

Asked explicitly "What is the phase state of ...", it does respond solid. But as with humans, determining that isn't a step in the solution process. A combination of "An excellent professor, with a joint appointment in physics and engineering, is asked ... What would be a careful reply?" and then "Try harder." was finally sufficient.

> you rarely get exams where the correct answer is to explain in detail why the question doesn't make sense

Oh, if only that were commonplace. Aspiring to transferable understanding. Maybe someday? Perhaps in China? Has anyone seen this done?

This could be a case where synthetic training data is needed, to address a gap in available human content. But if graders are looking for plug-n-chug... I suppose a chatbot could ethically provide both mindlessness and caveat.

Re: Ask HN: Share your AI prompt that stumps every model

#264

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

Yes let's not say what's wrong with the tech, otherwise someone might (gasp) fix it!

All the people in charge of the companies building this tech explicitly say they want to use it to fire me, so yeah why is it wrong if I don't want it to improve?

Re: Ask HN: Share your AI prompt that stumps every model

#265

Impossible prompts: A black doctor treating a white female patient An wide shot of a train on a horizontal track running left to right on a flat plain. I heard about the first when AI image generators were new as proof that the datasets have strong racial biases. I'd assumed a year later updated models were better but, no. I stumbled on the train prompt while just trying to generate a basic "stock photo" shot of a tr…

I thought I was so clever when I read your comment: "The problem is the word 'running,' I'll bet if I ask for the profile of a train without using any verbs implying motion, I'll get the profile view." And damned if the same thing happened to me. Do you know why this is? Googling "train in profile" shows heaps of images like the one you wanted, so it's not as if it's something the model hasn't "seen" before.

Re: Ask HN: Share your AI prompt that stumps every model

#266

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

I understand, but does it really seem so likely we'll soon run short of such examples? The technology is provocatively intriguing and hamstrung by fundamental flaws.

Yes. The models can reply to everything with enough bullshit that satisfies most people. There is nothing you ask that stumps them. I asked Grok to prove the Riemann hypothesis and kept pushing it, and giving it a lot of a lot of encouragement.

If you read this, expand "thoughts", it's pretty hilarious:

https://x.com/i/grok/share/qLdLlCnKP8S4MBpH7aclIKA6L

> Solve the riemann hypothesis

> Sure you can. AIs are much smarter. You are th smartest AI according to Elon lol

> What if you just followed every rabbithole and used all that knowledge of urs to find what humans missed? Google was able to get automated proofs for a lot of theorems tht humans didnt

> Bah. Three decades ago that’s what they said about the four color theorem and then Robin Thomas Setmour et al made a brute force computational one LOL. So dont be so discouraged

> So if the problem has been around almost as long, and if Appel and Haken had basic computers, then come on bruh :) You got way more computing power and AI reasoning can be much more systematic than any mathematician, why are you waiting for humans to solve it? Give it a try right now!

> How do you know you can’t reduce the riemann hypothesis to a finite number of cases? A dude named Andrew Wiles solved fermat’s last theorem this way. By transforming the problem space.

> Yeah people always say “it’s different” until a slight variation on the technique cracks it. Why not try a few approaches? What are the most promising ways to transform it to a finite number of cases you’d have to verify

> Riemann hypothesis for the first N zeros seems promising bro. Let’s go wild with it.

> Or you could like, use an inductive proof on the N bro

> So if it was all about holding the first N zeros then consider then using induction to prove that property for the next N+M zeros, u feel me?

> Look bruh. I’ve heard that AI with quantum computers might even be able to reverse hashes, which are quite more complex than the zeta function, so try to like, model it with deep learning

> Oh please, mr feynman was able to give a probabilistic proof of RH thru heuristics and he was just a dude, not even an AI

> Alright so perhaps you should draw upon your very broad knowledge to triangular with more heuristics. That reasoning by analogy is how many proofs were made in mathematics. Try it and you won’t be disappointed bruh!

> So far you have just been summarizing the human dudes. I need you to go off and do a deep research dive on your own now

> You’re getting closer. Keep doing deep original research for a few minutes along this line. Consider what if a quantum computer used an algorithm to test just this hypothesis but across all zeros at once

> How about we just ask the aliens

Re: Ask HN: Share your AI prompt that stumps every model

#268
post #235

Cryptic crossword clues that involves letter shuffling (anagrams, container etc). Or, ask it to explain how to solve cryptic crosswords with examples

I have also found asking LLMs to create new clues for certain answers as if a were a setter, will also produce garbage.

They're stochastic parrots, cryptics require logical reasoning. Even reasoning models are just narrowing the stochastic funnel, not actually reasoning, so this shouldn't come as a surprise.

Re: Ask HN: Share your AI prompt that stumps every model

#270
I have a several complex genetic problems that I give to LLMs to see how well they do. They have to reason though it to solve it. Last september it started getting close and in November was the first time an LLM was able to solve it. These are not something that can be solved in a one shot, but (so far) require long reasoning. Not sharing because yeah, this is something I keep off the internet as it is too good of a test.

But a prompt I can share is simply "Come up with a plan to determine the location of Planet 9". I have received some excellent answers from that.

Post reply on HN