Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

221–230 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#221
post #209

Earlier quoted context omitted.

Tuning the model output to perform better on certain prompts is not the same as improving the model. It's valid to worry that the model makers are gaming the benchmarks. If you think that's happening and you want to personally figure out which models are really the best, keeping some prompts to yourself is a great way to do that.

Who’s going out of their way to optimize for random HNers informal benchmarks?

Probably anyone training models who also browses HN?

So I would guess every single AI being made currently

Re: Ask HN: Share your AI prompt that stumps every model

#224

Earlier quoted context omitted.

The difference is consistency. You can read a manual and know exactly how to oil and refill the tank on a chainsaw. You can inspect the blades to see if they are worn. You can listen to it and hear how it runs. If a part goes bad, you can easily replace it. If it's having troubles, it will be obvious - it will simply stop working - cutting wood more slowly or not at all. The situation with an LLM is completely differ…

The thing is, nothing you've said is untrue for any search engine or user-driven web site. Only a reckless moron would paste code they find on Stack Overflow or Github into their project without at least looking it over. Same with code written by LLMs. The difference is, just as the LLM can write unit tests to help you deal with uncertainty, it can also cross-check the output of other LLMs. You have to be careful whe…

The nice thing about SO and Github is that there's little to no reason there for things to not work, at least in the context where you found the code. The steps are getting the context, assuming it's true based on various indicators (mostly reputation) and then continuing on with understanding the snippet.

But with LLMs, every word is a probability factor. Assuming the first paragraph is true has no impact on the rest.

Re: Ask HN: Share your AI prompt that stumps every model

#225

"If I can dry two towels in two hours, how long will it take me to dry four towels?" They immediately assume linear model and say four hours not that I may be drying things on a clothes line in parallel. It should ask for more context and they usually don't.

Claude 3.7, Grok 3 DeepThink and QwQ-32B Thinking stil get it wrong!

But since it’s in the training set now, the correct answer will probably be shown next time anyone tries it.

Re: Ask HN: Share your AI prompt that stumps every model

#226

"If I can dry two towels in two hours, how long will it take me to dry four towels?" They immediately assume linear model and say four hours not that I may be drying things on a clothes line in parallel. It should ask for more context and they usually don't.

Gemini 2.5 Pro nails it, as far as I'm concerned:

This sounds like a bit of a riddle! Here's how to think about it:

If you can dry all the towels at the same time: For example, if you have a large enough dryer or enough space on a clothesline, adding more towels doesn't necessarily increase the drying time.

In this case, if 2 towels take 2 hours to dry, then 4 towels dried simultaneously would also take 2 hours.

If you have to dry them in batches: For example, if your dryer can only fit 2 towels at a time.

You would dry the first 2 towels (2 hours). Then you would dry the second 2 towels (another 2 hours). In this scenario, it would take 4 hours in total.

Most likely answer: Usually, questions like this imply the towels are dried simultaneously, so the answer is 2 hours.

Re: Ask HN: Share your AI prompt that stumps every model

#227

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

Yes let's not say what's wrong with the tech, otherwise someone might (gasp) fix it!

"Fix".

So long as the grocery store has groceries, most people will not care what a chat bot spews.

This forum is full of syntax and semantics obsessed loonies who think the symbolic logic represents the truth.

I look forward to being able to use my own creole to manipulate a machine's state to act like a video game or a movie rather than rely on the special literacy of other typical copy-paste middle class people. Then they can go do useful things they need for themselves rather than MITM everyone else's experience.

Re: Ask HN: Share your AI prompt that stumps every model

#228

Write 20 sentences that end with "p"

https://chatgpt.com/share/680a3da0-b888-8013-9c11-42c22a642b...

>20 sentences that end in 'o'

>They shouted cheers after the winning free throw.

good attempt by ChatGPT tho imo

Post reply on HN