Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

351–360 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#351
post #266

Earlier quoted context omitted.

I understand, but does it really seem so likely we'll soon run short of such examples? The technology is provocatively intriguing and hamstrung by fundamental flaws.

Yes. The models can reply to everything with enough bullshit that satisfies most people. There is nothing you ask that stumps them. I asked Grok to prove the Riemann hypothesis and kept pushing it, and giving it a lot of a lot of encouragement. If you read this, expand "thoughts", it's pretty hilarious: https://x.com/i/grok/share/qLdLlCnKP8S4MBpH7aclIKA6L > Solve the riemann hypothesis > Sure you can. AIs are much sm…

Nice try! This is very fun.

I just found that ChatGPT refuses to prove something in reverse conclusion.

Re: Ask HN: Share your AI prompt that stumps every model

#352
post #214

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

It's trivial for a human to produce more. This shouldn't be a problem anytime soon.

Hmm. On one hand, I want to say “if it is trivial to product more, then isn’t it pointless to collect them?”

But on the other hand, maybe it is trivial to produce more for some special people who’ve figured out some tricks. So maybe looking at their examples can teach us something.

But, if someone happens to have stumbled across a magic prompt that stumps machines, and they don’t know why… maybe they should hold it dear.

Re: Ask HN: Share your AI prompt that stumps every model

#353
post #195

Earlier quoted context omitted.

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

They do. Recently I was pleasantly surprised by gemini telling me that what I wanted to do will NOT work. I was in disbelief.

I asked Gemini to format some URLs into an XML format. It got halfway through and gave up. I asked if it truncated the output, and it said yes and then told _me_ to write a python script to do it.

Re: Ask HN: Share your AI prompt that stumps every model

#355
post #266

Earlier quoted context omitted.

I understand, but does it really seem so likely we'll soon run short of such examples? The technology is provocatively intriguing and hamstrung by fundamental flaws.

Yes. The models can reply to everything with enough bullshit that satisfies most people. There is nothing you ask that stumps them. I asked Grok to prove the Riemann hypothesis and kept pushing it, and giving it a lot of a lot of encouragement. If you read this, expand "thoughts", it's pretty hilarious: https://x.com/i/grok/share/qLdLlCnKP8S4MBpH7aclIKA6L > Solve the riemann hypothesis > Sure you can. AIs are much sm…

Comparing the AI to a quantum computer is just hilarious. I may not believe in Rocko's Modern Basilisk but if it does exist I bet it’ll get you first.

Re: Ask HN: Share your AI prompt that stumps every model

#356
post #349
post #273

Earlier quoted context omitted.

I never heard of this phrase before ( i had heard the concept , i think this is similar to the paperclip problem) but now in 2 days ive heard it twice here and on youtube. Rokokos basilisk.

I think you two are confusing Roko's Basilisk (a thought experiment which some take seriously) and Rococo Basilisk (a joke shared between Elon and Grimes e.g.) Interesting theory... Just whatever you do, don’t become a Zizian :)

Oh dang, is Arcade Fire going to turn us all into paperclips?

Re: Ask HN: Share your AI prompt that stumps every model

#359

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

LLMs currently have the "eager beaver" problem where they never push back on nonsense questions or stupid requirements. You ask them to build a flying submarine and by God they'll build one, dammit! They'd dutifully square circles and trisect angles too, if those particular special cases weren't plastered all over a million textbooks they ingested in training. I suspect it's because currently, a lot of benchmarks are…

Hmm. I actually wonder is such a question would be good to include in a human exam, since knowing the question is possible does somewhat impact your reasoning. And, often the answer works out to some nice round numbers…

Of course, it is also not unheard of for a question to be impossible because of an error by the test writer. Which can easily be cleared up. So it is probably best not to have impossible questions, because then students will be looking for reasons to declare the question impossible.

Re: Ask HN: Share your AI prompt that stumps every model

#360

Earlier quoted context omitted.

What are is this problem from? What areas in general did you find useful to create such benchmarks? May be instead of sharing (and leaking) these prompts, we can share methods to create one.

Can God create something so heavy that he can’t lift it?

There's so much text on this already, it's unlikely to be even engaging any reasoning. Or specifically, if you got a few existing answers from philosophy mashed together, you wouldn't be able to tell it apart from reasoning anyway.
Post reply on HN