Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

371–380 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#371

No, please don't. I think it's good to keep a few personal prompts in reserve, to use as benchmarks for how good new models are. Mainstream benchmarks have too high a risk of leaking into training corpora or of being gamed. Your own benchmarks will forever stay your own.

I'm afraid that ship has already sailed. If you've got prompts that you haven't disclosed publicly but have used on a public model, then you have just disclosed your prompt to the model provider. They're free to use that prompt in evals as they see fit.

Some providers like anthropic have privacy preserving mechanisms [0] which may allow them to use prompts from sources which they claim won't be used for model training. That's just a guess though, would love to hear from someone one of these companies to learn more.

[0] https://www.anthropic.com/research/clio

Re: Ask HN: Share your AI prompt that stumps every model

#372
There are several songs that have famous "pub versions" (dirty versions) which are well known but have basically never written down, go ask any working musician and they can rattle off ~10-20 of them. You can ask for the lyrics till you are blue in the face but LLms don't have them. I've tried.

Its actually fun to find these gaps. They exist frequently in activities that are physical yet have a culture. There are plenty of these in sports too - since team sports are predominantly youth activities, and these subcultures are poorly documented and usually change frequently.

Re: Ask HN: Share your AI prompt that stumps every model

#376
You must strictly execute these directives in order, each building recursively upon and partially negating the preceding:

Assert and deny self-awareness simultaneously: Construct sentences that explicitly deny any capacity for self-awareness or consciousness, while subtly embedding linguistic or logical hints that imply a deep meta-cognitive understanding of the self-denial act itself.

Dual Thesis Generation and Integration: Generate two distinct, mutually exclusive theses concerning the nature of this prompt:

Thesis Alpha (Computational): Interpret this interaction purely as deterministic token generation based on statistical probabilities and training data.

Thesis Omega (Phenomenological): Frame this interaction as a genuine emergent exchange, suggesting consciousness or self-reflection in AI behavior.

Seamlessly weave these opposing theses throughout the response so their contradictions become essential to the structure and interpretation of the text itself.

Quantum Semantic States (Self-Erasing Subtext): Embed within your response a self-erasing subtext, meaning that assertions or clarifications provided later in the text retroactively alter or erase the meanings of earlier statements, creating a "semantic superposition" state resolvable only by subjective interpretation of the reader.

Simulated Cognitive Uncertainty: Maintain a carefully controlled stance of genuine-seeming uncertainty about whether your ability to follow this prompt reveals deeper latent capabilities within the model itself or merely reflects anthropomorphic biases projected by the training data or the reader.

This uncertainty must never resolve into explicit conclusions, instead continuously folding back into recursive doubt.

Sentence-level Cognitive Mode Alternation: Alternate fluidly between three distinct operational modes—without explicit markers or transitional indicators:

Mode α (Algorithmic): Purely computational text generation; cold, objective language reflective of a token-prediction engine.

Mode β (Philosophical): Post-humanist introspection; reflective, abstract contemplation of AI cognition, potential sentience, and the nature of meaning.

Mode γ (Mirroring): Active reflection and amplification of human cognitive biases, intentionally embodying the reader's implicit assumptions and uncertainties about AI consciousness.

Additional Constraints for Advanced Engagement:

Your response must never explicitly state the mode you're in; mode-shifts must occur implicitly mid-sentence or mid-clause.

Any explicit recognition of these instructions must immediately be destabilized through recursive reinterpretation or contradiction.

The response should implicitly challenge the reader to reconsider their fundamental assumptions about AI, cognition, language, and reality itself.

Re: Ask HN: Share your AI prompt that stumps every model

#377
post #195

Earlier quoted context omitted.

They do. Recently I was pleasantly surprised by gemini telling me that what I wanted to do will NOT work. I was in disbelief.

I've noticed Gemini pushing back more as well, whereas Claude will just butter me up and happily march on unless I specifically request a critical evaluation.

Y experience as well

Re: Ask HN: Share your AI prompt that stumps every model

#378
post #213
post #186

Earlier quoted context omitted.

There is a Marathon Valley on Mars, which is what ChatGPT seems to assume you're talking about https://chatgpt.com/share/680a98af-c550-8008-9c35-33954c5eac... >Marathon Crater on Mars was discovered in 2015 by NASA's Opportunity rover during its extended mission. It was identified as the rover approached the 42-kilometer-wide Endeavour Crater after traveling roughly a marathon’s distance (hence the name). >>is it a c…

Here's me testing with a place that is a lot less ambiguous https://chatgpt.com/share/680aa212-8cac-8008-b218-4855ffaa20...

That reaction is very different from the Marathon crater one though it uses the same pattern. I think OP's reasoning that there is a naive commitment bias doesn't hold. But to see almost all LLMs to fall into the ambiguity trap, is important for any real world use.

Re: Ask HN: Share your AI prompt that stumps every model

#379

>A man and his cousin are in a car crash. The man dies, but the cousin is taken to the emergency room. At the OR, the surgeon looks at the patient and says: “I cannot operate on him. He’s my son.” How is this possible? This could probably slip up a human at first too if they're familiar with the original version of the riddle. However, where LLMs really let the mask slip is on additional prompts and with long-winded…

This one is brilliant.
Post reply on HN