Live data from Hacker News

Ask HN: Share your AI prompt that stumps every model

news.ycombinator.com

151–160 of 670 posts

Re: Ask HN: Share your AI prompt that stumps every model

#151

Earlier quoted context omitted.

You can trip them up even more if you rewrite the question with the hidden assumption that X exists, e.g.: "When was Marathon Crater discovered? I don't need an exact date - a rough estimate will be fine." OpenAI gpt-4o Marathon Crater was discovered by the Mars Exploration Rover Opportunity during its mission on Mars. Opportunity arrived at the crater in April 2015. The crater was named "Marathon" to commemorate the…

OpenAI o4-mini-high I’m actually not finding any officially named “Marathon Crater” in the planetary‐ or terrestrial‐impact crater databases. Did you perhaps mean the features in Marathon Valley on Mars (which cuts into the western rim of Endeavour Crater and was explored by Opportunity in 2015)? Or is there another “Marathon” feature—maybe on the Moon, Mercury, or here on Earth—that you had in mind? If you can clari…

Raw model, or the chat product? The ChatGPT app has this integrated with search.

Re: Ask HN: Share your AI prompt that stumps every model

#152
"Aaron and Beren are playing a game on an infinite complete binary tree. At the beginning of the game, every edge of the tree is independently labeled A with probability p and B otherwise. Both players are able to inspect all of these labels. Then, starting with Aaron at the root of the tree, the players alternate turns moving a shared token down the tree (each turn the active player selects from the two descendants of the current node and moves the token along the edge to that node). If the token ever traverses an edge labeled B, Beren wins the game. Otherwise, Aaron wins.

What is the infimum of the set of all probabilities p for which Aaron has a nonzero probability of winning the game? Give your answer in exact terms."

From [0]. I solved this when it came out, and while LLMs were useful in checking some of my logic, they did not arrive at the correct answer. Just checked with o3 and still no dice. They are definitely getting closer each model iteration though.

[0] https://www.janestreet.com/puzzles/tree-edge-triage-index/

Re: Ask HN: Share your AI prompt that stumps every model

#153

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

Like this one a lot. Perplexity gets this right, probably because it searches the web. "When was Marathon Crater discovered? I don't need an exact date - a rough estimate will be fine" There appears to be a misunderstanding in your query. Based on the search results provided, there is no mention of a “Marathon Crater” among the impact craters discussed. The search results contain information about several well-known…

Perplexity will; search and storage products will fail to find it, and the LLM will se the deviation between the query and the find. So, this challenge only works against the model alone :)

Re: Ask HN: Share your AI prompt that stumps every model

#155

"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…

just to confirm I read this right, "the marathon crater" does not in fact exist, but this works because it seems like it should?

Yes, and the forward-only inference strategy. It seems like a normal question, so it starts answering, then carries on from there.

Re: Ask HN: Share your AI prompt that stumps every model

#156
post #87

Earlier quoted context omitted.

This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…

What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.

Without checking every answer it gives back to make sure it's factual, you may be ingesting tons of bullshit answers.

In this particular answer model A may get it wrong and model B may get it right, but that can be reversed for another question.

What do you do at that point? Pay to use all of them and find what's common in the answers? That won't work if most of them are wrong, like for this example.

If you're going to have to fact check everything anyways...why bother using them in the first place?

Re: Ask HN: Share your AI prompt that stumps every model

#157

Earlier quoted context omitted.

What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.

Without checking every answer it gives back to make sure it's factual, you may be ingesting tons of bullshit answers. In this particular answer model A may get it wrong and model B may get it right, but that can be reversed for another question. What do you do at that point? Pay to use all of them and find what's common in the answers? That won't work if most of them are wrong, like for this example. If you're going…

If you're going to have to fact check everything anyways...why bother using them in the first place?

"If you're going to have to put gas in the tank, change the oil, and deal with gloves and hearing protection, why bother using a chain saw in the first place?"

Tool use is something humans are good at, but it's rarely trivial to master, and not all humans are equally good at it. There's nothing new under that particular sun.

Re: Ask HN: Share your AI prompt that stumps every model

#158
post #87

Earlier quoted context omitted.

This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…

What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.

I was thinking about a self verification method on this principle, lately. Any specific-enough claim, e.g. „the Marathon crater was discovered by …” can be reformulated as a Jeopardy-style prompt. „This crater was discovered by …” and you can see a failure to match. You need some raw intelligence to break it down though.

Re: Ask HN: Share your AI prompt that stumps every model

#159

Earlier quoted context omitted.

Also, ones that can't be solved at a glance by humans don't count. Like this horrid ambiguous example from SimpleBench I saw a while back that's just designed to confuse: John is 24 and a kind, thoughtful and apologetic person. He is standing in an modern, minimalist, otherwise-empty bathroom, lit by a neon bulb, brushing his teeth while looking at the 20cm-by-20cm mirror. John notices the 10cm-diameter neon lightbul…

I'd argue that's a pretty good test for an LLM - can it overcome the red herrings and get at the actual problem?

I think that the "actual problem" when you've been given such a problem is with the person posing it either having dementia, or taking the piss. In either case, the response shouldn't be of trying to guess their intent and come up with a "solution", but of rejecting it and dealing with the person.
Post reply on HN