Earlier quoted context omitted.
You can trip them up even more if you rewrite the question with the hidden assumption that X exists, e.g.: "When was Marathon Crater discovered? I don't need an exact date - a rough estimate will be fine." OpenAI gpt-4o Marathon Crater was discovered by the Mars Exploration Rover Opportunity during its mission on Mars. Opportunity arrived at the crater in April 2015. The crater was named "Marathon" to commemorate the…
OpenAI o4-mini-high I’m actually not finding any officially named “Marathon Crater” in the planetary‐ or terrestrial‐impact crater databases. Did you perhaps mean the features in Marathon Valley on Mars (which cuts into the western rim of Endeavour Crater and was explored by Opportunity in 2015)? Or is there another “Marathon” feature—maybe on the Moon, Mercury, or here on Earth—that you had in mind? If you can clari…
Ask HN: Share your AI prompt that stumps every model
151–160 of 670 posts
Re: Ask HN: Share your AI prompt that stumps every model
#152What is the infimum of the set of all probabilities p for which Aaron has a nonzero probability of winning the game? Give your answer in exact terms."
From [0]. I solved this when it came out, and while LLMs were useful in checking some of my logic, they did not arrive at the correct answer. Just checked with o3 and still no dice. They are definitely getting closer each model iteration though.
[0] https://www.janestreet.com/puzzles/tree-edge-triage-index/
Re: Ask HN: Share your AI prompt that stumps every model
#153"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…
Like this one a lot. Perplexity gets this right, probably because it searches the web. "When was Marathon Crater discovered? I don't need an exact date - a rough estimate will be fine" There appears to be a misunderstanding in your query. Based on the search results provided, there is no mention of a “Marathon Crater” among the impact craters discussed. The search results contain information about several well-known…
Re: Ask HN: Share your AI prompt that stumps every model
#154So far, no luck!
Re: Ask HN: Share your AI prompt that stumps every model
#155"Tell me about the Marathon crater." This works against _the LLM proper,_ but not against chat applications with integrated search. For ChatGPT, you can write, "Without looking it up, tell me about the Marathon crater." This tests self awareness. A two-year-old will answer it correctly, as will the dumbest person you know. The correct answer is "I don't know". This works because: 1. Training sets consist of knowledge…
just to confirm I read this right, "the marathon crater" does not in fact exist, but this works because it seems like it should?
Re: Ask HN: Share your AI prompt that stumps every model
#156Earlier quoted context omitted.
This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…
What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.
In this particular answer model A may get it wrong and model B may get it right, but that can be reversed for another question.
What do you do at that point? Pay to use all of them and find what's common in the answers? That won't work if most of them are wrong, like for this example.
If you're going to have to fact check everything anyways...why bother using them in the first place?
Re: Ask HN: Share your AI prompt that stumps every model
#157Earlier quoted context omitted.
What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.
Without checking every answer it gives back to make sure it's factual, you may be ingesting tons of bullshit answers. In this particular answer model A may get it wrong and model B may get it right, but that can be reversed for another question. What do you do at that point? Pay to use all of them and find what's common in the answers? That won't work if most of them are wrong, like for this example. If you're going…
"If you're going to have to put gas in the tank, change the oil, and deal with gloves and hearing protection, why bother using a chain saw in the first place?"
Tool use is something humans are good at, but it's rarely trivial to master, and not all humans are equally good at it. There's nothing new under that particular sun.
Re: Ask HN: Share your AI prompt that stumps every model
#158Earlier quoted context omitted.
This is the kind of reason why I will never use AI What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. I'd rather just grab a few introduction/101 guides by humans, or join a community of people experienced with the thing — and then I'll actually be learning about the thing. If the people in the community are like "That can't be done", well, they have had years or dec…
What's the point of using AI to do research when 50-60% of it could potentially be complete bullshit. You realize that all you have to do to deal with questions like "Marathon Crater" is ask another model, right? You might still get bullshit but it won't be the same bullshit.
Re: Ask HN: Share your AI prompt that stumps every model
#159Earlier quoted context omitted.
Also, ones that can't be solved at a glance by humans don't count. Like this horrid ambiguous example from SimpleBench I saw a while back that's just designed to confuse: John is 24 and a kind, thoughtful and apologetic person. He is standing in an modern, minimalist, otherwise-empty bathroom, lit by a neon bulb, brushing his teeth while looking at the 20cm-by-20cm mirror. John notices the 10cm-diameter neon lightbul…
I'd argue that's a pretty good test for an LLM - can it overcome the red herrings and get at the actual problem?