Earlier quoted context omitted.
This is my biggest peeve when people say that LLMs are as capable as humans or that we have achieved AGI or are close or things like that. But then when I get a subpar result, they always tell me I'm "prompting wrong". LLMs may be very capable of great human level output, but in my experience leave a LOT to be desired in terms of human level understanding of the question or prompt. I think rating an LLM vs a human or…
I mentioned this in another thread, but this is genuinely demonstrating a known issue with ambiguous prompts. You might be inclined to say, "a human would always interpret the question as having the car nearby the speaker, 50m away from the carwash." But this is objectively untrue. There are people in this comments section and on the Mastodon thread that found the question to be somewhat confusing. In other words, th…
So it feels like a big area of limitation or a big bottleneck towards getting a good answer.