ChatGPT now performs well at GeoGuesser
101–110 of 152 posts
Re: ChatGPT now performs well at GeoGuesser
#102Earlier quoted context omitted.
I uploaded this image that I screenshotted off Google street view (no metadata) and it got with 200m. https://chatgpt.com/share/6801bbf7-fd40-8008-985d-75c8813f55... There is the chat. Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods."
> Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods." That's slightly creepy!
Re: ChatGPT now performs well at GeoGuesser
#103Earlier quoted context omitted.
I uploaded this image that I screenshotted off Google street view (no metadata) and it got with 200m. https://chatgpt.com/share/6801bbf7-fd40-8008-985d-75c8813f55... There is the chat. Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods."
This is the most impressive ChatGPT chat I’ve seen yet. While I theoretically can accept how large-scale probabilistic text generation can lead to this chain of “reasoning”, it really feels like actual intelligence.
Re: ChatGPT now performs well at GeoGuesser
#104Earlier quoted context omitted.
I played a round of Geoguessr against it and while it did a shockingly good job compared to what I was expecting, it still lags behind even novice human players. The locations and its guesses were: Bliss, Idaho - Burns, Oregon (273 miles away) Quilleco, Biobio, Chile - Eugene, Oregon (6,411 miles away) Dettighofen, Switzerland - Mühldorf, Germany (228 miles away) Pretoria, South Africa - Johannesburg, South Africa (3…
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0
Re: ChatGPT now performs well at GeoGuesser
#105Re: ChatGPT now performs well at GeoGuesser
#106Chatbots appear to have some amount of fluid intellgence so they can do impressive tasks with this information, the impressiveness of these tasks will likely increase in the future. But for simply getting a good score on Geoguesser it's not even close to hobby projects let alone state of the art.
Re: ChatGPT now performs well at GeoGuesser
#107Earlier quoted context omitted.
I uploaded this image that I screenshotted off Google street view (no metadata) and it got with 200m. https://chatgpt.com/share/6801bbf7-fd40-8008-985d-75c8813f55... There is the chat. Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods."
> Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods." That's slightly creepy!
Tell it your name and then it just looks you up and street views your house, and puts that all into memory.
Re: ChatGPT now performs well at GeoGuesser
#108Meanwhile O3 can not even count rocks in a picture. This is a commonly recurring theme -- ChatGPT does really well at few things considered hard by us but fails miserably at things even a child could do.
Is that really so surprising, considering the tremendous differences in underlying hardware and training process?
Re: ChatGPT now performs well at GeoGuesser
#109It wouldn’t shock me if multimodal LLMs were good at GeoGuesser [0], but if we’re being picky, it takes more than a few examples to demonstrate a game is “solved.” I also wonder what kind of data leakage might have been at play, like other people have suggested. To be clear, my point is not that this is unimpressive, just that this doesn’t demonstrate much. (Edit: I should have said, it doesn’t demonstrate what the t…
The examples are cherry-picked. I took a photo outside my office window in a built-up area, o3 thought for 5m 7s (!), and it got the location wrong by 40km. Doesn't look solved to me.
Re: ChatGPT now performs well at GeoGuesser
#110Meanwhile O3 can not even count rocks in a picture. This is a commonly recurring theme -- ChatGPT does really well at few things considered hard by us but fails miserably at things even a child could do.
It's almost as if this is a non-human intelligence, which presents different strengths and weaknesses than human intelligence. Is that really so surprising, considering the tremendous differences in underlying hardware and training process?