It wouldn’t shock me if multimodal LLMs were good at GeoGuesser [0], but if we’re being picky, it takes more than a few examples to demonstrate a game is “solved.” I also wonder what kind of data leakage might have been at play, like other people have suggested. To be clear, my point is not that this is unimpressive, just that this doesn’t demonstrate much. (Edit: I should have said, it doesn’t demonstrate what the t…
The examples are cherry-picked. I took a photo outside my office window in a built-up area, o3 thought for 5m 7s (!), and it got the location wrong by 40km. Doesn't look solved to me.
ChatGPT now performs well at GeoGuesser
131–140 of 152 posts
Re: ChatGPT now performs well at GeoGuesser
#132Re: ChatGPT now performs well at GeoGuesser
#133I've been telling women to keep copies of all the dick pics they get sent. Since you can tell by the characteristic noise of a cameras sensor which other pictures were taken with the same camera. All missing is a search engine capable of doing this. I feel with AI, we are 2-3 years away from people uploading a dick pic to AI and getting the social media profile of that person...
C'mon boys. Start uploading those dick pics for research purposes.
Re: ChatGPT now performs well at GeoGuesser
#134Earlier quoted context omitted.
One problem is how can you even set up a "fair" competition between an AI and Rainbolt? He does ones where it flashes for a fraction of a second and then he guesses the country. How do you simulate "only saw it for a fraction of a second" to an AI?
Maybe limit the time the AI is allowed to think? In the post it showed the AI thought for almost a minute. I’ve seen Rainbolt ID an image based on some dirt and nothing else. I’d want to see AI be able to do that before saying it’s a solved problem.
Re: ChatGPT now performs well at GeoGuesser
#135Earlier quoted context omitted.
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0
What about 04-mini-high ?
Re: ChatGPT now performs well at GeoGuesser
#136Earlier quoted context omitted.
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0
How does Google Lens compare?
Re: ChatGPT now performs well at GeoGuesser
#137Re: ChatGPT now performs well at GeoGuesser
#138Earlier quoted context omitted.
The military/infosec uses of this are not real time. You can wait 5 minutes for a drone strike.
Yeah but you also cant be off by 40km with your drone strike.
Re: ChatGPT now performs well at GeoGuesser
#139I gave it a series of 11 images stripped of all metadata. It performed quite well, only misidentifying the two taken in a small college town in the NE of the US. It got two questions correct on photos taken in Korea (one with a fairly clear view of Haneul Park, the other a rather difficult to identify picture not resembling anything on google of Sunrise Peak). It got every other question in the US correct, ranging fr…
It eventually decided that the photo was taken outside the Scottish Exhibition and Conference Centre in Glasgow. It actually generally considered Scottish locations more than others.
The picture was actually taken in Plymouth (so pretty much as far from Scotland as you can get in Britain), on Charles Street looking south-east[2]. The building on the right is Drake Circus, and the one on the left is the Arts University. It actually did consider Plymouth, but decided it didn't match.
[0] This image with the "university plymouth" on the left cropped out, just to make it harder: https://www.facebook.com/photo/?fbid=9719044988151697&set=gm...
[1] https://chatgpt.com/share/68024c91-61d0-800c-99b1-fcecf0bfe8...
Re: ChatGPT now performs well at GeoGuesser
#140Meanwhile O3 can not even count rocks in a picture. This is a commonly recurring theme -- ChatGPT does really well at few things considered hard by us but fails miserably at things even a child could do.
It's almost as if this is a non-human intelligence, which presents different strengths and weaknesses than human intelligence. Is that really so surprising, considering the tremendous differences in underlying hardware and training process?