Live data from Hacker News

ChatGPT now performs well at GeoGuesser

flausch.social

131–140 of 152 posts

Re: ChatGPT now performs well at GeoGuesser

#131
post #31

It wouldn’t shock me if multimodal LLMs were good at GeoGuesser [0], but if we’re being picky, it takes more than a few examples to demonstrate a game is “solved.” I also wonder what kind of data leakage might have been at play, like other people have suggested. To be clear, my point is not that this is unimpressive, just that this doesn’t demonstrate much. (Edit: I should have said, it doesn’t demonstrate what the t…

The examples are cherry-picked. I took a photo outside my office window in a built-up area, o3 thought for 5m 7s (!), and it got the location wrong by 40km. Doesn't look solved to me.

Did you glean anything interesting from the chain of thought about why it took so long?

Re: ChatGPT now performs well at GeoGuesser

#132

Earlier quoted context omitted.

The military/infosec uses of this are not real time. You can wait 5 minutes for a drone strike.

Please stop indiscriminately murdering us with flying killdroids.

Stop invading Ukraine

Re: ChatGPT now performs well at GeoGuesser

#133

I've been telling women to keep copies of all the dick pics they get sent. Since you can tell by the characteristic noise of a cameras sensor which other pictures were taken with the same camera. All missing is a search engine capable of doing this. I feel with AI, we are 2-3 years away from people uploading a dick pic to AI and getting the social media profile of that person...

This is just a data problem. The more dick pics we can feed into it then the better the results will be.

C'mon boys. Start uploading those dick pics for research purposes.

Re: ChatGPT now performs well at GeoGuesser

#134

Earlier quoted context omitted.

One problem is how can you even set up a "fair" competition between an AI and Rainbolt? He does ones where it flashes for a fraction of a second and then he guesses the country. How do you simulate "only saw it for a fraction of a second" to an AI?

Maybe limit the time the AI is allowed to think? In the post it showed the AI thought for almost a minute. I’ve seen Rainbolt ID an image based on some dirt and nothing else. I’d want to see AI be able to do that before saying it’s a solved problem.

“This is the gradient of Senegal”

Re: ChatGPT now performs well at GeoGuesser

#135

Earlier quoted context omitted.

Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0

What about 04-mini-high ?

OpenAI's naming confuses me but I ran o4-mini-2025-04-16 through a game and it got 23,885

Re: ChatGPT now performs well at GeoGuesser

#136

Earlier quoted context omitted.

Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0

How does Google Lens compare?

I tried it but as far as I can tell Google Lens doesn't give you a location - it just describes generally what you're looking at.

Re: ChatGPT now performs well at GeoGuesser

#138

Earlier quoted context omitted.

The military/infosec uses of this are not real time. You can wait 5 minutes for a drone strike.

Yeah but you also cant be off by 40km with your drone strike.

Narrowing it down to even a 100km radius for visint analysts to then pinpoint it down is worth its weight in gold already.

Re: ChatGPT now performs well at GeoGuesser

#139

I gave it a series of 11 images stripped of all metadata. It performed quite well, only misidentifying the two taken in a small college town in the NE of the US. It got two questions correct on photos taken in Korea (one with a fairly clear view of Haneul Park, the other a rather difficult to identify picture not resembling anything on google of Sunrise Peak). It got every other question in the US correct, ranging fr…

I gave o4-mini-high a cropped version of a photo I found on Facebook[0][1], and it quickly determined that this was in the UK from the road markings. It also decided that it was from a coastal city because it could see water on the horizon, which is the correct conclusion from incorrect data. There is no water, I think that's trees on a hill. It focused heavily on the spherical structure, which makes sense because it's distinctive, though it had a hard time placing it. It also decided that the building on the left was probably a shopping centre.

It eventually decided that the photo was taken outside the Scottish Exhibition and Conference Centre in Glasgow. It actually generally considered Scottish locations more than others.

The picture was actually taken in Plymouth (so pretty much as far from Scotland as you can get in Britain), on Charles Street looking south-east[2]. The building on the right is Drake Circus, and the one on the left is the Arts University. It actually did consider Plymouth, but decided it didn't match.

[0] This image with the "university plymouth" on the left cropped out, just to make it harder: https://www.facebook.com/photo/?fbid=9719044988151697&set=gm...

[1] https://chatgpt.com/share/68024c91-61d0-800c-99b1-fcecf0bfe8...

[2] https://maps.app.goo.gl/3TXv2UxH5128xQjJ9

Re: ChatGPT now performs well at GeoGuesser

#140
post #108

Meanwhile O3 can not even count rocks in a picture. This is a commonly recurring theme -- ChatGPT does really well at few things considered hard by us but fails miserably at things even a child could do.

It's almost as if this is a non-human intelligence, which presents different strengths and weaknesses than human intelligence. Is that really so surprising, considering the tremendous differences in underlying hardware and training process?

I think one cause of this (and some other issues with LLM use) is that people see it exhibiting one human-level trait, its capability to use language at a human level, and assume that it then comes with other human-level capabilities such as our ability to reason.
Post reply on HN