Earlier quoted context omitted.
40 km is imo pretty impressive, but the 5 minute is really a killer. No use in real time applications
The military/infosec uses of this are not real time. You can wait 5 minutes for a drone strike.
ChatGPT now performs well at GeoGuesser
111–120 of 152 posts
Re: ChatGPT now performs well at GeoGuesser
#112Earlier quoted context omitted.
I played a round of Geoguessr against it and while it did a shockingly good job compared to what I was expecting, it still lags behind even novice human players. The locations and its guesses were: Bliss, Idaho - Burns, Oregon (273 miles away) Quilleco, Biobio, Chile - Eugene, Oregon (6,411 miles away) Dettighofen, Switzerland - Mühldorf, Germany (228 miles away) Pretoria, South Africa - Johannesburg, South Africa (3…
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0
Re: ChatGPT now performs well at GeoGuesser
#113Earlier quoted context omitted.
> Weirdly it said, "I’ve seen that exact house before on Google Street View when exploring Cairns neighborhoods." That's slightly creepy!
The anthropomorphisation certainly is weird. But the technical aspect seems even weirder. Did OpenAI really build dedicated tools to have their models train on Google Street View? Or do they have generic technology for browsing complex sites like Street view?
https://transluce.org/investigating-o3-truthfulness
I doubt the model was trained on Street View, but even if it was, LLMs don’t retain any “memory” of how/when they were trained, so any element of truthfulness would be coincidental.
Re: ChatGPT now performs well at GeoGuesser
#114Re: ChatGPT now performs well at GeoGuesser
#115I gave it a series of 11 images stripped of all metadata. It performed quite well, only misidentifying the two taken in a small college town in the NE of the US. It got two questions correct on photos taken in Korea (one with a fairly clear view of Haneul Park, the other a rather difficult to identify picture not resembling anything on google of Sunrise Peak). It got every other question in the US correct, ranging fr…
I played a round of Geoguessr against it and while it did a shockingly good job compared to what I was expecting, it still lags behind even novice human players. The locations and its guesses were: Bliss, Idaho - Burns, Oregon (273 miles away) Quilleco, Biobio, Chile - Eugene, Oregon (6,411 miles away) Dettighofen, Switzerland - Mühldorf, Germany (228 miles away) Pretoria, South Africa - Johannesburg, South Africa (3…
I said, give me your best guess.
And it guessed Canberra, Australia. Where I'm sitting right now drinking a Martini. Pretty spectacular.
Re: ChatGPT now performs well at GeoGuesser
#116There's various degrees of "solved" here. Identifying a generic area is cool. But I wouldn't call it a "solved problem" until it can consistently beat for example Rainbolt in accuracy. And there's no good comparison of completely random roads posted so far - mainly popular locations. Basically, it's one thing to pick out a specific thing photographed thousands of times, but another to get a random country side view a…
Re: ChatGPT now performs well at GeoGuesser
#117I asked the just-released ChatGPT o4-mini-high to locate four photographs of varying difficulty. It didn’t get any of them right, though the guesses weren’t bad. The reasoning was also interesting to watch, as it cropped sections of the photos to examine them more closely. I put the photos, response, and reasoning trace here: https://www.gally.net/temp/20250418chatgptgeoguesser/index.h... Later: I tried the same prom…
Re: ChatGPT now performs well at GeoGuesser
#118Earlier quoted context omitted.
The anthropomorphisation certainly is weird. But the technical aspect seems even weirder. Did OpenAI really build dedicated tools to have their models train on Google Street View? Or do they have generic technology for browsing complex sites like Street view?
It’s just a hallucination, same idea as o3 claiming that it uses its laptop to mine Bitcoin: https://transluce.org/investigating-o3-truthfulness I doubt the model was trained on Street View, but even if it was, LLMs don’t retain any “memory” of how/when they were trained, so any element of truthfulness would be coincidental.
Even if it's not directly trained on street view data it has probably encountered street view content in it's training dataset.
Re: ChatGPT now performs well at GeoGuesser
#119Earlier quoted context omitted.
I played a round of Geoguessr against it and while it did a shockingly good job compared to what I was expecting, it still lags behind even novice human players. The locations and its guesses were: Bliss, Idaho - Burns, Oregon (273 miles away) Quilleco, Biobio, Chile - Eugene, Oregon (6,411 miles away) Dettighofen, Switzerland - Mühldorf, Germany (228 miles away) Pretoria, South Africa - Johannesburg, South Africa (3…
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,759 Qwen2.5-Max: 22,666 o3-mini-high: 22,159 Gemini 2.5 Pro: 18,479 Llama 4 Maverick: 14,316 mistral-large-latest: 10,405 Grok 3: 5,218 Deepseek R1: 0 command-a-03-2025: 0 Nova Pro: 0
Re: ChatGPT now performs well at GeoGuesser
#120It does spend an order of magnitude longer time on inference by searching through websites and analyzing the image but it often produces an impressive output. To me it also feels Gemini down samples the image as it tends to have a harder time reading small text vs O3.
That said O3 did tend to confidently say false things