I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…
Watching o3 guess a photo's location is surreal, dystopian and entertaining
351–360 of 453 posts
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#352Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#353I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…
Did it mention it in its chain of thought? Otherwise, it could definitely output something because of X and then when asked why “rationalize” that it did it because Y
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#354Earlier quoted context omitted.
I've been surprised that so much focus was put on generative uses for LLMs and similar ML tools. It seems to me like they have a way better chance of being useful when tasked with interpreting given information rather than generating something meant to appear new.
Yeah, the "generative" in "generative AI" gives a little bit of a false impression. I like Laurie Voss's take on this: https://seldo.com/posts/what-ive-learned-about-writing-ai-ap... > Is what you're doing taking a large amount of text and asking the LLM to convert it into a smaller amount of text? Then it's probably going to be great at it. If you're asking it to convert into a roughly equal amount of text it will b…
I have been very pleased with responses to things like: "explain x", "summarize y", "make up a parody dog about A to the tune of B", "create a single page app that does abc".
The response is 1000x more text than the prompt.
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#355Feels a bit boring to use a picture of California. That's just a little too "in distribution" for a feature likely developed and tested there.
I agree, that's why I followed up with a photo from Madagascar.
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#356Earlier quoted context omitted.
Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.
This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#357Earlier quoted context omitted.
Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.
This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#358Earlier quoted context omitted.
Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.
This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…
Relatively speaking we live in a bubble, there are still broad swaths of the economy that operate with pen and paper. Another broad swath that migrated off 1980s era AS/400 in the last few years. Even if we had ASI available literally today (And we don’t) I’d give it 20-30 years until the guy that operates your corner market or the local auto repair shop has any use in the world for it.
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#359Earlier quoted context omitted.
This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…
I don’t know if most software engineers build toy CRUD apps all day? I have found the state of the art models to be almost completely useless in a real large codebase. Tried Claude and Gemini latest since the company provides them but they couldn’t even write tests that pass after over a day of trying
Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining
#360I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…
> when I asked it how, it mentioned that it knows I live nearby. > The process for how it arrives at the conclusion is somewhat similar to humans. It looks at vegetation, terrain, architecture, road infrastructure, signage, and it just knows seemingly everything about all of them. Can we trust what the model says when we ask it about how it comes up with an answer?