Live data from Hacker News

Watching o3 guess a photo's location is surreal, dystopian and entertaining

simonwillison.net

351–360 of 453 posts

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#351
post #63

I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…

Have you gleaned anything watching o3 make decisions on a photo? ( i.e. have you noticed if it has thought of anything you.. and other higher level players similar to you... have not? )

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#353
post #63

I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…

> when I asked it how, it mentioned that it knows I live nearby

Did it mention it in its chain of thought? Otherwise, it could definitely output something because of X and then when asked why “rationalize” that it did it because Y

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#354
post #268

Earlier quoted context omitted.

I've been surprised that so much focus was put on generative uses for LLMs and similar ML tools. It seems to me like they have a way better chance of being useful when tasked with interpreting given information rather than generating something meant to appear new.

Yeah, the "generative" in "generative AI" gives a little bit of a false impression. I like Laurie Voss's take on this: https://seldo.com/posts/what-ive-learned-about-writing-ai-ap... > Is what you're doing taking a large amount of text and asking the LLM to convert it into a smaller amount of text? Then it's probably going to be great at it. If you're asking it to convert into a roughly equal amount of text it will b…

This quote sounds clever, but is very different than my experience.

I have been very pleased with responses to things like: "explain x", "summarize y", "make up a parody dog about A to the tune of B", "create a single page app that does abc".

The response is 1000x more text than the prompt.

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#355
post #309

Feels a bit boring to use a picture of California. That's just a little too "in distribution" for a feature likely developed and tested there.

I agree, that's why I followed up with a photo from Madagascar.

Ah, must've missed that. Either way I appreciate your frequent posts testing new AI models.

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#356

Earlier quoted context omitted.

Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.

This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…

I wonder what the impact will be when replicating the same thing becomes machine readable with near 100% accuracy.

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#357

Earlier quoted context omitted.

Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.

This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…

I don’t know if most software engineers build toy CRUD apps all day? I have found the state of the art models to be almost completely useless in a real large codebase. Tried Claude and Gemini latest since the company provides them but they couldn’t even write tests that pass after over a day of trying

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#358

Earlier quoted context omitted.

Which also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.

This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…

I agree but the reason it won’t be an apocalypse is the same reason economists get most things wrong, it’s not an efficient market.

Relatively speaking we live in a bubble, there are still broad swaths of the economy that operate with pen and paper. Another broad swath that migrated off 1980s era AS/400 in the last few years. Even if we had ASI available literally today (And we don’t) I’d give it 20-30 years until the guy that operates your corner market or the local auto repair shop has any use in the world for it.

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#359

Earlier quoted context omitted.

This is also why I buy the apocalyptic headlines about AI replacing white collar labor - most white collar employment is mostly creating the same things (a CRUD app, a landing page, a business plan) with a few custom changes Not a lot of labor is actually engaged in creating novel things. The marketing plan for your small business is going to be the same as the marketing plan for every other small business with some…

I don’t know if most software engineers build toy CRUD apps all day? I have found the state of the art models to be almost completely useless in a real large codebase. Tried Claude and Gemini latest since the company provides them but they couldn’t even write tests that pass after over a day of trying

Same. Like Claude code for example will write some tests. But what they are testing is often incorrect

Re: Watching o3 guess a photo's location is surreal, dystopian and entertaining

#360
post #63

I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby. However, I've given it vacation photos from ages ago, and not only in tourist destinations eit…

> when I asked it how, it mentioned that it knows I live nearby. > The process for how it arrives at the conclusion is somewhat similar to humans. It looks at vegetation, terrain, architecture, road infrastructure, signage, and it just knows seemingly everything about all of them. Can we trust what the model says when we ask it about how it comes up with an answer?

You're just asking the left-brain interpreter [1] its opinion about what the right-brain did.

[1] https://en.wikipedia.org/wiki/Left_brain_interpreter

Post reply on HN