Live data from Hacker News

First Impressions with GPT-4V(ision)

blog.roboflow.com

1–10 of 345 posts

Re: First Impressions with GPT-4V(ision)

#2
I’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with.

Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it doesn’t do as well.

I tried having it play Geoguessr via screenshots & it wasn’t good at it.

Re: First Impressions with GPT-4V(ision)

#3
post #2

I’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with. Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it…

I wonder how many images from Street View it has been trained on.

I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.

Re: First Impressions with GPT-4V(ision)

#4
post #3
post #2

I’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with. Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it…

I wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.

It's been done recently! It's a bit better than (but competitive with) top players.

https://www.youtube.com/watch?v=ts5lPDV--cU

Re: First Impressions with GPT-4V(ision)

#5
post #3
post #2

I’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with. Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it…

I wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.

> I would assume training an LLM to do the same would definitely be doable.

I wouldn't be so sure. The reasoning process of Geoguessr pros is symbolic, not statistical inference.

/edit: as other commenters pointed out, something similar was done. While this wasn't an LLM, it was a deep learning model, so not symbolic -> https://www.theregister.com/2023/07/15/pigeon_model_geolocat...

Re: First Impressions with GPT-4V(ision)

#6
post #3
post #2

I’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with. Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it…

I wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.

Yep, some CS/AI grads from Stanford trained an AI on loads of Street View images and built a bot that is able to beat some of the best Geoguessr players: https://www.youtube.com/watch?v=ts5lPDV--cU

Re: First Impressions with GPT-4V(ision)

#10
Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe. AWE.

Let me state the obvious, in case anyone here isn't clear about the implications:

If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your dishwasher, your home, your office, etc.

UIs to many apps, services, and devices -- and many apps themselves -- will be replaced by an AI that does what you want when you want it.

A lot of people don't want this to happen -- it is kind of scary -- but to me it looks inevitable.

Also inevitable in my view is that eventually we'll give these AI models robotic bodies (think: "computer, make me my favorite breakfast").

We live in interesting times.

--

EDITS: Changed "every single thing" to "almost every thing," and elaborated on the original comment to convey my thoughts more accurately.

Post reply on HN