First Impressions with GPT-4V(ision)
81–90 of 345 posts
Re: First Impressions with GPT-4V(ision)
#82Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…
I don’t know, I hate the idea of having to hold a natural-language conversation with a computer in order to make use of its functionality. It feels like being one of those Futurama heads in a jar that can’t do anything by themselves.
For now almost all applications of ChatGPT happen in chat windows because it requires no further integration, but there's no reason to expect things will always be this way.
Re: First Impressions with GPT-4V(ision)
#83Earlier quoted context omitted.
> Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? OpenAI’s example included bike repair and toolkit choice Allot of people could use this even if they aren’t right now
Don’t be ridiculous. They’ll use YouTube, just like they do right now. Maybe if it could watch the video, then step you through it step by step. …but it cant , with what they’ve actually released here. Oh whatever. If I’m wrong, I’m wrong. Time will tell.
and ad block doesn't work on mobile
if you have a case that wasn't covered by that video? you have to go to another or continue searching all while wishing you could just talk to someone about it. if you don't know the word for what you're looking for, all the search engines lack utility.
ChatGPT4 with image recognition and conversation solves all of that use case and people already use it, so now they'll just start sending it pictures from the phone already in their hand that they're already using to chat with
there are plenty of times over the last year that would have been useful for me. plenty of times over the last year I just didn't continue being interested in that problem
it just seems kind of…. late ?… for that “dont be ridiculous” reaction. classic dropbox moment
Re: First Impressions with GPT-4V(ision)
#84Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…
I agree. I think apps that would initially benefit from LLM-powered conversational interfaces are those that have the following traits: - constrained context - part of a hands-free workflow A couple use-cases I have been pondering are driving assistant and cooking assistant. People are already used to using their phone or car's nav system to give them directions to an unfamiliar place. But even with such a system it'…
Re: First Impressions with GPT-4V(ision)
#85The "Why is this image funny?" test reminds me of https://karpathy.github.io/2012/10/22/state-of-computer-visi... In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"
Re: First Impressions with GPT-4V(ision)
#86So a jumble of chair legs is “NVIDIA burger” and it did say GPU was a “bun” so it thinks the flat thing (chicken?) is some sort of bread. If GPT-4V was “aware”, it would say “it’s funny because I won’t get it right but you will use it get a bunch of $VC, and that is funny, kinda”.
Re: First Impressions with GPT-4V(ision)
#87I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…
This is likely how we'll communicate with information systems: throw some hand-wavy question at it, and refine your query based on its output using natural language until you find the answer (or even the question) you were looking for.
Re: First Impressions with GPT-4V(ision)
#88Graph analysis is impressive (last example) - https://imgur.com/a/iOYTmt0 Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469 Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?... Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac…
Oh wow, I'm completely fucked as a front end developer.
However, keep in mind that these are cherry-picked. If someone just took that output and stuck onto a website, it'd be a pretty horrible website. There's always going to be someone who manages the code and actually interacts with the AI, so there will still be some jobs.
And your boss isn't going to be doing any coding. I'm pretty sure that role is still loaded and they'll still be managing people rather than coding, and maybe sometimes engaging with an AI.
Another prediction: I'm pretty sure specialists are going to be significantly more important as your job will be to identify the AI's deficiencies and improve on it.
Re: First Impressions with GPT-4V(ision)
#89Earlier quoted context omitted.
So you won't be able to do anything without Internet connection to the AI mainframe? No thanks.
Until the AI mainframe runs on your $device
Re: First Impressions with GPT-4V(ision)
#90Earlier quoted context omitted.
I wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.
Yep, some CS/AI grads from Stanford trained an AI on loads of Street View images and built a bot that is able to beat some of the best Geoguessr players: https://www.youtube.com/watch?v=ts5lPDV--cU