Live data from Hacker News

First Impressions with GPT-4V(ision)

blog.roboflow.com

81–90 of 345 posts

Re: First Impressions with GPT-4V(ision)

#81
So a jumble of chair legs is “NVIDIA burger” and it did say GPU was a “bun” so it thinks the flat thing (chicken?) is some sort of bread. If GPT-4V was “aware”, it would say “it’s funny because I won’t get it right but you will use it get a bunch of $VC, and that is funny, kinda”.

Re: First Impressions with GPT-4V(ision)

#82
post #28
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

I don’t know, I hate the idea of having to hold a natural-language conversation with a computer in order to make use of its functionality. It feels like being one of those Futurama heads in a jar that can’t do anything by themselves.

There's nothing stopping developers from taking a prompt to GPT and sticking it behind a button or command line, with options in the UI interpolated into the prompt.

For now almost all applications of ChatGPT happen in chat windows because it requires no further integration, but there's no reason to expect things will always be this way.

Re: First Impressions with GPT-4V(ision)

#83

Earlier quoted context omitted.

> Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? OpenAI’s example included bike repair and toolkit choice Allot of people could use this even if they aren’t right now

Don’t be ridiculous. They’ll use YouTube, just like they do right now. Maybe if it could watch the video, then step you through it step by step. …but it cant , with what they’ve actually released here. Oh whatever. If I’m wrong, I’m wrong. Time will tell.

the best case scenario is a 30 second youtube video with an ad that lasts 15 seconds followed by a 2 minute ad that I can skip in 5 more seconds

and ad block doesn't work on mobile

if you have a case that wasn't covered by that video? you have to go to another or continue searching all while wishing you could just talk to someone about it. if you don't know the word for what you're looking for, all the search engines lack utility.

ChatGPT4 with image recognition and conversation solves all of that use case and people already use it, so now they'll just start sending it pictures from the phone already in their hand that they're already using to chat with

there are plenty of times over the last year that would have been useful for me. plenty of times over the last year I just didn't continue being interested in that problem

it just seems kind of…. late ?… for that “dont be ridiculous” reaction. classic dropbox moment

Re: First Impressions with GPT-4V(ision)

#84
post #71
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

I agree. I think apps that would initially benefit from LLM-powered conversational interfaces are those that have the following traits: - constrained context - part of a hands-free workflow A couple use-cases I have been pondering are driving assistant and cooking assistant. People are already used to using their phone or car's nav system to give them directions to an unfamiliar place. But even with such a system it'…

In the cooking example, you either need the AI to have full awareness of the step you are at or you need to describe the step you are at, which could be cumbersome ("I did ..., how much sugar do I need now"). I venture, having the recipe projected in front of you would be much faster.

Re: First Impressions with GPT-4V(ision)

#85
post #37

The "Why is this image funny?" test reminds me of https://karpathy.github.io/2012/10/22/state-of-computer-visi... In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"

Has anyone tried GPT-4V on that image?

Re: First Impressions with GPT-4V(ision)

#86

So a jumble of chair legs is “NVIDIA burger” and it did say GPU was a “bun” so it thinks the flat thing (chicken?) is some sort of bread. If GPT-4V was “aware”, it would say “it’s funny because I won’t get it right but you will use it get a bunch of $VC, and that is funny, kinda”.

[deleted]

Re: First Impressions with GPT-4V(ision)

#87

I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…

This is mostly useless. Essentially a toy. I am not that much hyped by AI tools either, but come on. This is clearly the future of human-computer interaction.

This is likely how we'll communicate with information systems: throw some hand-wavy question at it, and refine your query based on its output using natural language until you find the answer (or even the question) you were looking for.

Re: First Impressions with GPT-4V(ision)

#88
post #18

Graph analysis is impressive (last example) - https://imgur.com/a/iOYTmt0 Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469 Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?... Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac…

Oh wow, I'm completely fucked as a front end developer.

If AI continues to get better it won't just be you who's in trouble.

However, keep in mind that these are cherry-picked. If someone just took that output and stuck onto a website, it'd be a pretty horrible website. There's always going to be someone who manages the code and actually interacts with the AI, so there will still be some jobs.

And your boss isn't going to be doing any coding. I'm pretty sure that role is still loaded and they'll still be managing people rather than coding, and maybe sometimes engaging with an AI.

Another prediction: I'm pretty sure specialists are going to be significantly more important as your job will be to identify the AI's deficiencies and improve on it.

Re: First Impressions with GPT-4V(ision)

#89
post #65

Earlier quoted context omitted.

So you won't be able to do anything without Internet connection to the AI mainframe? No thanks.

Until the AI mainframe runs on your $device

GPT-3 requires 700 gigabytes of GPU RAM. I'm looking at my cheapest computer components retailer listing a 48 gigabyte GPU at $5k. So to run the previous generation of GPT would cost me about $70k right now. When do you think I can expect to run GPT-4 on my consumer $device? :)

Re: First Impressions with GPT-4V(ision)

#90
post #6
post #3

Earlier quoted context omitted.

I wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.

Yep, some CS/AI grads from Stanford trained an AI on loads of Street View images and built a bot that is able to beat some of the best Geoguessr players: https://www.youtube.com/watch?v=ts5lPDV--cU

IIRC it wasn't that impressive in the end as instead of recognizing the places the AI apparently learnt to recognize subtle differences in street view cameras used in different locations? I might be wrong / thinking of the wrong model l and I'm on mobile without my browsing history so hard to check, but I think it was putting a lot of weight on some pixels that are noisy
Post reply on HN