Live data from Hacker News

First Impressions with GPT-4V(ision)

blog.roboflow.com

11–20 of 345 posts

Re: First Impressions with GPT-4V(ision)

#11
Graph analysis is impressive (last example) - https://imgur.com/a/iOYTmt0

Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469

Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?...

Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac1/i_just_got...

It's Computer Vision on Steroids basically.

Multi-modality is pretty low hanging fruit so i'm glad we're finally getting started on that. Imagine if GPT-4 could manipulate sound and images even half as well as it could manipulate text. We still don't have a large scale multi-modal model trained from scratch so a lot of possible synergistic effects are still unknown.

Re: First Impressions with GPT-4V(ision)

#13
> The bounding box coordinates returned by GPT-4V did not match the position of the dog.

I suppose it just doesn't take image dimensions into consideration, and needs to be provided with max dimensions, or prompted to give percentages or other absolute values instead of pixels.

Re: First Impressions with GPT-4V(ision)

#16
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

Agree and the next big step may well be human computer interface. Speech is starting point for input. At some point output will change also and if think it out longer term perhaps a future where instead of reading information we install knowledge, including the stored memory of actual experience. If I want to do pottery, I could think this, download the experience and then be competent at it.

Re: First Impressions with GPT-4V(ision)

#18

Graph analysis is impressive (last example) - https://imgur.com/a/iOYTmt0 Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469 Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?... Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac…

Oh wow, I'm completely fucked as a front end developer.

Re: First Impressions with GPT-4V(ision)

#19
post #12
post #9

I'd be interested in knowing how good it is at solving visual captchas, do we foresee a huge rise in automated bypasses?

Solving CAPTCHAs at the moment is more inexpensive using humans than using GPT-4 API.

If true, this is wild.

I suppose a human could spend 10 seconds per Captcha, so they could do 360 per hour. Add some overhead for not being operating at peak performance every minute of every hour & call it 250. Let's say you can hire someone for $2, that works out to a bit over a penny per Captcha.

I don't think OpenAI has published pricing for GPT-4 Vision yet, but if we assume it's on par with GPT-4, and uses only 1000 of the 8000 possible tokens to process an image that's 3 cents per Captcha.

Doesn't seem completely unreasonable that at-scale humans may actually be cheaper than LLMs at this point. My mind is a little blown.

Re: First Impressions with GPT-4V(ision)

#20
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

Agree and the next big step may well be human computer interface. Speech is starting point for input. At some point output will change also and if think it out longer term perhaps a future where instead of reading information we install knowledge, including the stored memory of actual experience. If I want to do pottery, I could think this, download the experience and then be competent at it.

Even more impressive would be if I don't want to know pottery anymore, and I can delete that knowledge to make room for something else.
Post reply on HN