Live data from Hacker News

First Impressions with GPT-4V(ision)

blog.roboflow.com

261–270 of 345 posts

Re: First Impressions with GPT-4V(ision)

#261

I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…

Image understanding -> scene understanding -> world understanding -> intelligence

Re: First Impressions with GPT-4V(ision)

#262
post #37

The "Why is this image funny?" test reminds me of https://karpathy.github.io/2012/10/22/state-of-computer-visi... In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"

Has anyone tried GPT-4V on that image?

Another response for this image a friend sent:

"The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly unaware or unbothered by the presence of the others. The contrast between the formal setting and the informal action makes the image amusing. Additionally, the reactions of the other individuals, particularly the man looking at the scale, add to the comedic element."

It got the humor but didn't identify Obama.

Re: First Impressions with GPT-4V(ision)

#263

Can somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?

How what works? Could you elaborate?

As far as I understand, these multi-modal models work by embedding the text/image in a shared representation space. To perform OCR on such an embedding, it would require extracting every letter, in the correct order, from the embedding. But given the embedding is a fixed size, and therefor necessarily compressed, I would expect it to loose the exactness of the underlying input, especially with images containing a lot of text. So assuming GPT-V can effectively perform OCR, how is this being done given the constraints?

Or is my understanding completely off? Perhaps it's "Translating" the image to text, by outputting a sequence of text tokens as it scans the image regions, and then the text queries (e.g. "whats funny about this") uses this translation as the context? Presumably, this is how the model handles audio input.

Re: First Impressions with GPT-4V(ision)

#264
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

> they will become a superior user interface to almost every thing you want to do No they won't. They're actually a pretty terrible user interface from a design perspective. Primarily because they provide zero affordances, but also because of speed. UX is about providing an intuitive understanding of available capabilities at a glance, and allowing you to do things with a single tap that then reflect the new state ba…

There’s no reason chatbots have to be the interface to an LLM. Imagine dynamically generated interfaces redesigning themselves to your needs as you work through a task.

Re: First Impressions with GPT-4V(ision)

#265

Can somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?

I wouldn’t be surprised if they do an actual OCR pass for every input image and just pass in the raw text as a part of the prompt. That plus the embedding should work well.

Re: First Impressions with GPT-4V(ision)

#266

Earlier quoted context omitted.

But what is bad about that? Why shouldn't they be included?

It's a race to the bottom. You build an idiot-proof UI, Mother Nature builds a better idiot.

This agrees with my experience. When I have begun to automate complex manual systems (nibbling from each end, typically), I note when I watch people use the finished product (at each step) that they simply find some other facet of the job to pay less attention to. The eventual error rate just returns to what it was before.

Re: First Impressions with GPT-4V(ision)

#267
post #262

Earlier quoted context omitted.

Has anyone tried GPT-4V on that image?

Another response for this image a friend sent: " The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly…

It doesn't get the humor. The whole "foot on the scale" aspect and how it increases weight.

Re: First Impressions with GPT-4V(ision)

#268
post #264

Earlier quoted context omitted.

> they will become a superior user interface to almost every thing you want to do No they won't. They're actually a pretty terrible user interface from a design perspective. Primarily because they provide zero affordances, but also because of speed. UX is about providing an intuitive understanding of available capabilities at a glance, and allowing you to do things with a single tap that then reflect the new state ba…

There’s no reason chatbots have to be the interface to an LLM. Imagine dynamically generated interfaces redesigning themselves to your needs as you work through a task.

Exactly. Chat is just the first, most general purpose UI. Like a terminal was for computers.

Re: First Impressions with GPT-4V(ision)

#269
post #210

Earlier quoted context omitted.

This isn't true. There's plenty of people who are verbally fine but can't read or write. Spoken language is a far more common and fundamental skill than reading or writing.

Plus the LLM could adapt its language and dialect to say Appalachia or Compton, etc.

Ha. Can you imagine an AI speaking in colloquial Black American or Appalachian dialect? People's minds would short circuit, not knowing whether to be offended or approving.

Re: First Impressions with GPT-4V(ision)

#270
post #262

Earlier quoted context omitted.

Another response for this image a friend sent: " The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly…

It doesn't get the humor. The whole "foot on the scale" aspect and how it increases weight.

lol even the HN user didn’t get the humor
Post reply on HN