I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…
First Impressions with GPT-4V(ision)
261–270 of 345 posts
Re: First Impressions with GPT-4V(ision)
#262The "Why is this image funny?" test reminds me of https://karpathy.github.io/2012/10/22/state-of-computer-visi... In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"
Has anyone tried GPT-4V on that image?
"The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly unaware or unbothered by the presence of the others. The contrast between the formal setting and the informal action makes the image amusing. Additionally, the reactions of the other individuals, particularly the man looking at the scale, add to the comedic element."
It got the humor but didn't identify Obama.
Re: First Impressions with GPT-4V(ision)
#263Can somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?
How what works? Could you elaborate?
Or is my understanding completely off? Perhaps it's "Translating" the image to text, by outputting a sequence of text tokens as it scans the image regions, and then the text queries (e.g. "whats funny about this") uses this translation as the context? Presumably, this is how the model handles audio input.
Re: First Impressions with GPT-4V(ision)
#264Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…
> they will become a superior user interface to almost every thing you want to do No they won't. They're actually a pretty terrible user interface from a design perspective. Primarily because they provide zero affordances, but also because of speed. UX is about providing an intuitive understanding of available capabilities at a glance, and allowing you to do things with a single tap that then reflect the new state ba…
Re: First Impressions with GPT-4V(ision)
#265Can somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?
Re: First Impressions with GPT-4V(ision)
#266Earlier quoted context omitted.
But what is bad about that? Why shouldn't they be included?
It's a race to the bottom. You build an idiot-proof UI, Mother Nature builds a better idiot.
Re: First Impressions with GPT-4V(ision)
#267Earlier quoted context omitted.
Has anyone tried GPT-4V on that image?
Another response for this image a friend sent: " The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly…
Re: First Impressions with GPT-4V(ision)
#268Earlier quoted context omitted.
> they will become a superior user interface to almost every thing you want to do No they won't. They're actually a pretty terrible user interface from a design perspective. Primarily because they provide zero affordances, but also because of speed. UX is about providing an intuitive understanding of available capabilities at a glance, and allowing you to do things with a single tap that then reflect the new state ba…
There’s no reason chatbots have to be the interface to an LLM. Imagine dynamically generated interfaces redesigning themselves to your needs as you work through a task.
Re: First Impressions with GPT-4V(ision)
#269Earlier quoted context omitted.
This isn't true. There's plenty of people who are verbally fine but can't read or write. Spoken language is a far more common and fundamental skill than reading or writing.
Plus the LLM could adapt its language and dialect to say Appalachia or Compton, etc.
Re: First Impressions with GPT-4V(ision)
#270Earlier quoted context omitted.
Another response for this image a friend sent: " The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly…
It doesn't get the humor. The whole "foot on the scale" aspect and how it increases weight.