> With that said, GPT-4V did make a mistake. The model said the fried chicken was labeled “NVIDIA BURGER” instead of “GPU”. Any midwesterner could tell you that CLEARLY it's a tenderloin :) https://www.seriouseats.com/best-breaded-pork-tenderloin-san...
First Impressions with GPT-4V(ision)
251–260 of 345 posts
Re: First Impressions with GPT-4V(ision)
#252Earlier quoted context omitted.
If the rate of improvement continues at the current pace - which is GPT 1 to 2 to 3 to 4 in the last five years - we are just one or two improvements away from a full blown AGI/superintelligence/singularity/etc. At that point, a superior user interface is probably the least interesting (or scary) thing that would happen. I personally doubt GPT-5 will be as much of an improvement over GPT-4 as GPT-4 was over GPT-3, bu…
chat-gpt at the end is a language model, not an real AI, it have limits and are huge
Re: First Impressions with GPT-4V(ision)
#253I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…
Re: First Impressions with GPT-4V(ision)
#254Earlier quoted context omitted.
No, it does not. It's still the same words, in a different medium. If you are unable to write, you'll probably be unable to speak your ideas.
This isn't true. There's plenty of people who are verbally fine but can't read or write. Spoken language is a far more common and fundamental skill than reading or writing.
Re: First Impressions with GPT-4V(ision)
#255Earlier quoted context omitted.
I don’t know, I hate the idea of having to hold a natural-language conversation with a computer in order to make use of its functionality. It feels like being one of those Futurama heads in a jar that can’t do anything by themselves.
> I hate the idea of having to hold a natural-language conversation with a computer in order to make use of its functionality. I hate the idea of having to use a mouse to click on a visual GUI to navigate a file system in order to make use of its functionality. It's less the case today, even among developers, but it wasn't that long ago that I remember that any serious technical user of a computer took it as a point…
Plus I think there's a nuance to what you're saying:
UX is not just about making the best channel surfing interface, which is essentially what phones/tablets are. We need UIs that are capable of rich interaction and expression of ideas, creation, etc.
Re: First Impressions with GPT-4V(ision)
#256Earlier quoted context omitted.
The more people say that, the less convincing it is There is no way I would have a UI developer onboarded when I can generate many iterations of layouts in midjourney, copy them into chatgpt4 and get code in NextJS with Typescript instantly non devs will have trouble doing this or thinking of the prompts to ask, but the dev team asking for headcount simply wont ask for headcount, and the engineering manager is going…
Yeah, I'm also skeptical about the actual value of specialists in the future. To me, AI brings a ton of power to generalists, who now have access to very powerful tools that would have taken them a long time to learn otherwise.
I would even go further and say the generalist gains a powerful tool belt that previously could not have existed. Not enough hours in the day or years in a lifetime.
Re: First Impressions with GPT-4V(ision)
#257Earlier quoted context omitted.
No, it does not. It's still the same words, in a different medium. If you are unable to write, you'll probably be unable to speak your ideas.
...do you know that illiterate people exist, right? Do you understand that people were illiterate for thousands of years and still managed to speak their ideas, right? Right?
Re: First Impressions with GPT-4V(ision)
#258Earlier quoted context omitted.
Roughly half of people in most developed countries are not functionally articulate: meaning, they can read functionally, but struggle to articulate what they want with the written word. LLM-based chatbots can be extremely attractive to the top 30% literacy users in the developed world. They are not a good universal UI. You still need to provide pathways for the user to follow to get done what they need without forcin…
> Roughly half of people in most developed countries are not functionally articulate Where did you get this idea? I found this article ( https://www.uxtigers.com/post/ai-articulation-barrier , is this you?), but it makes a leap from literacy to articulacy that I don't understand. It's not obvious to me why an illiterate person would be "functionally inarticulate" assuming they can speak instead of write. Also, I'm no…
I do however run a company that employs lots of blue collar, non-college-educated people, in manufacturing. And although this is in no way scientific, my experience matches this: most people are much more uncomfortable writing than they are reading. Even with reading, most strongly resist reading documentation unless they have to, and prefer trial and erroring their own gut instinct until they happen to find something that works or they give up. (This is less true of the most highly skilled technicians, such as those who troubleshoot robots and low voltage control systems.) The official statistics on literacy are absolutely not a good indicator of how comfortable people are articulating themselves with the written word, much less reading.
This is generally met with disbelief by most people in tech I talk with about this, because for the most part they have nearly zero interaction with this large portion of the population. From their daily experience, 98%+ people can make effective use of these tools.
But almost nobody in this partially literate population wants to write in an empty text box to ask an AI to do things. They can learn to visually navigate a simple UI, especially if it's well-designed, because they can effectively make decisions about what of several paths to take.
Some others here have brought up voice, and I do agree that voice is a more promising avenue, although I think it'll still take carefully constructed conversational experiences to work well (i.e., free form 'tell it what you want' will still not work).
Re: First Impressions with GPT-4V(ision)
#259Re: First Impressions with GPT-4V(ision)
#260Can somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?