Live data from Hacker News

First Impressions with GPT-4V(ision)

blog.roboflow.com

191–200 of 345 posts

Re: First Impressions with GPT-4V(ision)

#191

Earlier quoted context omitted.

If something is a repetitive habit that you can do almost without thinking, there is a good chance an AI could infer that entire chain. I think what's more likely is that an AI based interface will end up being superior after it has had a chance to observe your personal preferences and approach on a conventional UI. So both will still be needed, with an AI helping at the low end and high end of experience and the mid…

I think I wasn't clear enough -- these habits I'm talking about are things like "press cold water button, press start" or "press warm water button, press start" or "tap 'News' app grouping, tap 'NY Times' icon". There's nothing to infer. The sequence is already short. There are no benefits from AI here. But you raise a good point, which is that there are occasionally things like 15-step processes that people repeat a…

I totally get your point, but I think that AI will allow much "smarter" behavior. Where every appliance is an expert in doing what it is intended to do.

So sure, it will still have buttons, but those buttons are really just preset AI prompts on the backend. You can also just talk to your appliance and nuance your request however you want to.

A TV with a remote whose channel button just prompts "Next channel" but if you want you would just talk to your TV and say "Skip 10 channels" or "make the channel button do (arbitrary behavior)"

The shortcuts will definitely stay, but they will behave closer to "ring bell for service" than "press selection to vend".

Re: First Impressions with GPT-4V(ision)

#192
post #65

Earlier quoted context omitted.

So you won't be able to do anything without Internet connection to the AI mainframe? No thanks.

Until the AI mainframe runs on your $device

By the time the current AI mainframe runs on your device, there will be new, better models that still require the mainframe.

I think AI fundamentally favors centralization. Except for narrow tasks and domains, there's no such thing as "enough" intelligence. For general purpose AI, you'll always want the best and most intelligent model available, which means cloud rather than local.

Re: First Impressions with GPT-4V(ision)

#193
post #10

Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…

I share your awe. I feel like a kid in a candy store with all these incredible AI breakthroughs coming out these days! There's a place for cynicism and pessimism, but the kid in me who loves technology for its own sake is just absolutely on cloud 9.

Re: First Impressions with GPT-4V(ision)

#195
post #181

Earlier quoted context omitted.

People keep saying "ah but it can't do X!". So what? Most of us have multiple decades until we can retire. This AI is getting better every few months. It will be able to do it even faster, better and more cheaply than a human can.

> It will be able to do it even faster, better and more cheaply than a human can. Take what you did in the past year. Write down every product decision taken, every interaction with other teams figuring out APIs you had, all the infra where your code is running and how it was setup and changed, all the design iterations and changes that had to be implemented (especially if you have external partners demanding it). Ye…

So ... in this example, your job is continually feeding information to the AI from various sources. Why would the AI not be automatically hooked up to all those sources? Building a system that can do that is essentially trivial with the OpenAI API.

Re: First Impressions with GPT-4V(ision)

#196
post #194

Could someone with access tell me what GPT-4V has to say of this image? http://karpathy.github.io/assets/obamafunny.jpg Andrej Karpathy used it in 2012 as an example of an image he thought would be extremely hard for a model to interpret. I'm wondering how this holds 11 years later...

[deleted]

Re: First Impressions with GPT-4V(ision)

#197
post #181

Earlier quoted context omitted.

People keep saying "ah but it can't do X!". So what? Most of us have multiple decades until we can retire. This AI is getting better every few months. It will be able to do it even faster, better and more cheaply than a human can.

> It will be able to do it even faster, better and more cheaply than a human can. Take what you did in the past year. Write down every product decision taken, every interaction with other teams figuring out APIs you had, all the infra where your code is running and how it was setup and changed, all the design iterations and changes that had to be implemented (especially if you have external partners demanding it). Ye…

We'll have jobs, but they sure as shit won't be worth $150k anymore.

Any grunt can feed meeting notes into an AI. And frankly, and AI can parse an audio recording on a meeting.

Re: First Impressions with GPT-4V(ision)

#198
post #47

The discrepancy between the two answers regarding the set of coins is jarring. From the answer to the first question, one would assume that it can’t tell the currency. The answer to the second question shows that it actually can. The fact that LLMs don’t reflect a consistent inner model in that way, and hence the users’ inability to adequately reason about their AI interlocutor, is currently a severe usability issue.

I've heard that it is because the AI outputs what it is thinking as it is thinking it. It doesn't really reflect, it sort of does the equivalent of just verbal thought streaming right onto the screen.

So when you ask it to reflect on what it said, that's when it actually looks at it and reflects on it.

Re: First Impressions with GPT-4V(ision)

#199

Earlier quoted context omitted.

Roughly half of people in most developed countries are not functionally articulate: meaning, they can read functionally, but struggle to articulate what they want with the written word. LLM-based chatbots can be extremely attractive to the top 30% literacy users in the developed world. They are not a good universal UI. You still need to provide pathways for the user to follow to get done what they need without forcin…

Audio to text solves written word articulation, right? Besides this post is about vision, which also solves it.

No, it does not. It's still the same words, in a different medium. If you are unable to write, you'll probably be unable to speak your ideas.

Re: First Impressions with GPT-4V(ision)

#200

Earlier quoted context omitted.

If something is a repetitive habit that you can do almost without thinking, there is a good chance an AI could infer that entire chain. I think what's more likely is that an AI based interface will end up being superior after it has had a chance to observe your personal preferences and approach on a conventional UI. So both will still be needed, with an AI helping at the low end and high end of experience and the mid…

I think I wasn't clear enough -- these habits I'm talking about are things like "press cold water button, press start" or "press warm water button, press start" or "tap 'News' app grouping, tap 'NY Times' icon". There's nothing to infer. The sequence is already short. There are no benefits from AI here. But you raise a good point, which is that there are occasionally things like 15-step processes that people repeat a…

Most user interfaces already have a much finer granularity and number of options than your examples.

When taking a shower, I would like fine control over the water temperature, preferably with a feedback loop regulating the temperature. (Preferably also the regulation changes over the duration of the showering.)

Choosing to read the NY times indeed is only a few taps away, but navigating through and within its list of articles is nowadays done quite fast and intuitively thanks to quite a lot of UI advancements.

My point being, short sequences are a very limited set within a vast UI space.

People go for convenience and speed, oftentimes even if there's some accuracy cost. AI fulfills this preference, especially because it can learn on the go.

Post reply on HN