Earlier quoted context omitted.
I agree. I think apps that would initially benefit from LLM-powered conversational interfaces are those that have the following traits: - constrained context - part of a hands-free workflow A couple use-cases I have been pondering are driving assistant and cooking assistant. People are already used to using their phone or car's nav system to give them directions to an unfamiliar place. But even with such a system it'…
In the cooking example, you either need the AI to have full awareness of the step you are at or you need to describe the step you are at, which could be cumbersome ("I did ..., how much sugar do I need now"). I venture, having the recipe projected in front of you would be much faster.
First Impressions with GPT-4V(ision)
171–180 of 345 posts
Re: First Impressions with GPT-4V(ision)
#172Sure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe . AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your…
Roughly half of people in most developed countries are not functionally articulate: meaning, they can read functionally, but struggle to articulate what they want with the written word. LLM-based chatbots can be extremely attractive to the top 30% literacy users in the developed world. They are not a good universal UI. You still need to provide pathways for the user to follow to get done what they need without forcin…
Re: First Impressions with GPT-4V(ision)
#173Earlier quoted context omitted.
Also, keep in mind that Robots may be scifi level in 2 years. Like, able to cook and clean with hands! Crazy, but I think we will see it happen so fast: https://www.tri.global/news/toyota-research-institute-unveil...
I'm not sure if we watched the same video, but I saw a robot fumble though a few mechanical motions with dexterity and speed of a toddler to achieve a few isolated, best-case tasks where all the hard parts were taken care of by a human. Cool demo, I suppose, but nobody is going to buy this as anything other than a toy.
Re: First Impressions with GPT-4V(ision)
#174Earlier quoted context omitted.
Generally agree. Just to play devils advocate: If you want something done right, sometimes you have to do it yourself. Employees are sort of a universal UI. But you will always know more about what you want done than your agent, whether it’s human or computer. That’s even before considering the principal agent problem.
Just to play double-devils advocate: If you want something done right, other times you will have to get someone else to do it. You know what you want, but you might not have the skills to do it. I can't represent myself well in court, do a good job of plumbing or cut my own hair, so I would ask for experts to do that for me. Plus if someone is capable, it's often quicker to delegate than do, and if you are delegating…
Currently ChatGPT doesn't know it's bad at math, so it can convert a story problem into an equation better than a human but then mess up the arithmetic or forget a step in the straightforward part.
But if you specifically give ChatGPT access to Mathematica and an appropriate prompt, it can leverage a good math engine to get the right answer nearly every time.
Before long, I don't think that extra step will be necessary. It will know its limits and have dozens of other services that it can delegate to.
Re: First Impressions with GPT-4V(ision)
#175I’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not s…
> who takes a photo then wants to know it was a photo of? I find myself doing this rather frequently. The scenario described in the article is quite common for me: capturing a photo of a plant and utilizing an existing classification service to determine its identity. It could be driven by mere curiosity or practical concerns like identifying whether a plant is poison ivy. Wildlife identification also falls into this…
FYI Cornell Lab's Merlin app is fantastic at this, and its bird call audio identification is even better. They obviously have some top-notch machine learning going on there, and I'm really curious to see how both they and other services innovate on this front in the months to come.
Re: First Impressions with GPT-4V(ision)
#176It didn't successfully explain the NVIDIA burger joke though? The image is making fun of how nvidia has implemetned price discrimination by releasing consumer gpu's that don't have as much vram as they should so that they can sell the fully loaded datacenter gpu's at insane markup without pissing off gamers. The explanation didn't even come close to getting that.
Re: First Impressions with GPT-4V(ision)
#177Earlier quoted context omitted.
GPT-3 requires 700 gigabytes of GPU RAM. I'm looking at my cheapest computer components retailer listing a 48 gigabyte GPU at $5k. So to run the previous generation of GPT would cost me about $70k right now. When do you think I can expect to run GPT-4 on my consumer $device? :)
That doesn't seem to make sense. I can run Llama 2 on my 12-year-old desktop PC with no compatible GPU and only 16GB of system RAM. It ain't quick, but it runs.
Re: First Impressions with GPT-4V(ision)
#178Earlier quoted context omitted.
Until the AI mainframe runs on your $device
GPT-3 requires 700 gigabytes of GPU RAM. I'm looking at my cheapest computer components retailer listing a 48 gigabyte GPU at $5k. So to run the previous generation of GPT would cost me about $70k right now. When do you think I can expect to run GPT-4 on my consumer $device? :)
Second, I don't think it will be that long. There are already LLMs as good as GPT-3 running on average laptops and even phones.
In the next couple of years, you'll see:
- Ordinary PCs, tablets, and phones with dedicated AI chips, like TPUs - they'll be more tuned specifically for LLMs
- Mathematical and algorithmic optimizations will make existing LLMs faster on the same hardware
- Newer generations of LLMs will get even more useful with fewer parameters
The combination of all of these means that it's not at all unreasonable to expect that today's top-of-the-line LLM will be running locally on your device within just a couple of years.
Of course, LLMs in the cloud will advance even further, so there will always be a tradeoff, and there will always be demand for cloud AI, depending on the application.
Re: First Impressions with GPT-4V(ision)
#179Earlier quoted context omitted.
I think I wasn't clear enough -- these habits I'm talking about are things like "press cold water button, press start" or "press warm water button, press start" or "tap 'News' app grouping, tap 'NY Times' icon". There's nothing to infer. The sequence is already short. There are no benefits from AI here. But you raise a good point, which is that there are occasionally things like 15-step processes that people repeat a…
I don't know - the timer app on my oven is trivial too. But I always, always use Alexa to start timers. My hands are busy, so I can just ask "How many minutes left on the tea timer?" Voice is not really clumsy, compared to finding a device, browsing to an app, remembering the interface etc. Already when we meet a new app, we (I) often ask someone to show me around or tell me where the feature is that I want. Not any…
Everyone agrees setting timers in the kitchen via voice is great precisely because your hands are occupied. It's a special case. (And often used as the example of the only thing people end up consistently using their voice assistant for.)
And asking an AI where a feature is in an app -- that's exactly what I was describing. The app still has its UX though. But this is exactly the learning assistance I was describing.
And as for searching with Alexa, of course -- but that's just voice dictation instead of typing. Nothing to do with LLM's or interfaces.
Re: First Impressions with GPT-4V(ision)
#180Earlier quoted context omitted.
I don’t know, I hate the idea of having to hold a natural-language conversation with a computer in order to make use of its functionality. It feels like being one of those Futurama heads in a jar that can’t do anything by themselves.
UIs being dumbed down for average users was already annoying. Apparently the process won't stop until the illiterate are included too.