Earlier quoted context omitted.
Moreover you keep saying we can't scale infinitely. Sure...but nobody is saying we have to. 4 is not as scaled from 3 as 3 was from 2. Doesn't matter, still massive gap.
As I said already, the gap from 3 to 4 was substantially smaller than the gap between 2 to 3, and all indications are that the gap from 4 to 5 will also be further smaller than that.
We are beginning to roll out new voice and image capabilities in ChatGPT
761–770 of 914 posts
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#762Earlier quoted context omitted.
>The LLM concept can only take you so far, and we're approaching the limits of what an LLM is capable of. You don't know that. This is literally just an assertion. An unfounded one at that. If you couldn't predict how far in 2017 the LLM concept would take us today, then you definitely have no idea how far it could actually go. >believes we are approaching the limits of LLM size for size’s sake Nothing to do with thi…
Er, that's not how arguments work. What we can't know is that those trends will continue, so it's on you to demonstrate that they will, despite evidence suggesting they won't. As for as what you linked, Altman is saying the same thing I'm saying: > That doesn’t mean that OpenAI won't continue to try to make the models bigger, it just means they will likely double or triple in size each year rather than increasing by…
Gpt-4 can perform nearly all tasks you throw at it with well above average human performance. There literally isn't any testable definition of intelligence it fails that a big chunks of humans wouldn't also fail. You seem to keep missing the fact that We do not need an exponential improvement from 4.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#763Earlier quoted context omitted.
> how you can talk to something that doesn't provide factual information and just take it at face value Like talking to most people you mean?
Hopefully people know not to ask others for factual information (unless it's an area they're actually well educated/knowledgeable in), but for opinions and subjective viewpoints. "How's your day going", "How are you feeling", "What did you think of X", etc, not "So what was the deal with the Hundred Year's War?" or whatever. If people are treating LLMs like a random stranger and only making small talk, fair enough, b…
That's on them. I mean, people need to figure out that LLMs aren't random strangers, they're unfiltered inner voices of random strangers, spouting the first reaction they have to what you say to them.
Anyway, there is a middle ground. I like to ask GPT-4 questions within my area of expertise, because I'm able to instantly and instinctively - read: effortlessly - judge how much to trust any given reply. It's very useful this way, because rating an answer in your own field takes much less work than coming up with it on your own.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#764Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…
Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#765Earlier quoted context omitted.
The bike shown in the first image is Specialized Sirrus X. You can make out from the image of the manual that it says "spacer/axle/bolt specifications". Searching for this yields the following Specialized bike manual which is similar: https://www.manualslib.com/manual/1974494/Specialized-Epic-E... -- there are some notable differences, but the Specialized Sirrus X manuals that are online aren't in the same style. The…
Ah good find. yah, I tried bing and it is able to read a photo of that manual page and understand that the seat collar takes a 4mm hex wrench (though hallucinated and told me the torque was 5 Nm, unlike the correct 6.2, suggesting table reading is imperfect). Toolbox: I just found it too strong to claim you have the right tool, when it really doesn't know that. :) In the end it does feel like the image reader is just…
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#766The most important question for me: did it stop inventing facts?
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#767Earlier quoted context omitted.
> This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. Yep - it needs to be ready as soon as I'm done talking and I need to be able to interrupt it. If those things can be done then it can also start tentatively talking if I pause and immediately stop if I continue. I don't want to have to think about how to structure t…
The interruption is an important point yeah. It's so annoying when Siri misunderstands again and starts rattling off a whole host of options. And keeps getting stuck in a loop if you don't respond. In fact I'm really surprised these assistants are still as crap as they are. Totally scripted, zero AI. It seems low hanging fruit to implement an LLM but none of the big three have done so. Not even sure about the fringe…
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#768Earlier quoted context omitted.
Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…
Since this is getting a bit of interest, here's one more demo of this https://youtu.be/cvKUa5JpRp4 This demo shows even lower latency, plus the ability to handle very large menus with lots of complicated sub-options (this restaurant has over a billion option combinations to order a coffee). The latency is negative in some places, meaning the system finishes predicting before I finish speaking.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#769I like how they silently removed the web browsing (Bing browsing) chat feature after first having it disabled for several months. A proper notice about them removing the feature would've been nice. Maybe I missed it (someone please correct me if wrong), but the last I heard officially it was temporarily disabled while they fix something. Next thing I know, it's completely gone from the platform without another peep.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#770Earlier quoted context omitted.
Would you mind sharing some threads where you thought ChatGPT was useful? These discussions always feel like I’m living on a different planet with a different implementation of large language models than others who claim they’re great. The problems I run into seem to stem from the fundamental nature of this class of products.
The usefulness of ChatGPT is a bit situational, in my experience. But in the right situations it can be pretty powerful. Take a look at https://chat.openai.com/share/41bdb053-facd-448b-b446-1ba1f1... for example.
Context: had a bunch of photos and videos I wanted to share with a colleague, without uploading them to any cloud. I asked GPT-4 to write me a trivial single-page gallery that doesn't look like crap, feeding it the output of `ls -l` on the media directory, got it on first shot, copy-pasted and uploaded the whole bundle to a personal server - all in few minutes. It took maybe 15 minutes from the idea of doing it first occurring to me, to a private link I could share.
I have plenty more of those touching C++, Emacs Lisp, Python, generating vCARD and iCalendar files out of blobs of hastily-retyped or copy-pasted text, etc. The common thread here is: one-off, ad-hoc requests, usually underspecified. GPT-4 is quite good at being a fully generic tool for one-off jobs. This is something that never existed before, except in form of delegating a task to another human.