Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

771–780 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#771

Earlier quoted context omitted.

Pay for the Plus version. Then it makes stuff up far less frequently. If the next version has the same step up in performance, I will no longer consider inaccuracy an issue - even the best books have mistakes in them, they just need to be infrequent enough.

> Pay for the Plus version. > Then it makes stuff up far less frequently. Now there's a business model for a ChatGPT-like service. $1/month: Almost always wrong $10/month: 50/50 chance of being right or wrong $100/month: right 95% of the time

You make it sound like business shenanigans, but the truth is, it's a natural fit for now, as performance of LLMs improves with their size, but costs of training (up-front investment) and inference (marginal, per-query) also go up.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#773

Earlier quoted context omitted.

OpenAI isn't marketing ChatGPT as, "infotainment."

now that you mention it, a big "for entertainment purposes only" banner like they use to have on all the psychic commercials on tv would not be inappropriate. it's incredible that LLMs are being marketed as general purpose assistants with a tiny asterisk, "may contain inaccuracies" like it's a walnut contamination

Not sure what's being incredible here. GPT-4 is a stellar general-purpose assistant, that shines when you stop treating it as encyclopedia, and start using it as an assistant. That is, give it tasks, like summarizing, or writing code, or explaining code, or rewriting prose. Ask for suggestions, ideas. You can do that to great effect, even when your requests are underspecified and somewhat confused, and it still works.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#774

Earlier quoted context omitted.

YMMV. For my case on software development, I don't even look on stackoverflow anymore. Just type the tech question, start refining into what is needed and get a snippet of code tailored for what is needed. What previously would take 30 to 60 minutes of research and testing is now less than a couple of minutes.

Fortunately it's not like StackOverflow has been used as training data for LLMs, right?

Well, yes. Point is, GPT-4 read the entire StackOverflow and then some, comprehended it, and now is a better interface to it, more specific and free of all the bullshit that's part of the regular web.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#775

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

I wonder when computers will start taking our intonation into account too. That would really help with understanding the end of a phrase. And there’s SO MUCH information in intonation that doesn’t exist in pure text. Any AI that doesn’t understand that part of language will always still be kinda dumb, however clever they are.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#776
post #473

Earlier quoted context omitted.

Yep. This example basically convinced me that they were unable to figure out anything actually useful to do with the model's new capabilities. Which makes me wonder how capable the new model in fact is.

Yah, pretty sure it is the same feature that's been in Bing Chat for 2 months now. Which feels really like there's only one pass of feature extraction from the image, preventing any detailed analysis beyond a course "what do you see". (Follow-up questions of things it likely didn't parse are highly hallucinated). This is why they can't extract the seat post information directly from the bike when the user asks. There…

>Yah, pretty sure it is the same feature that's been in Bing Chat for 2 months now.

It's not. Feel free to try these queries:

https://twitter.com/ComicSociety/status/1698694653845848544?... (comic book page in particular, from a be my eyes user)

Or these https://imgur.com/a/iOYTmt0 (graph analysis in particular, last example) and see Bing fail them.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#777

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

I wonder when computers will start taking our intonation into account too. That would really help with understanding the end of a phrase. And there’s SO MUCH information in intonation that doesn’t exist in pure text. Any AI that doesn’t understand that part of language will always still be kinda dumb, however clever they are.

Don’t they do it already? There are a lot of languages where intonation is absolutely necessary to distinguish between some words, so I would be surprised that this not already taken into account by the major voice assistants.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#778

Earlier quoted context omitted.

Ah good find. yah, I tried bing and it is able to read a photo of that manual page and understand that the seat collar takes a 4mm hex wrench (though hallucinated and told me the torque was 5 Nm, unlike the correct 6.2, suggesting table reading is imperfect). Toolbox: I just found it too strong to claim you have the right tool, when it really doesn't know that. :) In the end it does feel like the image reader is just…

Like a basic CLIP description: Tools, yellow toolbox, DEWALT, Allen wrenches, instruction manual. And then just using those keywords in the prompt. Yes, you’re right, it does feel like that.

A few of these wouldn't be possible with something like that. Look at the last picture, the graph analysis.

https://imgur.com/a/iOYTmt0

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#779

Earlier quoted context omitted.

Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…

Since this is getting a bit of interest, here's one more demo of this https://youtu.be/cvKUa5JpRp4 This demo shows even lower latency, plus the ability to handle very large menus with lots of complicated sub-options (this restaurant has over a billion option combinations to order a coffee). The latency is negative in some places, meaning the system finishes predicting before I finish speaking.

This is cool. But I want to see how it handles you going back one and tweaking it.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#780
post #258

Earlier quoted context omitted.

At what point to you go from startup to not when you have 10 billion invested and countless employees and is practically a sub branch of microsoft. Sounds cooler though I guess

I think you stop being a startup when there are engineers who do not know the CEO. I would guess OpenAI is still a startup by that definition (they don't have that many engineers IIRC) but I don't actually know.

That's really a function of what kind of CEO the company has, and what do you mean by "know". I worked remotely for a company employing hundreds of people, around for couple decades and with offices in different regions of the world, and I still got to talk to the CEO a couple times, and he knows me by name, all by virtue of bumping into him a couple times on corridor while on one of my infrequent visits to the office.
Post reply on HN