Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

751–760 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#751

Earlier quoted context omitted.

> just surrender to it all, believe in all the machine tells you unquestionably, forget the fact checking, it feels good to be ignorant... it will be fine... It's the same issue with Google Search, any web page, or, heck, any book. Fact checking gets you only so far. You need critical thinking. It's okay to "learn" wrong facts from time to time as long as you are willing to be critical and throw the ideas away if the…

The issue is far more serious with ChatGPT/similar models because things that are laughably untrue are delivered exactly the same as something that's solidly true. When doing a normal search I can make some assessment on the quality of the source and the likelihood the source is wrong. People should be able "throw the ideas away if they turn out to be wrong" but the problem is these ideas unconsciously or not help bu…

> Once you find out something isn't true it's hard to unpick your mental model of the world.

Intuitively, I would think the same, but a book about education research that I read and my own experience taught me that new information is surprisingly easy to unlearn. It’s probably because new information sits at the edges of your neural networks and do not yet provide a foundation for other knowledge. This will only happen if the knowledge stands the test of time (which is exactly how it should be according to Popper). If a counterexample is found, then the information can easily be discarded since it’s not foundational anyway and the brain learns the counterexample too (the brain is very good in remembering surprising things).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#752
Demos are underwhelming, but the potential is huge

Patiently awaiting rollout so I can chat about implementing UIs I like, and have GPT4 deliver a boilerplate with an implemented layout... Figma/XD plugins will probably arrive very soon too.

UX/UI Design is probably solved reached this point

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#753

Earlier quoted context omitted.

Definitely depends on the application, agreed. The more open ended the application the more dependent it is on larger LLMs (and other systems) that don't easily fit on edge. At the same time, progress is happening that is increasing the size of LLM that can be ran on edge. I imagine we end up in a hybrid world for many applications, where local models take a first pass (and also handle speech transcription) and only…

Can you share the source code? What did you do to improve the latency?

Lots of work around speculative decoding, optimizing across the ASR->LLM->TTS interfaces, fine-tuning smaller models while maintaining accuracy (lots of investment here), good old fashioned engineering around managing requests to the GPU, etc. We're considering commercializing this so I can't open source just yet, but if we end up not selling it I'll definitely think about opening it up.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#754

Earlier quoted context omitted.

GPT-4 is not the same product. I know it seems like it due to the way they position 3.5 and 4 on the same page, but they are really quite separate things. When I signed up for ChatGPT plus I didn't even bother using 3.5 because I knew it would be inferior. I still have only used it a handful of times. GPT-4 is just so much farther ahead that using 3.5 is just a waste of time.

Would you mind sharing some threads where you thought ChatGPT was useful? These discussions always feel like I’m living on a different planet with a different implementation of large language models than others who claim they’re great. The problems I run into seem to stem from the fundamental nature of this class of products.

I agree that none of the problems people have mentioned above happen with GPT4.

It used to be more reliable when web browsing worked, but it's still pretty reliable.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#755
post #153

Earlier quoted context omitted.

> ...talk to something that doesn't provide factual information and... Ah yes, I dont understand how to talk to people either!

I always thought a better future would be full of more and more distilled, accurate, useful knowledge and truthful people to promote that. Comments like yours make me think that no one cares about this...and judging by a lot of the other comments, I guess they don't. Probably going to be people, wading through a sea of AI generated shit, and the individual is supposed to just forever "apply critical thinking" to it a…

There aren't any real world sources of truth you can avoid applying critical thinking to. Much published research is false, and when it isn't, you need to know when it's expired or what context it's valid in.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#756

Earlier quoted context omitted.

It’s just a hiring article.

Hiring to produce more demos, to hire more to produce even more demos...

Yes. As long as the hirees do some actual work in between producing demos, this even makes sense as a hiring approach.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#757

Earlier quoted context omitted.

Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…

This is a very slick demo. Nice job!

Thanks! It's a lot of fun building with these new models and recent AI approaches.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#758

This is the dagger that will make online schooling unviable. ChatGPT already made it so that you could easily copy & paste any full-text questions and receive an answer with 90% accuracy. The only flaw was that problems that also used diagrams or figures would be out of the domain of ChatGPT. With image support, students could just take screenshots or document scans and have ChatGPT give them a valid answer. From wha…

It's true. I mean what is the point of doing schoolwork when some of the greatest minds of our time have decided the best way for the species to progress is to be replaced by machines? Imagine you're 16 years old right now, you know about ChatGPT, you know about OpenAI and their plans, and you're being told you need to study hard to get a good career..., but you're also reading up on what the future looks like accord…

I'm in my mid 30s and even I have some amount of apathy for the remainder of my career. I feel pretty confident my software and product experience is going to be not-so-useful in 15 years as it is today.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#759

This is the dagger that will make online schooling unviable. ChatGPT already made it so that you could easily copy & paste any full-text questions and receive an answer with 90% accuracy. The only flaw was that problems that also used diagrams or figures would be out of the domain of ChatGPT. With image support, students could just take screenshots or document scans and have ChatGPT give them a valid answer. From wha…

Use online for training, real life for testing/grading. That way cheating at home will only hurt yourself.

The problem here is that homework is designed to provide the structure kids need to apply themselves and actually learn. If you don't provide structure for this, they will simply never study and accept failure. They frequently don't have the self-discipline and mindfulness and long-term vision to study "because it's the right thing to do". I know my entire education, even with college, was "why do i need to know this?" and being wildly bored with it all as a result.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#760

Earlier quoted context omitted.

Voice assistants have always been a half complete product. They were shown off as a cool feature, then they were never integrated so they were useful. The two biggest features I want are for the voice assistants to read something for me, and to do something on google/Apple Maps hand free. Neither of these ever work. “Siri/ ok google add the next gas station on the route” or “take me to the Chinese restaurant in Hobok…

In the current world: Me: “OK Google, take me to the Chinese restaurant in Hoboken” Google Assistant: “Calling Jessica Hobkin”.

You forgot the third brand name.

The pattern for current world's voice assistants is: ${brand 1}, ${action} ${brand 2} ${joiner} ${brand 3}.

So, "OK Google, take me to Chinese restaurant in Hoboken using Google Maps".

Which is why I refuse to use this technology until the world gets its shit together.

Post reply on HN