Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

831–840 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#832

I've been making a few hobby projects that consolidate different AI services to achieve this, so I look forward to the reduced complexity and latency from all those trips. If the API is available in time (halloween), my multi-modal talking skeleton head with an ESP32 camera that makes snarky comments about your costume just got slightly easier on the software side.

If you make this, please share some steps/details! It sounds super cool and I'd love to make something like this!

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#833

I've been making a few hobby projects that consolidate different AI services to achieve this, so I look forward to the reduced complexity and latency from all those trips. If the API is available in time (halloween), my multi-modal talking skeleton head with an ESP32 camera that makes snarky comments about your costume just got slightly easier on the software side.

> I've been making a few hobby projects that consolidate different AI services to achieve this, so I look forward to the reduced complexity and latency from all those trips.

ironically this is basically the exact line of reasoning for why i didn't embark on any such endeavors

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#834

I like how they silently removed the web browsing (Bing browsing) chat feature after first having it disabled for several months. A proper notice about them removing the feature would've been nice. Maybe I missed it (someone please correct me if wrong), but the last I heard officially it was temporarily disabled while they fix something. Next thing I know, it's completely gone from the platform without another peep.

Just made an account to say that I currently have this feature. It was gone for a few months but it came back to me I think this past week. Not as a plugin, either, it is its own “model” to select.

Since so many others including myself don't see it, I guess that means it is getting a slow rollout which they are being extra cautious with this time.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#835

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

Also curious to hear about your setup. Using whisper too? When I was experimenting with it there was still a lot of annoyance about hallucinations and I was hard coding some "if last phrase is 'thanks for watching', ignore last phrase" I was just googling a bit to see what's out there now for whisper/llama combos and came across this: https://github.com/yacineMTB/talk There's a demo linked on the github page that see…

Turn the volume on your microphone down and watch as Whisper just starts SCREAMING.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#836
post #147

Earlier quoted context omitted.

Another option is that this doesn't replace the student's work, but the teacher's. The single greatest use I have found for ChatGPT is in educating myself on various topics, hosting a socratic seminar where I am questioning ChatGPT in order to learn about X. Of course this could radically change a student's ability to generate homework etc, but this could also radically change how the student learns in the first plac…

I agree, but typical GPT use is actually the opposite of the traditional Socratic mode in which the teacher uses questions to guide the student to understanding. But I wonder how it would do if it was prompted to use the Socratic method.

Duolingo is experimenting with a GPT-4 fine tune/wrapper which makes it act as a Socratic method teacher.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#837

Earlier quoted context omitted.

That demo is pretty slick. What happens when you go totally off book? Like, ask it to recite the numbers of pi? Or if you become abusive? Will it call the cops?

It's trained to ignore everything else. That way background conversations are ignored as well (like your kids talking in the back of the car while you order).

How do you train for this?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#838

Earlier quoted context omitted.

I work as a ethical hacker, so I'm well aware of the phishing and impersonation possibilities. But the net positive is so, so much bigger for society that I'm sure we'll figure it out. And yes, in 20 years you can tell your kids that 'back in my day' support consisted of real people. But truthfully, as someone who worked on a ISP helpdesk it's much better for society if these people move on to more productive areas.

> But truthfully, as someone who worked on a ISP helpdesk it's much better for society if these people move on to more productive areas. But is it, though? I started my career in customer support for a server hosting company, and eventually worked my way up to sysadmin-type work. I would not have been qualified for the position I eventually moved to at the start, I learned on the job. Is it really better for society…

Historically this exact same thing has happened, it was one of the bigger arguments against the abolition of child labour. "How will they grow up to be workers if they're not doing these jobs where they can learn the skills they'll need?"

The answer then was extending schooling, so that people (children at the time) could learn those skills without having their labour exploited. I would argue we should consider that today, extend mandatory free schooling. The economic purpose of education is that at the end of it the person should be able to have a job, removing entry level jobs doesn't change the economic purpose of education, so extend education until the person is able to have a job at the end of it again.

The social purpose of schooling is to make good members of society, and I don't think that cause would be significantly harmed by extending schooling in order for students to have learned enough to be more capable than an LLM in the job market.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#839

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

Do you have a rough design outline of what you built? I feel like we're on the cusp of something like this and it sounds amazing.

I'm using Llama2-chat-13B via mlc-llm @ 4bit quantization + whisper-streaming + coqui TTS, all running simultaneously on one 4090 in real time.

It didn't take long to prototype. Polishing and shipping it to non-expert users would take much longer than I've spent on it so far. I'd have to test for and solve a ton of installation problems, find better workarounds for whisper-streaming's hallucination issues, improve the heuristics for controlling when to start and stop talking, tweak the prompts to improve the suitability of the LLM responses for speech, fixup the LLM context when the LLM's speech is interrupted, probably port the whole thing to Windows for broader reach in the installed base of 4090s, possibly introduce a low-memory mode that can support 12GB GPUs that are much more common, document the requirements and installation process, and figure out hosting for the ginormous download it would be. I'd estimate at least 10x the effort I've spent so far on the prototype before I'd really be satisfied with the result.

I'd honestly love to do all that work. I've been prioritizing other projects because I judged that it was so obvious as a next step that someone else was probably working on the same thing with a lot more resources and would release before I could finish as a solo dev. But maybe I'm wrong...

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#840
post #600

Earlier quoted context omitted.

The bike shown in the first image is Specialized Sirrus X. You can make out from the image of the manual that it says "spacer/axle/bolt specifications". Searching for this yields the following Specialized bike manual which is similar: https://www.manualslib.com/manual/1974494/Specialized-Epic-E... -- there are some notable differences, but the Specialized Sirrus X manuals that are online aren't in the same style. The…

It bugged me that they made no mention of torque. The manual is really clear on that part with a big warning: > WARNING! Correct tightening force on fasteners (nuts, bolts, screws) on your bicycle is important for your safety. If too little force is applied, the fastener may not hold securely. If too much force is applied, the fastener can strip threads, stretch, deform or break. Either way, incorrect tightening forc…

The seat collar also probably has the max torque printed on it. <<<< Nope. There's no need for a torque wrench on that one.
Post reply on HN