Live data from Hacker News

OpenAI rolls out Advanced Voice Mode with more voices and a new look

techcrunch.com

51–60 of 85 posts

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#51

I was hoping to find a voice like Microsoft's "Guy" (a name, not referring to the gender) or Google Assistant's "Pink". An unambiguously white, masculine, American "radio voice" or "audiobook narrator" voice. ChatGPT describes this as "A rich, deep, and smooth tone that is pleasant to listen to for extended periods. This often comes from good control over pitch and timbre, creating a voice that resonates well." If yo…

What could be the reason for excluding a white male voice, do you think?

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#52

> Advanced Voice is not yet available in the EU, the UK, Switzerland, Iceland, Norway, and Liechtenstein. That's disappointing. I wonder if it's related to legal issues, technical issues, or just doing a phased rollout?

Sam Altman only said this: https://x.com/sama/status/1838864011321872407 >except for jurisdictions that require additional external review

...I'm not at all clear what this external review would be for either the UK or Switzerland.

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#53

I was hoping to find a voice like Microsoft's "Guy" (a name, not referring to the gender) or Google Assistant's "Pink". An unambiguously white, masculine, American "radio voice" or "audiobook narrator" voice. ChatGPT describes this as "A rich, deep, and smooth tone that is pleasant to listen to for extended periods. This often comes from good control over pitch and timbre, creating a voice that resonates well." If yo…

[deleted]

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#54

Earlier quoted context omitted.

But that assumes that the speech recognition is perfect. Which at least for those of us speaking non-US English is never the case. And you have only to ring your bank and try and transfer money between accounts with it reading out every account number and asking for confirmation every step of the way. Versus a few clicks with a mouse to see that for almost all operational tasks voice is cumbersome and inefficient.

Whisper is incredibly robust, with a vast amount of language. I use it in German as well as English and it's incredibly reliable. Modern, transformer based ASR is a different ballgame.

Time for the mandatory mention of the Scottish voice recognition elevator sketch: https://www.youtube.com/watch?v=MNuFcIRlwdc

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#55

I have played with it for 20 minutes and here’s my review: 1. The low latency responses do make a difference. It feels miles better than any other voice chat out there. 2. Its pronunciation is excellent and very human like but it is not quite there. Somehow I can tell instantly that it’s a chatbot, it feels firmly in the uncanny valley. 3. On the same note if I was on call and there was a chatbot on the other side of…

> I find it harder to use than just typing to it Systems like this have existed since the 90s e.g. Dragon albeit far more rudimentary. And the issues are exactly the same: (a) discoverability, (b) efficiency and (c) recoverability. It is so much easier to have a screen with fixed options that you interact with, can easily see your journey and can go back for any mistakes. Versus with our voice which is the clunkiest,…

> Versus with our voice which is the clunkiest, slowest and least precise input method we have.

Some related issues I have:

- my thoughts always seem to be jumbled when talking to AI

- I rush to talk quickly because any pause seems to trigger a response

- I worry words or DSL I use won’t be interpreted properly

This all leads to a pretty poor voice experience for me, and I usually forget half of what I want to talk about.

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#56

I have played with it for 20 minutes and here’s my review: 1. The low latency responses do make a difference. It feels miles better than any other voice chat out there. 2. Its pronunciation is excellent and very human like but it is not quite there. Somehow I can tell instantly that it’s a chatbot, it feels firmly in the uncanny valley. 3. On the same note if I was on call and there was a chatbot on the other side of…

I generally like it, but anytime I bump up against the guidelines, which you do if you want to do basically anything fun at all (singing!) is the most obnoxious experience, because it feels 10x worse coming from a fake smiling personality that sounds almost human than it does over text.

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#57
post #43

Earlier quoted context omitted.

Well the. Surely OpenAI just need to make a consent popup for the UK they should do that asap

The blocker definitely isn't the current users of the application, but all the people involved in the creation of the training data.

I dont understand what this statement has to do with the conversation regarding releasing a new voice mode over an existing model thats allowed in the UK

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#58

Earlier quoted context omitted.

Sam Altman only said this: https://x.com/sama/status/1838864011321872407 >except for jurisdictions that require additional external review

...I'm not at all clear what this external review would be for either the UK or Switzerland.

Under a strict reading of the AI Act, advanced voice mode would be illegal because it is the "use of AI systems to infer emotions of a natural person in the areas of workplace and education institutions, except where the use of AI system is intended to be put in place or into the market for medical or safety reasons"

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#59
post #24

Earlier quoted context omitted.

> I find it harder to use than just typing to it Systems like this have existed since the 90s e.g. Dragon albeit far more rudimentary. And the issues are exactly the same: (a) discoverability, (b) efficiency and (c) recoverability. It is so much easier to have a screen with fixed options that you interact with, can easily see your journey and can go back for any mistakes. Versus with our voice which is the clunkiest,…

I understand how (a) discoverability and (c) recoverability are a problem, but what do you mean with (b) efficiency? Most people talk faster than they can type.

That assumes perfect accuracy. If a command is misheard then you probably need to correct whatever is now in the wrong state, and then definitely reissue the original command. If its text input then you have to do some select/correct dance. Both of these things take a lot of time.

Re: OpenAI rolls out Advanced Voice Mode with more voices and a new look

#60
Some review bullet points:

1. It's a bit too agreeable, example: "thats an excellent point" etc every single time.

2. It understands surprisingly well. example: from experience, when I explain something vaguely, my expectation is that it would not understand, but it does most of the time. It removes the frustration of needing to spell out in much more detail.

3. It feels like talking to a real person, but the way the AI talks in a sort of monotonic ways. Example: it would respond with similar tones/excitement every time.

4. Very useful if you need to chat but doesn't want to chat with humans about some subjects like ideas, and explainations.

Post reply on HN