Live data from Hacker News

GPT‑Live

openai.com

461–470 of 556 posts

Re: GPT‑Live

#461
I just tried this and I find the constant interruptions infuriating: ah, alright, mh, mh-mh, go on, etc. This means the model already reliably detected my point/question isn’t finished. Please give me an option or dial to tone down the back channel noise.

Re: GPT‑Live

#462

Earlier quoted context omitted.

Hey! Bit of an unusual question maybe: if this stuff further exarcerbates the loneliness epidemic and atomization of society, will you be able to live with yourself you think? If you hear about teenagers only spending time with your chatbot in 5 years, will you feel some amount of personal responsibility or not? Always curious to hear you guys' perspective on that kind of stuff!

People who do this kind of stuff are very irritating. You clearly have some problem with the work they do. Instead of saying and approaching that outright, you pass it in some passive aggressive fake bullshit. Makes you sound like the kind of person I would much rather not be speaking to, which is kind of ironic given your comment.

[deleted]

Re: GPT‑Live

#463

Earlier quoted context omitted.

Hey! Bit of an unusual question maybe: if this stuff further exarcerbates the loneliness epidemic and atomization of society, will you be able to live with yourself you think? If you hear about teenagers only spending time with your chatbot in 5 years, will you feel some amount of personal responsibility or not? Always curious to hear you guys' perspective on that kind of stuff!

> if this stuff further exarcerbates the loneliness epidemic and atomization of society, will you be able to live with yourself you think? Looking at the 30,000-foot view of how society is set up: laws, economic system, employee incentives, etc, do you suppose it matters what the individual contributors think? I say this not to absolve anyone of responsibility, but to point out the obvious outcomes of our incentives…

[dead]

Re: GPT‑Live

#464

Earlier quoted context omitted.

People who do this kind of stuff are very irritating. You clearly have some problem with the work they do. Instead of saying and approaching that outright, you pass it in some passive aggressive fake bullshit. Makes you sound like the kind of person I would much rather not be speaking to, which is kind of ironic given your comment.

i mean you chose to be irritated by it ¯\_(ツ)_/¯

[deleted]

Re: GPT‑Live

#465
post #19

I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects. The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier. I did report a fun bug with it though: it…

But does it generate good pelicans?

I had an enjoyable conversation with it just now where it repeatedly told me it had generated an SVG of a pelican riding a bicycle along with a clickable data: link. Obviously, this is just harness polish. Overall, it's quite impressive.

Someday we'll have a tokeniser that works directly on human speech instead of doing speech-to-text and text-to-speech steps in between.

Re: GPT‑Live

#466

watched the live translation video very impressive Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly

I've wondered if this would happen, although doing inference directly on speech tokens would seem to imply an entirely different model (trained on lots and lots of actual speech).

Re: GPT‑Live

#467
post #388

Much better than it was before but it’s still significantly weaker than a direct chat. For example I asked “Why should LLM attention use dot product instead of cosine similarity, being that we often hear vector magnitude does not encode most of the useful information needed”? The voice response was directionally right but lacked detail and was a little hand wavy. The answer to the same question in a text chat was muc…

You can read far more text than listening. If the voice response is too long, people lose the patience quickly. So it is better to show it on screen if we have too much text.

I think the bare truth is that the target audience for this product is not people who are highly particular about terminology in answers involving vector mathematics.

It’s a different set of tradeoffs for users that don’t already have strong engagement or interest in existing AI products.

Do you realize how many more people prefer to chat over the phone and watch television or videos in their free time vs type multiple paragraphs of text into a chat window and then read 3x more back?

I’m not even talking about grandma here, it’s a non starter for the vast majority of humans who don’t spend their free time writing and reading tech news. To most people, having to write out a bunch of words describing their problem/goals, then sift through pages and pages of detailed response to get an answer, feels overwhelming and not worth doing.

Re: GPT‑Live

#468

What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an a…

Funny enough I have only encountered two voice modes that were okay to use tools and they are so far out there that you wouldn't even think to try. Bixby and Perplexity (in android phone digital assistant mode) both seem happy to use whatever the account is connected to. I mainly use it for managing my local phone's calendar. Claude chat can interact with it but voice can't which is frustrating.

Perplexity can use “some” tools (mostly built-in) but custom connectors were from my experience not available in voice.

I usually do a very simple test and tell it to verbatim tell me all tools that are connected

Re: GPT‑Live

#469
post #118

What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an a…

Because they take too long to run, and have an unpredictable latency and success rate. Seeing loading spinners and error messages in a visual interface is fine, but it would firmly put a natural language conversation in uncanny valley territory. Regular chat already supports voice input, so might as well use that.

It can simply say “Okay let me try to connect to Notion and add this, just a sec”

Re: GPT‑Live

#470
post #19

I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects. The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier. I did report a fun bug with it though: it…

I love the UX of voice mode. I’ve always hated voice models. Gemini’s being particularly egregious (always ending in some deranged question, can not reliably be prompted away) prompted me to build my own client for my real harness that simply does STT -> model -> TTS (both being independently useful). I guess I see some value in a model responding quickly and with more nuance, but it’s not much. I can wait for it to…

I’ve spent my time very similarly working on my own voice stack project, but having also seen how non-developers use AI or experience technology in general, I truly think they are better served with a different UX and product than what we have.

In other words, if you’re building your own voice inference tooling you’re just about the polar opposite user demographic than the one that truly needs and will value this. You’re using voice as a medium of convenience doing what existing models are technically and practically “shaped” to be able to do, knowing how they work well enough that conversation is more like typing/prompting with your voice than a natural interface. I’m guilty of this myself but have you ever even paid for a voice/audio model or hardware?

Compare that to the millions of people with an Alexa device in their home who buy products through it, or who prefer calling support to get a human over poring over technical documentation. They’re actually very close to finally getting a version of “Alexa” that lives up to its promise and I’m happy for them

Post reply on HN