Live data from Hacker News

GPT‑Live

openai.com

451–460 of 556 posts

Re: GPT‑Live

#451

Like a lot of AI things, this seems both cool and kind of creeps me out. I've never used voice interfaces in the past (siri, the google one, whatever is on my tv) so I'm probably not the target market, but this does seem like an improvement. The part that creeps me out is, we're living in an era where we're more disconnected from each other than ever before. Do we really need to be replacing conversations?! The demon…

None of the things I'm talking to ChatGPT about are replacing human conversations. It's replacing having to type a question into Google on my phone.

But you do bring up a good point with old people with no one to talk to. This technology would be a god send for them. And if you think otherwise I hope your next words are that you routinely visit nursing homes and talk with old people.

Re: GPT‑Live

#452
post #397

Earlier quoted context omitted.

Honestly I find it distracting when people do the mmmhmm thing. I don’t find it encouraging or helpful or whatever.

When I go into listening node, I listen . Actively and intensely. Apparently it unnerves people. My wife coaches me to give occasional Yes,Go On, or UhHuh. But it's a conscious, learned, active mechanism for me. Intuitively, I'll tell you if we've weered off my desired conversation path and I'd ask the same courtesy. I've learned very late in life it's apparently an autism thing. Either way, I don't seek mmms and umm…

TCP vs UDP. ;) Both have their uses.

Re: GPT‑Live

#453
post #19

I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects. The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier. I did report a fun bug with it though: it…

Can it in some way do some action based on the talk? For example, can it summarize the brainstorming into a text file, post it somewhere etc.?

Re: GPT‑Live

#454

Earlier quoted context omitted.

A year? Is it that hard, or do they not see the value?

A timer tool with a callback feature would take like a couple of hours to implement, a model which has a native internal ability to know the time would take ages to make

I don't think it needs an 'internal clock'. It needs two things. (1) the timer tool, (2) to know not to lie about having a tool or using a tool. I think (2) is the hard thing.

Re: GPT‑Live

#455
Once this gets video capabilities and is ported to glasses, it'll be a major revolution for blind people (and I say this as a blind person).

People have tried "smart that helps blind people navigate" since the 80s, many, many, many times, and all such projects failed. The cycle of "wow, blind people could benefit from a navigation aid, why don't I make one, if there's none around, I must surely have been the first bright university student to think of this idea" is pretty well known in the community, and I'm personally quite tired of it. Nevertheless, I think this may be the one.

Circa 2020, I have said that people who are getting a guide dog now are probably getting their last one. I think we aren't far off from that prediction coming true.

Re: GPT‑Live

#456

A nightmare scenario would be that people become so accustomed to talking to agreeable AI, that they lose the inability to talk to anything that disagrees with them or has different perspectives with responses that don't include stroking the ego.

It's a very interesting phenomenon. It's often said that famous stars become delusional exactly because they are surrounded by yes-men who are too afraid of rocking the boat and of losing their relationship with the star.

But if I have to think of my own experiences here, I would say that overly agreeable people become tiresome pretty quickly. To me their perspective does not offer anything of value without any pushback, as it certainly doesn't help with grounding my own thoughts. Perhaps it's why being too nice makes it difficult to form deeper bonds, and maybe paradoxically it is therefore a good thing that LLMs are overly agreeable.

It probably also depends on one's mindset - those who are interested in growing could be more likely interested in opposing views, while those who perceive that they have already "made it" (e.g. stars) perhaps don't care so much and prefer an agreeable tone.

Re: GPT‑Live

#457
At first I was really impressed, I thought granny was the voice. But it turns out it's the same kind of annoying voice and tone. And then Constance starts talking to it and it immediately cuts her off at 1:05, after they just explained it was better at conversation flow.

Seems a bit disappointing, but the 3 overlapping questions example was impressive.

Re: GPT‑Live

#458

What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an a…

If you’re serious about this, let me know!

This is something I built for myself, and to experiment with inference stacks. You can obviously just transcribe audio and hand it off to frontier models, so all you really need is a good voice stack and a “driver” for the interaction (like a phone call, place to see their work).

There are two big problems with this space IMO. One isn’t that you can’t get this to work but that people generally aren’t willing to pay for it for themselves, rather as a way to screen or automate stuff to be used by other people. Did you know Claude Code has a voice mode and that openai launched whisper a year ago, both of which have positive sentiment and adoption in heavy ai tool users? Yet it’s a blip in their marketing or why people use their products, meanwhile outside of coding, most of the biggest and highest earning AI product companies so far are voice agents targeting customer service, sales, business processes, etc.

The second is related: voice is genuinely a low-bandwidth medium, so as a primary interface for interacting with AI there is not a lot you can get out of it compared to eg complex technical work or visualizations or interactive applications. It is physically and mentally demanding to speak-aloud a highly detailed prompt fast enough that VAD won’t cut you off and you have something with comparable information density or specificity vs text. But to keep up a shorter and more natural cadence you’ll not be able to wait on a lot of thinking/tool unless you play UI tricks (ums and fillers, two models in a trench coat), break the illusion of a single coherent conversation, or take a lot of long pauses.

That’s why for the supplementary coding use case it’s mostly used for remote steering, and for general use marketed towards the large and very not-online group of people for whom typing is not a natural or common thing for them to spend their time on. Now that so much spend goes through heavily used token subscriptions and they’ve proven that kind of product, they’re not marketing “tool to get the most tokens per $ running your subscription 24/7” anymore lol.

What I’m most interested in is true “ambient” tool use against my own data or work, and for-later (or pushed live via your phone) visualizations or “five models in a trench coat but still coherent” UX, which you probably are too. But I think unless you work a lot with AI tools already it’s hard to understand how that’s any different from asking Alexa to set a timer, and either way something you’re not so desperate to have that you go looking for it, or pay smaller vendors/set up yourself.

Re: GPT‑Live

#459

(Atty from OpenAI here) GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction. Would love to hear your feedback!

Any feedback from different locations or cultural groups?

One group's expectation of interruption for pleasant conversational flow can be just as off-putting as another's expectation of patient silence.

Re: GPT‑Live

#460

(Atty from OpenAI here) GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction. Would love to hear your feedback!

I like it! I was watching a YouTube video of someone driving in a big city. I saw an interesting skyscraper and asked Chatgpt what it was. After Chatgpt answered, I asked for a photo of the building to confirm, and Chatgpt helpfully showed it in-chat. It was the correct building!

And because the voice is so frictionless to talk to, I asked about what company owns the building, then that company's industry, then how that industry works in this particular country etc. I probably wouldn't have bothered going down a rabbit hole like this if I'd had to type. Voice is much easier than typing.

Anyhow it's fun! Thanks for making it!

Post reply on HN