Live data from Hacker News

GPT‑Live

openai.com

291–300 of 556 posts

Re: GPT‑Live

#292
post #210

What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an a…

There's a tiktok of Sam Altman reacting to a viral clip of someone using Voice to time themselves on a mile run (it hilariously failed). Sam's reaction was "Yea, it doesn't have access to tools like a timer. It's a known issue. Should be coming in about a year" Edit: here's the clip: https://www.youtube.com/shorts/Py2YgJe8fqQ

One of the videos on the announcement page shows someone making coffee and it _appears_ that the agent counts 30 seconds in real-time. Curious to know how they made it do that without tool support.

Re: GPT‑Live

#293

I for one am greatly looking forward to the day these kind of voice models can be run locally. It seems like the gap between open-weight and frontier is way larger for voice models than coding/language models.

Probably not _as_ good, but you can run gemma 4 for the ears/brain (accepts audio input), and kokoro TTS for the mouth. You want something like silero VAD sitting in front of the LLM so you aren't passing wasteful audio to gemma 4, so you send only voice activity segments. I type all this because I tried it out recently as a Zork experience, seeing how creative Gemma 4 12B could be. It was surprisingly good!

Re: GPT‑Live

#294
post #267
post #19

I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects. The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier. I did report a fun bug with it though: it…

Does it have full access to your chat history, project files, etc? That's the biggest limitation I have with voice mode right now, if I ask it about something I chatted about before (even in the same conversation) it has zero recollection of it.

You can start a chat in an existing session after pasting data into that session and it can then talk about that content. I haven't tried it with projects.

Re: GPT‑Live

#295
It was obvious from the live demo that this thing still hasn't learned when to shut up. When it stops tacking on "I'm here when you need me" to every response that could have just been "ok" or simply silence, maybe they'll have something. I think voice remains OpenAI's most disappointing product.

Re: GPT‑Live

#296

(Atty from OpenAI here) GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction. Would love to hear your feedback!

I'm interested in how you can present simultaneous rich visual information about what is happening the side delegation work.

i.e. how will full duplex & delegation enable/enhance desktop flows w/o corresponding leaps in UI.

Re: GPT‑Live

#297
I don't have many opinions about how individuals use this tech (although the AI as friend trend is a bummer for many reasons), but have maaaany (negative) opinions about the customer service industrial complex that's already using this in what seems to be an attempt to fool people into thinking they're speaking to a real person. Which is why I now, like a freak I never thought I'd need become, always ask "Am I speaking to a bot or a human?" when dealing with CS. So far, it's worked, and the bot transfers me. But, I fear the bot will eventually be programmed to lie abiht that, as well.

Re: GPT‑Live

#298
Like a lot of AI things, this seems both cool and kind of creeps me out. I've never used voice interfaces in the past (siri, the google one, whatever is on my tv) so I'm probably not the target market, but this does seem like an improvement.

The part that creeps me out is, we're living in an era where we're more disconnected from each other than ever before. Do we really need to be replacing conversations?! The demonstration video of old ladies sort of hints at something for me, which I think we already have a societal problem with the way we treat the elderly (and a massive elderly-loneliness issue) and there's kind of a sadness of imagining people becoming really close with this machine that doesn't really think. Definite ick factor.

Re: GPT‑Live

#299
post #271

Earlier quoted context omitted.

Fewer people aren't staring into their phones or talking to them -- makes your social antennas pick up automatically on not wanting to disturb them (lest you draw their ire for not having the social antennas long enough to pick up on the fact they're "busy and don't want to engage with you" like a gymrat with AirPods to signal they're there to pump in peace and quiet listening to their favourite playlist, not talk to…

There are still lots of social people. I found a lot of people actually do want to talk but are just shy. I spent a few weeks at a hostel last year. It was always kind of depressing and tense in the shared kitchen, just this heavy silence. I don't feel comfortable around strangers, so I solved that problem by just saying hi to everyone. Most people didn't respond much, although most of them smiled and the tension was…

This is an individual solution to a systemic problem. On a personal level, it is possible to solve such problems, but generally no, the ship has sailed quite some time ago (I personally think cars are to blame).

Even if you do it, you are still swimming against the current.

Re: GPT‑Live

#300

I worry what this will do to human communication if it becomes commonplace. Will everyone learn to be a forceful speaker, speaking over anyone they want to stop speaking?

Ciao.
Post reply on HN