Live data from Hacker News

GPT-4o

openai.com

721–730 of 1001 posts

Re: GPT-4o

#721

Impressed by the model so far. As far as independent testing goes, it is topping our leaderboard for chess puzzle solving by a wide margin now: https://github.com/kagisearch/llm-chess-puzzles?tab=readme-o...

On one hand, some of these results are impressive; on the other, the illegal moves count is alarming - it suggests no reasoning ability as there should never be an illegal move? I mean, how could a violation of a fairly basic game (from a rules perspective) be acceptable in assigning any 'outcome' to a model other than failure?

Re: GPT-4o

#722

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

This is still gpt4. I don’t expect much more from this version than what previous version could do, in terms of reasoning abilities.

But everyone is expecting them to release gpt5 later this year, and it is a bit scary to think what it will be able to do.

Re: GPT-4o

#723

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

[flagged]

People with your sentiment said the same thing about all cool tech that changed the world. Doesn't change the reality, a lot of professions will need to adapt or they will go extinct.

Re: GPT-4o

#725
Does anyone have technical insight into how the screensharing in the math tutor video works? It looks like they start the broadcast from within the ChatGPT app, yet have no option to select which app will be the source of the stream. Or is that implied when both apps reside in the iPad's split view? And is this using regular ReplayKit or something new?

Re: GPT-4o

#727
post #569
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

But in this case you're not talking with a real person. Instinctively, I dislike a robot that pretends to be a real human being.

These sorts of comments are going to go in the annals with the hackernews people complaining about Dropbox when it first came out. This is so revolutionary. If you're not agog you're just missing the obvious.

Re: GPT-4o

#728
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

Somehow it also sounds almost like Dot Matrix from Spaceballs.

Joan Rivers!

Re: GPT-4o

#729

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

[flagged]

What do you mean? How would you define AI?

Re: GPT-4o

#730
Am I using it wrong? I have the gpt plus subscription, and can select "gpt4o" from the model list on ChatGPT, but whichever example I try from the example list under "Explorations of capabilities" on `https://openai.com/index/hello-gpt-4o/`, my results are worse:

* "Poetic typography" sample: I paste the prompt, and get an image with the typical lack of coherent text, just mangled letters.

* "Visual Narratives: Robot Writer's Block" - Mangled letters also

* "Visual Narratives: Sally the mailwoman" - not following instructions about camera angle. Sally looks different in each subsequent photo.

* "Meeting Notes with multiple speakers" - I uploaded the exact same audio file and used input 'How many speakers in this audio and what happened?'. gpt4o went off about about audio sample rates, speaker diarization models, torchaudio, and how its execution environment is broken and can't proceed to do it.

Post reply on HN