Live data from Hacker News

GPT-4o

openai.com

831–840 of 1001 posts

Re: GPT-4o

#831

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

All I could think about when watching this demo was how similar capabilities will work on the battlefield. Coordinated AIs look like they will be obscenely effective. Everything always starts as a toy.

The "Killer app" for AGI/ASI is, I suspect, going to be in robotics, even more so than in replacing "white collar workers".

That includes, beyond literal Killers, all kinds of manufacturing, construction and service work.

I would expect a LOT of funds to go into research all sorts of actuators, artificial muscles and any other technology that will be useful in building better robots.

Companies that can get and maintain a lead in such technologies may reach a position similar to what US Steel had in the 19th century.

That could be the next nvidia.

I would not be at all surprised if we will have a robot in the house in 10 years that can clean and do the dishes, and that is built using basically the same parts as the robots that replace our soldiers and the police.

Who will ultimately control them, though?

Re: GPT-4o

#832
What popped out to me in the "bunny ear" video, the bunny ears are not actually visible to the phone's camera Greg is holding. Are they in the background feeding the production camera and this is not really a live demo?

Re: GPT-4o

#833

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

Like another comment mentioned, sigmoid curves [1] are ubiquitous with neural network systems. Neural network systems can be intoxicating because it's so "easy" (relatively speaking) to go from nothing to 80% in extremely short periods of time. And so it seems completely obvious that hitting 100% is imminent. Yet it turns out that each percent afterwards starts coming exponentially more slowly, and we tend to just bu…

> So long as LLMs regularly hallucinate, they're not going to be useful for much other than tasks that can accept relatively high rates of failure.

But is this not what humans do, universally? We are certainly good at hiding it – and we are all good at coping with it – but my general sense when interacting with society is that there is a large amount of nonsense generated by humans that our systems must and do already have enormous flexibility for.

My sense is that's not an aspect of LLMs we should have any trouble with incorporating smoothly, just by adhering to the safety nets that we built in response to our own deficiencies.

Re: GPT-4o

#834

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

My new PC arrives tomorrow. Once I source myself two RTX 3060's I'll be an AI owner, no longer dependant on cloud APIs. Currently the bottleneck is Agents. If you want a large language model to actually do anything you need an Agent. Agents so far need a human in the loop to keep them sane. Until that problem is solved most human jobs are still safe.

GPT 4o incorporated multimodality directly in the neural network, while reducing inference costs to half.

I fully expect GPT 5 (or at the latest 6) to similarly have native inclusion of agentic capabilities either this year or next year, assuming it doesn't already, but is just kept from the public.

Re: GPT-4o

#836

I much prefer a GLADOS-type AI voice than one that approximates an endlessly happy chipper enthusiastic personal assistant. I think the AI tutor is probably the strongest for actual real-world value delivered the rest of them are cool but a bit questionable as far as actual pragmatic usefulness. It'd be cool if an AI calling the another AI would recognize it'd talking to an AI and then they agree to ditch the fake co…

> Or upgradable offensive capabilities to outmaneuver the customer service agents when they try to decline your warranty or whatever.

Yeah, OpenAI is not going to do that out of fear of liability. But that's where open source LLM projects will come into play, eg Dolphin.

Re: GPT-4o

#837

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

Like another comment mentioned, sigmoid curves [1] are ubiquitous with neural network systems. Neural network systems can be intoxicating because it's so "easy" (relatively speaking) to go from nothing to 80% in extremely short periods of time. And so it seems completely obvious that hitting 100% is imminent. Yet it turns out that each percent afterwards starts coming exponentially more slowly, and we tend to just bu…

> So long as LLMs regularly hallucinate, they're not going to be useful for much other than tasks that can accept relatively high rates of failure.

Yep. So basically they're useful for a vast, immense range of tasks today.

Some things they're not suited for. For example, I've been working on a system to extract certain financial "facts" across SEC filings. ChatGPT has not been helpful at all either with designing or implementing (except to give some broad, obvious hints about things like regular expressions), nor would it be useful if it was used for the actual automation.

But for many, many other tasks -- like design, architecture, brainstorming, marketing, sales, summarisation, step by step thinking through all sorts of processes, it's extremely valuable today. My list of ChatGPT sessions is so long already and I can't imagine life without it now. Going back to Google and random Quora/StackOverflow answers laced with adtech everywhere...

Re: GPT-4o

#838

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

> We’re moving toward a world where every job will be modeled After an OpenAI launch, I think it's important to take one's feelings about the future impact of the technology with a HUGE grain of salt. OpenAI are masters of hype. They have been generating hype for years now, yet the real-world impacts remain modest so far. Do you remember when they teased GPT-2 as "too dangerous" for public access? I do. Yet we now ha…

> Do you remember when they teased GPT-2 as "too dangerous" for public access? I do.

I can't help but notice the huge amount of hindsight and bad faith that it demonstrated here. Yes, now we are aware that the internet did not drown in a flood of bullshit (well, not noticeably more), when GPT-2 was released.

But was it obvious? I certainly thought that there was a chance that the amount of blog spam that could be generated effortlessly might just make internet search unusable. You are declaring "hype", when you could also say "very uncertain and conscientious". Is this not something we want people in charge to be careful with?

Re: GPT-4o

#839
That woman's voice intonation is just scary.Not because it talks really well, but because it is always happy, optimistic, enthusiastic. And this echoes to what several of my employers idealized as a good employee.

That's terrifying because those AI become what their master's think an engaging human should be. It's quite close to Bostondynamics di some years ago. what did they show ? You can hit a robot very hard while it does its job and then what ? It just goes on without complaining. A perfect employee again.

That's very dystopic to me.

(but I'm impressed by the technical achievement)

Re: GPT-4o

#840

I've worked quite a bit with STT and TTS over the past ~7 years, and this is the most impressive and even startling demo I've seen. But I would like to see how this is integrated into applications by third party developers where the AI is doing a specific job. Is it still as impressive? The biggest challenge I've had with building any autonomous "agents" with generic LLM's is they are overly gullible and accommodatin…

I think if you listen to the way it answers, it seems its using a technique trained speakers use. To buy itself time to think, it repeats/paraphrases the question/request before actually answering.

I'm sure you'll find this part is a lot quicker to process, giving the instant response (the old gpt4-turbo is generally very quick with simple requests like this). Rather impressively all it would need is an additional custom instruction.

Very clever and eerily human.

Post reply on HN