Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

141–150 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#141

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

> I'm unsure where this expectation of 100% absolute correctness comes from.

It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical jokes. Like I honestly feel the devs are fucking with us.

Re: Gemini "duck" demo was not done in realtime or with voice

#142

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

Honestly I agree. Humans make errors all the time. Perfection is not necessary and requiring perfection blocks deployment of systems that represent a substantial improvement over the status quo despite their imperfections.

The problem is a matter of degree. These models are substantially less reliable than humans and far below the threshold of acceptability in most tasks.

Also, it seems to me that AI can and will surpass the reliability of humans by a lot. Probably not by simply scaling up further or by clever prompting, although those will help, but by new architectures and training techniques. Gemini represents no progress in that direction as far as I can see.

Re: Gemini "duck" demo was not done in realtime or with voice

#143
post #131

Earlier quoted context omitted.

Is it possible for humans to be wrong about something, without lying?

Lying implies an intent to deceive despite, or giving a response despite having better knowledge, which I'd argue LLMs can't do, at least not yet. It just requires a more robust theory of mind than I'd consider them to realistically be capable of. They might have been trained/prompted with misinformation, but then it's the people doing the training/prompting who are lying, still not the LLM.

Not to say this example was lying but they can lie just fine - https://arxiv.org/abs/2311.07590

Re: Gemini "duck" demo was not done in realtime or with voice

#144

I was fooled. The model release announcement said it could accept video and audio multi-modal input. I understood that there was a lot of editing and cutting, but I really believed I was looking at an example of video and audio input. I was completely impressed since it’s quite a leap to go from text and still images to “eyes and ears.” There’s even the segment where instruments are drown and music was generated. I t…

Well put. I’m not touching anything Google does any more. They’re far too dishonest. This failed attempt at a release (which turns out was all sizzle and no steak) only underscored how far behind OpenAI they actually are. I’d love to have been a fly on the wall in the OAI offices when this demo video went live.

Re: Gemini "duck" demo was not done in realtime or with voice

#145

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

Is it possible for humans to be wrong about something, without lying?

Most humans I professionally interact with don't double down on their mistakes when presented with evidence to the contrary.

The ones that do are people I do my best to avoid interacting with.

LLMs act more like the latter, than the former.

Re: Gemini "duck" demo was not done in realtime or with voice

#146

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

I know exactly where the expectation comes from. The whole world has demanded absolute precision from computers for decades.

Of course, I agree that if we want computers to “think on their own“ or otherwise “be more human“ (whatever that means) we should expect a downgrade in correctness, because humans are wrong all the time.

Re: Gemini "duck" demo was not done in realtime or with voice

#147

I have used Swype texting since the t9 days. If I demoed swype texting as it functions in my day to day life to someone used to a querty keyboard they would never adopt it The rate at which it makes wrong assumptions about the word, or I have to fix it is probably 10% to 20% of the time However because it’s so easy to fix this is not an issue and it doesn’t slow me down at all. So within the context of the different…

I think you mean swipe. Swype was a brilliant third party keyboard app for Android which was better at text prediction and manual correction than Gboard is today. If however you really do still use Swype then please tell me how because I miss it.

Re: Gemini "duck" demo was not done in realtime or with voice

#148

Earlier quoted context omitted.

Let's see, so we exclude law, we exclude medical.. it's certainly not a "vast minority" and the failure cases are nothing at all like search or human experts.

Are you suggesting that failure cases are lower when interacting with humans? I don't think that's my experience at all. Maybe I've only ever seen terrible doctors but I always cross reference what doctors say with reputable sources like WebMD (which I understand likely contain errors). Sometimes I'll go straight to WebMD. This isn't a knock on doctors - they're humans and prone to errors. Lawyers, engineers, product…

You think you ask your legal assistant to find some precedents related to your current case and they will come back with an A4 page full of made up cases that sound vaguely related and convincing but are not real? I don't think you understand the failure case at all.

Re: Gemini "duck" demo was not done in realtime or with voice

#149

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

Aside: this is not what impartial means.

Re: Gemini "duck" demo was not done in realtime or with voice

#150

Earlier quoted context omitted.

Do you believe everything verbatim that companies tell you in advertising?

If they show a car driving I believe it's capable of self-propulsion and not just rolling downhill.

Hmm, might I interest you in a video of an electric semi-truck?
Post reply on HN