Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

241–250 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#241

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

LLMs do not lie, nor do they tell the truth. They have no goal as they are not agents.

With apologies to Dijkstra, the question of whether LLMs can lie is about as relevant as the question of whether submarines can swim.

Re: Gemini "duck" demo was not done in realtime or with voice

#242

Earlier quoted context omitted.

> I'm unsure where this expectation of 100% absolute correctness comes from. It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical j…

Speech to text is often wrong too. So is autocorrect. And object detection. Computers don't have to be 100% correct in order to be useful, as long as we don't put too much faith in them.

People put too much faith in conspiracy theories they find on YT, TikTok, FB, Twitter, etc. What you're claiming is already not the norm. People already put too much faith into all kinds of things.

Re: Gemini "duck" demo was not done in realtime or with voice

#243

Earlier quoted context omitted.

I'm not a fan of current approaches here. "Chain of thought" or other approaches where the model does all its thinking using a literal internal monologue in text seem like a dead end. Humans do most of their thinking non-verbally and we need to figure out how to get these models to think non-verbally too. Unfortunately it seems that Gemini represents no progress in this direction.

> Humans do most of their thinking non-verbally and we need to figure out how to get these models to think non-verbally too. That's a very interesting point, both technically and philosophically. Where Gemini is "multi-modal" from training, how close do you think that gets? Do we know enough about neurology to identical a native language in which we think? (not rhetorical questions, I'm really wondering)

Neural networks are only similar to brains on the surface. Their learning process is entirely different and their internal architecture is different as well.

We don’t use neural networks because they’re similar to brains. We use them because they are arbitrary function approximators and we have an efficient algorithm (backprop) coupled with hardware (GPUs) to optimize them quickly.

Re: Gemini "duck" demo was not done in realtime or with voice

#244

Earlier quoted context omitted.

There's nothing wrong with what you're saying, but what do you suggest? Factuality is an area of active research, and Deepmind goes into some detail in their technical paper. The models are too useful to say, "don't use them at all." Hopefully people will heed the warnings of how they can hallucinate, but further than that I'm not sure what more you can expect.

The problem is not with the model, but with its portrayal in the marketing materials. It's not even the fact that it lied, which is actually realistic. The problem is the lie was not called out as such. A better demo would have had the user note the issue and give the model the opportunity to correct itself.

But you yourself said that it was so convincing that the people doing the demo didn't recognize it as false, so how would they know to call it out as such?

I suppose they could've deliberately found a hallucination and showcased it in the demo. In which case, pretty much every company's promo material is guilty of not showcasing negative aspects of their product. It's nothing new or unique to this case.

Re: Gemini "duck" demo was not done in realtime or with voice

#245
post #86

The whole Gemini webpage and contents felt weird to me, it's in the uncanny valley of trying to look and feel like an Apple marketing piece. The hyperbolic language, surgically precise ethnic/gender diversity, unnecessary animations and the sales pitch from the CEO felt like a small player in the field trying to pass as a big one.

> surgically precise ethnic/gender diversity What does that mean and why is it bad? Diversity in marketing is used because, well, your desired market is diverse. I don't know what it means for it to be surgically precise, though.

Agreed with your comment. This is every marketing department on the planet right now, and it's not a bad thing IMO. Can feel a bit forced at times, but it's better than the alternative.

Re: Gemini "duck" demo was not done in realtime or with voice

#246
post #58

Earlier quoted context omitted.

Ask for information that only the actual person would know.

That will only work once if the channels are monitored.

You only know one piece of information about your family? I feel like I could reference many childhood facts or random things that happened years ago in social situations.

Re: Gemini "duck" demo was not done in realtime or with voice

#247

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

If a human expert gave wrong answers as often and as confidently as LLMs, most would consider no longer asking them. Yet people keep coming back to the same LLM despite the wrong answers to ask again in a different way (try that with a human).

This insistence on comparing machines to humans to excuse the machine is as tiring as it is fallacious.

Re: Gemini "duck" demo was not done in realtime or with voice

#248

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

EDIT: never mind, I missed the exact wording about being "made of a material..." which is definitely false then. Thanks for the correction below. Preserving the original comment so the replies make sense: --- I think it's a stretch to say that's false. In a conversational human context, saying it's made of rubber implies it's a rubber shell with air inside. It floats because it's rubber [with air] as opposed to being…

This is what the reply was:

> Oh, it it's squeaking then it's definitely going to float.

> It is a rubber duck.

> It is made of a material that is less dense than water.

Full points for saying if it's squeaking then it's going to float.

Full points for saying it's a rubber duck, with the implication that rubber ducks float.

Even with all that context though, I don't see how "it is made of a material that is less dense than water" scores any points at all.

Re: Gemini "duck" demo was not done in realtime or with voice

#249

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

EDIT: never mind, I missed the exact wording about being "made of a material..." which is definitely false then. Thanks for the correction below. Preserving the original comment so the replies make sense: --- I think it's a stretch to say that's false. In a conversational human context, saying it's made of rubber implies it's a rubber shell with air inside. It floats because it's rubber [with air] as opposed to being…

> In a conversational human context, saying it's made of rubber implies it's a rubber shell with air inside.

Disagree. It could easily be solid rubber. Also, it's not made of rubber, and the model didn't claim it was made of rubber either, so it's irrelevant.

> It floats because it's rubber [with air] as opposed to being a ceramic figurine or painted metal.

A ceramic figurine or painted metal in the same shape would float too. The claim that it floats because of the density of the material is false. It floats because the shape is hollow.

> It's not false to say a house is made of wood.

It's false to say a house is made of air simply because its shape contains air.

Re: Gemini "duck" demo was not done in realtime or with voice

#250

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

People seem to want to use LLMs to mine knowledge, when really it appears to be a next-gen word-processor.
Post reply on HN