Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

261–270 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#261

There was also the cringey "niiice!", "sweeeet!", "that's greaatt", "that's actually pretty good" responses from the narrator in a few of the demo videos that gave them the feel of a cheap 1980's TV ad.

It really reminds me of the Black Mirror episode Smithereens with the tech CEO talking with the shooter. Tech people really struggle with empathy, not just 1 on 1 but with the rest of the outside world which is predominantly low income relatively, with no college education. Paraphrased, Black Mirror ep was like:

[Tech CEO read instructions to "show empathy" from his assistant via Slack]

CEO: I hear you. It must be very hard for you.

Shooter: Of course you fucking hear me, we're on the phone! Talk like a normal person!

Re: Gemini "duck" demo was not done in realtime or with voice

#262
The hype really is drowning out the simple fact that basically no one really knows what these models are doing. Why does it matter so much that we include auto-correlation of embedding vectors as the "attention" mechanism in these models? And that we do this sufficiently many times across all the layers? And that we blindly smoosh values together with addition and call it a "skip" connection? Yes, you can tell me a bunch of stuff about gradients and residual information, but tell me why any of this stuff is or isn't a good model of causality.

Re: Gemini "duck" demo was not done in realtime or with voice

#263

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

EDIT: never mind, I missed the exact wording about being "made of a material..." which is definitely false then. Thanks for the correction below. Preserving the original comment so the replies make sense: --- I think it's a stretch to say that's false. In a conversational human context, saying it's made of rubber implies it's a rubber shell with air inside. It floats because it's rubber [with air] as opposed to being…

[deleted]

Re: Gemini "duck" demo was not done in realtime or with voice

#264

Earlier quoted context omitted.

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

> I'm unsure where this expectation of 100% absolute correctness comes from. It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical j…

"Computer says no" is not a meme for no reason.

Re: Gemini "duck" demo was not done in realtime or with voice

#265

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

>I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

As others hinted at, there's some bias because it's coming from a computer, but I think it's far more nuanced than that.

I've worked with many experts and professionals through my career ranging across medicine, various types of engineers, scientists, academics, researchers and so on and the pattern I often see is the level of certainty presented that always bothers me and the same is often embedded in LLM responses.

While humans don't typically quantify the certainty of their statements, the best SMEs I've ever worked with make it very clear what level of certainty they have when making professional statements. The SMEs who seem to be more often wrong than not speak in certainty quite often (some of this is due to cultural pressures and expectations surrounding being an "expert").

In this case, I would expect a seasoned scientist to say something in response to the duck question that: "many rubber ducks exist and are designed to float, this one very well might, we'd really need to test it or have far more information about the composition of the duck, the design, the medium we want it in (Water? Mecury? Helium?)" and so on. It's not an exact answer but you understand there's uncertainty there and we need to better clarify our question and the information surrounding that question. The fact is, it's really complex to know if it'll float or not from visual information alone.

It could have an osmimum ball inside that overcomes most the assumed buoyancy the material contains, including the air demonstrated to make it squeak. It's not transparent. You don't know for sure and the easiest way to alleviate uncertainty in this case is simply to test it.

There's so much uncertainty in the world, around what seem like the most certain and obvious things. LLMs seem to have grabbed some of this bad behavior from human language and culture where projecting confidence is often better (for humans) than being correct.

Re: Gemini "duck" demo was not done in realtime or with voice

#266

I have used Swype texting since the t9 days. If I demoed swype texting as it functions in my day to day life to someone used to a querty keyboard they would never adopt it The rate at which it makes wrong assumptions about the word, or I have to fix it is probably 10% to 20% of the time However because it’s so easy to fix this is not an issue and it doesn’t slow me down at all. So within the context of the different…

The insight here is that the speed of correction is a crucial component of the perceived long-term value of an interface technology. It is the main reason that handwriting recognition did not displace keyboards. Once the handwriting is converted to text, it’s easier to fix errors with a pointer and keyboard. So after a few rounds of this most people start thinking: might as well just start with the pointer and keyboa…

I think this is a good rebuttal.

Yeah the feedback loop with consumers has a higher likelihood of being detrimental, so even if the iteration rate is high, it’s potentially high cost at each step.

I think the current trend is to nerf the models or otherwise put bumpers on them so people can’t hurt themselves. That’s one approach that is brittle at best and someone with more risk tolerance (OpenAI) will exploit that risk gap.

It’s a contradiction then at best and depending on the level of unearned trust from the misleading marketing, will certainly lead to some really odd externalities

Think “man follows google maps directions into pond” but for vastly more things.

I really hated marketing before but yeah this really proves the warning I make in the AI addendum to my scarcity theory (in my bio).

Re: Gemini "duck" demo was not done in realtime or with voice

#267

Earlier quoted context omitted.

So what? People do this with Facebook news too. That's a people problem, not an LLM problem.

People on social media are absolutely 100% posting things deliberately to fuck with people. They are actively seeking to confuse people, cause chaos, divisiveness, and other ill intended purposes. Unless you're saying that the LLM developers are actively doing the same thing, I don't think comparing what people find on the socials vs getting back as a response from a chatBot is a logical comparison at all

How is that any different from what these AI chatbots are doing? They make stuff up that they predict will be rewarded highly by humans who look at it. This is exactly what leads to truisms like "rubber duckies are made of a material that floats over water" - which looks like it should be correct, even though it's wrong. It really is no different from Facebook memes that are devised to get a rise out of people and be widely shared.

Re: Gemini "duck" demo was not done in realtime or with voice

#269
post #35
post #4

The video itself and the video description give a disclaimer to this effect. Agreed that some will walk away with an incorrect view of how Gemini functions, though. Hopefully realtime interaction will be part of an app soon. Doesn’t seem like there would be too many technical hurdles there.

The entirety of the disclaimer is "sequences shortened throughout", in tiny text at the bottom for two seconds. They do disclose most of the details elsewhere, but the video itself is produced and edited in such a way that it's extremely misleading. They really want you to think that it's responding in complex ways to simple voice prompts and a video feed, and it's just not.

Yea, of all the edits in the video, the editing for timing is the least of concern. My gripe is that the prompting was different and in order to get that information you have to watch the video only on YouTube, expand the description and click on a link to a different blog article. Linking a "making of" video where they show this and interview some of the minds behind Gemini would have been better PR.

Re: Gemini "duck" demo was not done in realtime or with voice

#270

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

To be fair, one could describe the duck as being made of air and vinyl polymer, which in combination are less dense than water. That's not how humans would normally describe it, but that's kind of arbitrary; consider how aerogel is often described as being mostly made of air.
Post reply on HN