Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

171–180 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#171

Earlier quoted context omitted.

If they show a car driving I believe it's capable of self-propulsion and not just rolling downhill.

A marketing trick that has, in fact, been tried: https://arstechnica.com/cars/2020/09/nikola-admits-prototype...

Used to be "marketing tricks" were prosecuted as fraud.

Re: Gemini "duck" demo was not done in realtime or with voice

#172
post #156

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

I don't see it as a problem with most non-critical uses cases (critical being things like medical diagnoses, controlling heavy machinery or robotics, etc). LLMs right now are most practical for generating templated text and images, which when paired with an experienced worker, can make them orders of magnitude more productive. Oh, DALL-E created graphic images with a person with 6 fingers? How long would it have take…

>> Nothing there they couldn't fix in a few minutes and then SHIP.

If by ship, you mean put directly into the public domain then yes.

https://www.goodwinlaw.com/en/insights/publications/2023/08/...

and for more interesting takes: https://www.youtube.com/watch?v=5WXvfeTPujU&

Re: Gemini "duck" demo was not done in realtime or with voice

#173
post #161

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

I’m not an expert but I suspect that this aspect of lack of correctness in these models might be fundamental to how they work. I suppose there’s two possible solutions: one is a new training or inference architecture that somehow understand “facts”. I’m not an expert so I’m not sure how that would work, but from what I understand about how a model generates text, “truth” can’t really be a element in the training or i…

> this aspect of lack of correctness in these models might be fundamental to how they work.

Is there some sense in which this isn't obvious to the point of triviality? I keep getting confused because other people seem to keep being surprised that LLMs don't have correctness as a property. Even the most cursory understanding of what they're doing understands that it is, fundamentally, predicting words from other words. I am also capable of predicting words from other words, so I can guess how well that works. It doesn't seem to include correctness even as a concept.

Right? I am actually genuinely confused by this. How is that people think it could be correct in a systematic way?

Re: Gemini "duck" demo was not done in realtime or with voice

#174
post #86

The whole Gemini webpage and contents felt weird to me, it's in the uncanny valley of trying to look and feel like an Apple marketing piece. The hyperbolic language, surgically precise ethnic/gender diversity, unnecessary animations and the sales pitch from the CEO felt like a small player in the field trying to pass as a big one.

> surgically precise ethnic/gender diversity

What does that mean and why is it bad?

Diversity in marketing is used because, well, your desired market is diverse.

I don't know what it means for it to be surgically precise, though.

Re: Gemini "duck" demo was not done in realtime or with voice

#175

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

Most people I worked with either tell me "I don't know" or "I think x, but with not sure" when they are not sure about something, the issue with LLMs is they don't have this concept.

Re: Gemini "duck" demo was not done in realtime or with voice

#176

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

I find it ironic that computer scientists and technologists are frequently uberrationalists to the point of self parody but they get hyped about a technology that is often confidently wrong.

Just like the hype with AI and the billions of dollars going into it. There’s something there but it’s a big fat unknown right now whether any part of the investment will actually pay off - everyone needs it to work to justify any amount of the growth of the tech industry right now. When everyone needs a thing to work, it starts to really lose the fundamentals of being an actual product. I’m not saying it’s not useful, but is it as useful as the valuations and investments need it to be? Time will tell.

Re: Gemini "duck" demo was not done in realtime or with voice

#177

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

I think this problem needs to be solved at a higher level, and in fact Bard is doing exactly that. The model itself generates its output, and then higher-level systems can fact check it. I've heard promising things about feeding back answers to the model itself to check for consistency and stuff, but that should be a higher level function (and seems important to avoid infinite recursion or massive complexity stemming from the self-check functionality).

Re: Gemini "duck" demo was not done in realtime or with voice

#178

Earlier quoted context omitted.

You think you ask your legal assistant to find some precedents related to your current case and they will come back with an A4 page full of made up cases that sound vaguely related and convincing but are not real ? I don't think you understand the failure case at all.

That example seems a bit hyperbolic. Do you think lawyers who leverage ChatGPT will take the made up cases and present them to a judge without doing some additional research? What I'm saying is that the tolerance for mistakes is strongly correlated to the value ChatGPT creates. I think both will need to be improved but there's probably more opportunity in creating higher value. I don't have a horse in the race.

> Do you think lawyers who leverage ChatGPT will take the made up cases and present them to a judge without doing some additional research?

You don't?

https://fortune.com/2023/06/23/lawyers-fined-filing-chatgpt-...

Re: Gemini "duck" demo was not done in realtime or with voice

#179

Earlier quoted context omitted.

> I'm unsure where this expectation of 100% absolute correctness comes from. It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical j…

Speech to text is often wrong too. So is autocorrect. And object detection. Computers don't have to be 100% correct in order to be useful, as long as we don't put too much faith in them.

Your caveat is not the norm though, as everyone is putting a lot of faith in them. So, that's part of the problem. I've talked with people that aren't developers, but they are otherwise smart individuals that have absolutely not considered that the info is not correct. The readers here are a bit too close to the subject, and sometimes I think it is easy to forget that the vast majority of the population do not truly understand what is happening.

Re: Gemini "duck" demo was not done in realtime or with voice

#180

Earlier quoted context omitted.

You think you ask your legal assistant to find some precedents related to your current case and they will come back with an A4 page full of made up cases that sound vaguely related and convincing but are not real ? I don't think you understand the failure case at all.

That example seems a bit hyperbolic. Do you think lawyers who leverage ChatGPT will take the made up cases and present them to a judge without doing some additional research? What I'm saying is that the tolerance for mistakes is strongly correlated to the value ChatGPT creates. I think both will need to be improved but there's probably more opportunity in creating higher value. I don't have a horse in the race.

> Do you think lawyers who leverage ChatGPT will take the made up cases and present them to a judge without doing some additional research?

I generally agree with you, but it's funny that you use this as an example when it already happened. https://arstechnica.com/tech-policy/2023/06/lawyers-have-rea...

Post reply on HN