That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…
I’m not an expert but I suspect that this aspect of lack of correctness in these models might be fundamental to how they work. I suppose there’s two possible solutions: one is a new training or inference architecture that somehow understand “facts”. I’m not an expert so I’m not sure how that would work, but from what I understand about how a model generates text, “truth” can’t really be a element in the training or i…
Gemini "duck" demo was not done in realtime or with voice
281–290 of 683 posts
Re: Gemini "duck" demo was not done in realtime or with voice
#282That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…
Re: Gemini "duck" demo was not done in realtime or with voice
#283Earlier quoted context omitted.
To be fair, one could describe the duck as being made of air and vinyl polymer, which in combination are less dense than water. That's not how humans would normally describe it, but that's kind of arbitrary; consider how aerogel is often described as being mostly made of air.
Is an aircraft carrier made of a material that is less dense than water?
Re: Gemini "duck" demo was not done in realtime or with voice
#284I was fooled. The model release announcement said it could accept video and audio multi-modal input. I understood that there was a lot of editing and cutting, but I really believed I was looking at an example of video and audio input. I was completely impressed since it’s quite a leap to go from text and still images to “eyes and ears.” There’s even the segment where instruments are drown and music was generated. I t…
Re: Gemini "duck" demo was not done in realtime or with voice
#285The red flag for me was that they started that demo video with a background noise to make it seem like it's a raw video. A subtle manipulation for no reason, it's obviously not a raw video. The fact that they did not fact check the videos again makes me not particularly confident in the quality of Google's work. The bit where the model misinterpreted music notation (the circled area does not mean "piano"), and the "l…
Re: Gemini "duck" demo was not done in realtime or with voice
#286Re: Gemini "duck" demo was not done in realtime or with voice
#287They oversold its capabilities, but it does still seem that multi-modal models are going to be a requirement for AI to converge on a consistent idea of what kinds of phenomena are truly likely to be observed across modalities. So it's a good step forward. Now if they can just show us convincingly that a given architecture is actually modeling causality.
Re: Gemini "duck" demo was not done in realtime or with voice
#288Earlier quoted context omitted.
I’m not an expert but I suspect that this aspect of lack of correctness in these models might be fundamental to how they work. I suppose there’s two possible solutions: one is a new training or inference architecture that somehow understand “facts”. I’m not an expert so I’m not sure how that would work, but from what I understand about how a model generates text, “truth” can’t really be a element in the training or i…
> this aspect of lack of correctness in these models might be fundamental to how they work. Is there some sense in which this isn't obvious to the point of triviality? I keep getting confused because other people seem to keep being surprised that LLMs don't have correctness as a property. Even the most cursory understanding of what they're doing understands that it is, fundamentally, predicting words from other words…
This is maybe a pedantic "yes", but is also extremely relevant to the outstanding performance we see in tasks like programming. The issue is primarily the size of the correct output space (that is, the output space we are trying to model) and how that relates to the number of parameters. Basically, there is a fixed upper bound on the amount of complexity that can be encoded by a given number of parameters (obvious in principle, but we're starting to get some theory about how this works). Simple systems or rather systems with simple rules may be below that upper bound, and correctness is achievable. For more complex systems (relative to parameters) it will still learn an approximation, but error is guaranteed.
I am speculating now, but I seriously suspect the size of the space of not only one or more human language but also every fact that we would want to encode into one of these models is far too big a space for correctness to ever be possible without RAG. At least without some massive pooling of compute, which long term may not be out of the question but likely never intended for individual use.
If you're interested, I highly recommend checking out some of the recent work around monosemanticity for what fleshing out the relationship between model-size and complexity looks like in the near term.
Re: Gemini "duck" demo was not done in realtime or with voice
#289That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…
Well this seems like a huge nitpick. If a person said that, you would afford them some leeway, maybe they meant the whole duck, which includes the hollow part in the middle. As an example, when most people say a balloon's lighter than air, they mean an inflated balloon with hot air or helium, but you catch their meaning and don't rush to correct them.
Also, lighter-than-air balloons are intentionally filled with helium and sealed; rubber ducks are not sealed and contain air only incidentally. A balloon in a vacuum would still contain helium (if strong enough) but would not rise, while a rubber duck in a vacuum would not contain air but would still easily float on a liquid of similar density to water.
The reason why it seems like a nitpick is that this is such an inconsequential thing. Yeah, it's a false statement but it doesn't really matter in this case, nobody is relying on this answer for anything important. But the point is, in cases where it does matter these models cannot be trusted. A human would realize when the context is serious and requires accuracy; these models don't.
Re: Gemini "duck" demo was not done in realtime or with voice
#290Anyone with half a brain could see through this "demo." It was vastly too uncanny to be real, to the point that it was poorly setup. Google should be ashamed.