Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

201–210 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#201
Does it matter at all with regards to its AI capabilities though?

The video has a disclaimer that it was edited for latency.

And good speech-to-text and text-to-speech already exists, so building that part is trivial. There's no deception.

So then it seems like somebody is pressing a button to submit stills from a video feed, rather than live video. It's still just as useful.

My main question then is about the cup game, because that absolutely requires video. Does that mean the model takes short video inputs as well? I'm assuming so, and that it generates audio outputs for the music sections as well. If those things are not real, then I think there's a problem here. The Bloomberg article doesn't mention those, though.

Re: Gemini "duck" demo was not done in realtime or with voice

#202

I was fooled. The model release announcement said it could accept video and audio multi-modal input. I understood that there was a lot of editing and cutting, but I really believed I was looking at an example of video and audio input. I was completely impressed since it’s quite a leap to go from text and still images to “eyes and ears.” There’s even the segment where instruments are drown and music was generated. I t…

This kind of moral fraud - unethical behavior - is tolerated for some reason. It's almost like investors want to be fooled. There is no room for due diligence. They squeel like excited Taylor Swift fans as they are being lied to.

Re: Gemini "duck" demo was not done in realtime or with voice

#203

Earlier quoted context omitted.

> I'm unsure where this expectation of 100% absolute correctness comes from. It's a computer. That's why. Change the concept slightly: would you use a calculator if you had to wonder if the answer was correct or maybe it just made it up? Most people feel the same way about any computer based anything. I personally feel these inaccuracies/hallucinations/whatevs are only allowing them to be one rung up from practical j…

Speech to text is often wrong too. So is autocorrect. And object detection. Computers don't have to be 100% correct in order to be useful, as long as we don't put too much faith in them.

Call me old fashioned, but I would absolutely like to see autocorrect turned off in many contexts. I much prefer to read messages with 30% more transparent errors rather than any increase in opaque errors. I can tell what someone meant if I see "elephent in the room", but not "element in the room" (not an actual example, autocorrect would likely get that one right).

Re: Gemini "duck" demo was not done in realtime or with voice

#204
post #86

The whole Gemini webpage and contents felt weird to me, it's in the uncanny valley of trying to look and feel like an Apple marketing piece. The hyperbolic language, surgically precise ethnic/gender diversity, unnecessary animations and the sales pitch from the CEO felt like a small player in the field trying to pass as a big one.

It's funny because now the OpenAI keynote feels like it's emulating the Google keynotes from 5 years ago.

Google Keynote feels like it's emulating the Apple keynote from 5 years ago.

And the Apple keynote looks like robots just out of an uncanny valley pretending to be humans - just like keynotes might look in 5 years, but actually made by AI. Apple is always ahead of the curve in keynote trends.

Re: Gemini "duck" demo was not done in realtime or with voice

#205

Earlier quoted context omitted.

Okay, but search is done on a computer, and like the person you’re replying to said, we accept close enough. I don’t necessarily disagree with your interpretation, but there’s a revealed preference thing going on. The number of non-tech ppl I’ve heard directly reference ChatGPT now is absolutely shocking.

> The number of non-tech ppl I've heard directly reference ChatGPT now is absolutely shocking. The problem is that a lot of those people will take ChatGPT output at face value. They are wholly unaware that of its inaccuracies or that it hallucinates. I've seen it too many times in the relatively short amount of time that ChatGPT has been around.

So what? People do this with Facebook news too. That's a people problem, not an LLM problem.

Re: Gemini "duck" demo was not done in realtime or with voice

#206
post #146

Earlier quoted context omitted.

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

I know exactly where the expectation comes from. The whole world has demanded absolute precision from computers for decades. Of course, I agree that if we want computers to “think on their own“ or otherwise “be more human“ (whatever that means) we should expect a downgrade in correctness, because humans are wrong all the time.

> The whole world has demanded absolute precision from computers

The opposite. Far too tolerant of the excuse "sorry, computer mistake." (But yeah, just at the same time as "the computer says so".)

Re: Gemini "duck" demo was not done in realtime or with voice

#207

Earlier quoted context omitted.

I don't agree with the argument that "if a human can fail in this way, we should overlook this failing in our tooling as well." Because of course that's what LLMs are, tools, like any other piece of software. If a tool is broken, you seek to fix it. You don't just say "ah yeah it's a broken tool, but it's better than nothing!" All these LLM releases are amazing pieces of technology and the progress lately is incredib…

If a broken tool is useful, do you not use it because it is broken ? Overpowered LLMs like GPT-4 are both broken (according to how you are defining it) and useful -- they're just not the idealized version of the tool.

Maybe not if its the case that your use of the broken tool would result in the eventual undoing of your work. Like, lets say your staple gun is defective and doesn't shoot the staples deep enough, but it still shoots. You can keep using the gun, but it's not going to actually do its job. It seems useful and functional, but it isn't and its liable to create a much bigger mess.

Re: Gemini "duck" demo was not done in realtime or with voice

#208

Earlier quoted context omitted.

I'm a software engineer, and I more or less stopped asking ChatGPT for stuff that isn't mainstream. It just hallucinates answers and invents config file options or language constructs. Google will maybe not find it, or give you an occasional outdated result, but it rarely happens that it just finds stuff that's flat out wrong (in technology at least). For mainstream stuff on the other hand ChatGPT is great. And I'm s…

> it rarely happens that it just finds stuff that's flat out wrong "Flat out wrong" implies determinism. For answers which are deterministic such as "syntax checking" and "correctness of code" - this already happens. ChatGPT, for example, will write and execute code. If the code has an error or returns the wrong result it will try a different approach. This is in production today (I use the paid version).

Dollars to doughnuts says they are using GPT3.5.

Re: Gemini "duck" demo was not done in realtime or with voice

#209

Earlier quoted context omitted.

I'm an iOS user and prefer the swipe input implementation in GBoard over the one in the native keyboard. I'm not sure what the differences are, but GBoard just seems to overall make fewer mistakes and do a better job correcting itself from context.

As I was reading Andrew's comment to myself, I was trying to figure out when and why I stopped using swype typing on my phone. Then it hit me – I stopped after I switched from Android to iOS a few years ago. Something about the iOS implementation just doesn't feel right.

Apple's version is shit. Period. That's why.

Re: Gemini "duck" demo was not done in realtime or with voice

#210

Earlier quoted context omitted.

This seems to be a common view among some folks. Personally, I'm impartial. Search or even asking other expert human beings are prone to provide incorrect results. I'm unsure where this expectation of 100% absolute correctness comes from. I'm sure there are use cases, but I assume it's the vast minority and most can tolerate larger than expected inaccuracies.

I'm a software engineer, and I more or less stopped asking ChatGPT for stuff that isn't mainstream. It just hallucinates answers and invents config file options or language constructs. Google will maybe not find it, or give you an occasional outdated result, but it rarely happens that it just finds stuff that's flat out wrong (in technology at least). For mainstream stuff on the other hand ChatGPT is great. And I'm s…

The important thing is that with Web Search as a user you can learn to adapt to varying information quality. I have a higher trust for Wikipedia.org than I do for SEO-R-US.com, and Google gives me these options.

With a chatbot that's largely impossible, or at least impractical. I don't know where it's getting anything from - maybe it trained on a shitty Reddit post that's 100% wrong, but I have no way to tell.

There has been some work (see: Bard, Bing) where the LLM attempts to cite its sources, but even then that's of limited use. If I get a paragraph of text as an answer, is the expectation really that I crawl through each substring to determine their individual provenances and trustworthiness?

The shape of a product matters. Google as a linker introduces the ability to adapt to imperfect information quality, whereas a chatbot does not.

As an exemplar of this point - I don't trust when Google simply pulls answers from other sites and shows it in-line in the search results. I don't know if I should trust the source! At least there I can find out the source from a single click - with a chatbot that's largely impossible.

Post reply on HN