Live data from Hacker News

Gemini "duck" demo was not done in realtime or with voice

twitter.com

371–380 of 683 posts

Re: Gemini "duck" demo was not done in realtime or with voice

#371
post #293

Earlier quoted context omitted.

Microsoft eating Google's lunch on documents is laughable at best. Not to mention it confuses the entire timeline of office productivity software??

Is paid MS Teams is more or less common than paid GSuite? It's hard to find stats on this. GSuite is the better product IMO, but MS has a stronger b2b reputation, and anecdotally I hear more about people using Teams.

Nobody pays for Teams, but everyone pays for Office, and if you get Teams for free with it ...

Re: Gemini "duck" demo was not done in realtime or with voice

#372

Earlier quoted context omitted.

Is paid MS Teams is more or less common than paid GSuite? It's hard to find stats on this. GSuite is the better product IMO, but MS has a stronger b2b reputation, and anecdotally I hear more about people using Teams.

Does anyone use paid GSuite for anything other than docs/drive/Gmail ? In all companies I've worked at, we've used GSuite exclusively for those, and used slack/discord for chat, and zoom/discord for video/meetings. I know that MS Teams is a more full-featured product suite, but even at companies that used it, we still used Zoom for meetings.

My company uses Meet. It works great! I like it more than Zoom.

Re: Gemini "duck" demo was not done in realtime or with voice

#373

A big red flag for me was that Sundar was prompting the model to report lots of facts that can be either true or false. We all saw the benchmark figures that they published and the results mostly showed marginal improvements. In other words, the issue of hallucination has not been solved. But the demo seemed to imply that it had. My conclusion was that they had mostly cherry picked instances in which the model happen…

These LLMs do not have a concept of factual correctness and are not trained/optimized as such. I find it laughable that people expect these things to act like quiz bots - this misunderstands the nature of a generative LLM entirely. It simply spits out whatever output sequence it feels is most likely to occur after your input sequence. How it defines “most likely” is the subject of much research, but to optimize for f…

> It simply spits out whatever output sequence it feels is most likely to occur after your input sequence... but to optimize for factual correctness is a completely different endeavor

What if the input sequence says "the following is truth:", assuming it skillfully predicts following text, it would mean telling the most likely truth according to its training data.

Re: Gemini "duck" demo was not done in realtime or with voice

#374

That's not the only thing wrong. Gemini makes a false statement in the video, serving as a great demonstration of how these models still outright lie so frequently, so casually, and so convincingly that you won't notice, even if you have a whole team of researchers and video editors reviewing the output. It's the single biggest problem with LLMs and Gemini isn't solving it. You simply can't rely on them when correctn…

After asserting it's a rubber duck, there are some claims without follow-up:

- Just after that it doesn't translate the "rubber" part

- It states there's no land nearby for it to rest or find food in the middle of the ocean: if it's a rubber duck it doesn't need to rest nor feed. (That's a missed opportunity to mention the infamous "Friendly Floatees spill"[1] in 1992 as some rubber ducks floated to that map position). Although it seems to recognize geographical features of the map, it fails to mention Easter Island is relatively nearby. And if it were recognized as a simple duck — which it described as a bird swimming in the water — it seems oblivious to the fact that the duck might feed itself in the water. It doesn't mention either that the size of the duck seems abnormally big in that map context.

- The concept of friends and foes doesn't apply to a rubber duck either. Btw labeling the duck picture as a friend and the bear picture as a foe seems arbitrary (e.g. a real duck can be very aggressive even with other ducks.)

Among other things, the astronomical riddle seems also flawed to me: it answered "The correct order is Sun, Earth, Saturn".

I'd like for it to state :

- the premises it used, like "Assuming it depicts the Sun, Saturn and the Earth" (there are other stars, other ringed-planets, and the Earth similarity seems debatable)

- the sorting criteria it used (e.g. using another sorting key like the average distance from us "Earth, Sun, Saturn" can be a correct order)

[1] https://en.wikipedia.org/wiki/Friendly_Floatees_spill

Re: Gemini "duck" demo was not done in realtime or with voice

#375

Earlier quoted context omitted.

>Google Docs created in 2006 tech was based on an acquired company, Google just abused their search monopoly to make it more popular(same thing they did with YT). This has been the strategy for every service they've ever made, Google really hasn't launched a decent in-house product since Gmail and even that was grown using their search monopoly as free advertising >Google Docs originated from Writely, a web-based wor…

> Google really hasn't launched a decent in-house product since Gmail What about Chrome? And Chromebooks?

Sorry if this was a joke and I didn't spot it. Chrome was based on WebKit which was itself based on KHTML if memory serves. Chromebooks are based on a version of that outside engine running on top of Linux which they also didn't create.

Re: Gemini "duck" demo was not done in realtime or with voice

#377

Earlier quoted context omitted.

These LLMs do not have a concept of factual correctness and are not trained/optimized as such. I find it laughable that people expect these things to act like quiz bots - this misunderstands the nature of a generative LLM entirely. It simply spits out whatever output sequence it feels is most likely to occur after your input sequence. How it defines “most likely” is the subject of much research, but to optimize for f…

The first question I always ask myself in such cases: how much input data has a simple "I don't know" lines? This is clearly a concept (not knowing sth) that has to be learned in order to be expressed in the output.

What stops you from asking the same question multiple times, and seeing if the answers are consistent. I am sure the capital of France is always going to come out Paris, but the name of a river passing a small village might be hallucinated differently. Even better - use two different models, if they agree it's probably true. And probably the best - provide the data to the model in context, if you have a good source. Don't use the model as fact knowledge base, use RAG.

Re: Gemini "duck" demo was not done in realtime or with voice

#378
post #86

The whole Gemini webpage and contents felt weird to me, it's in the uncanny valley of trying to look and feel like an Apple marketing piece. The hyperbolic language, surgically precise ethnic/gender diversity, unnecessary animations and the sales pitch from the CEO felt like a small player in the field trying to pass as a big one.

> surgically precise ethnic/gender diversity What does that mean and why is it bad? Diversity in marketing is used because, well, your desired market is diverse. I don't know what it means for it to be surgically precise, though.

It's bad if the makeup of the company doesn't reflect the diversity seen in the marketing, because it doesn't reflect any genuine value and is just for show.

Now, I don't know how diverse the AI workforce is at Google, but the YT thumbnails show precisely 50% of white men. Maybe that's what the parent meant by "surgically precise".

Re: Gemini "duck" demo was not done in realtime or with voice

#379

Earlier quoted context omitted.

The voice interaction part didn't look a far cry from what we are doing with Dynamic Interaction at SoundHound. Because of this I assumed (like many it seems) that they had caught up. And it's dangerous to assume they can just "deliver later". It's not that simple. If it is why not bake it in right now instead of committing fraud? This is damaging to companies that walk the walk and then people have literally said to…

I feel that more than you realize That was basically what magic leap did to the whole AR development market. Everyone deep in it knew they couldn’t do it but they messed up so badly that it basically killed the entire industry

So let's not give big tech benefit of the doubt on this one. We have to call them out but even then the lie is already half way around the world...

Re: Gemini "duck" demo was not done in realtime or with voice

#380

A big red flag for me was that Sundar was prompting the model to report lots of facts that can be either true or false. We all saw the benchmark figures that they published and the results mostly showed marginal improvements. In other words, the issue of hallucination has not been solved. But the demo seemed to imply that it had. My conclusion was that they had mostly cherry picked instances in which the model happen…

The issue of hallucinations won't be solved with the RAG approach. It requires a fundamentally different architecture. These aren't my words but Yann LeCun's. You could easily understand if you spend some time playing around. The autoregressive nature won't allow the LLMs to create an internally consistent model before answering the question. We have approaches like Chain of Thought and others, but they are merely band-aids and superficially address the issue.
Post reply on HN