Live data from Hacker News

Gemini AI

deepmind.google

161–170 of 1001 posts

Re: Gemini AI

#162

There seems to be a small error in the reported results: In most rows the model that did better is highlighted, but in the row reporting results for the FLEURS test, it is the losing model (Gemini, which scored 7.6% while GPT4-v scored 17.6%) that is highlighted.

That row says lower is better. For "word error rate", lower is definitely better.

But they also used Large-v3, which I have not ever seen outperform Large-v2 in even a single case. I have no idea why OpenAI even released Large-v3.

Re: Gemini AI

#163

So just a bunch of marketing fluff? I can use GPT4 literally right now and it’s apparently within a few percentage points of what Gemini Ultra can do… which has no release date as far as I can tell. Would’ve loved something more substantive than a bunch of videos promising how revolutionary it is.

[flagged]

Re: Gemini AI

#164

There seems to be a small error in the reported results: In most rows the model that did better is highlighted, but in the row reporting results for the FLEURS test, it is the losing model (Gemini, which scored 7.6% while GPT4-v scored 17.6%) that is highlighted.

The text beside it says "Automatic speech recognition (based on word error rate, lower is better)"

Re: Gemini AI

#165

Apple lost the PC battle, MS lost the mobile battle, Google is losing the AI battle. You can't win everywhere.

Beautifully said. So basically: Apple lost the PC battle and won mobile, Microsoft lost the mobile battle and (seemingly) is winning AI, Google is losing the AI battle, but will win .... the Metaverse? Immersive VR? Robotics?

Maybe Google skips the LLM era and wins the AGI race?

Re: Gemini AI

#166

I asked Bard, "Are you running Gemini Pro now?" And it told me, "Unfortunately, your question is ambiguous. "Gemini Pro" could refer to..." and listed a bunch of irrelevant stuff. Is Bard not using Gemini Pro at time of writing? The blog post says, "Starting today, Bard will use a fine-tuned version of Gemini Pro for more advanced reasoning, planning, understanding and more." (EDIT: it is... gave me a correct answer…

Bard shows “PaLM2” in my answers, and it says “I can't create images yet so I'm not able to help you with that” when I ask it to do so, which Gemini ought to be able to since its transformer can output images.

I don’t think Bard is using Gemini Pro, perhaps because the rollout will be slow, but it is a bit of a blunder on Google’s part to indicate that it now uses it, since many will believe that this is the quality that Gemini assumes.

Re: Gemini AI

#167
Competition is good. Glad to see they are catching up with GPT4, especially with a lot of commentary expecting a plateau in Transformers.

Re: Gemini AI

#168

Interesting that they're announcing Ultra many months in advance of the actual public release. Isn't that just giving OpenAI a timeline for when they need to release GPT5? Google aren't going to gain much market share from a model competitive with GPT4 if GPT5 is already available.

I don't think there are a lot of surprises on either side about what's coming next. Most of this is really about pacifying shareholders (on Google's side) who are no doubt starting to wonder if they are going to fight back at all.

With either OpenAI and Google, or even Microsoft, the mid term issue is as much going to be about usability and deeper integration than it is about model fidelity. Chat gpt 4 turbo is pretty nice but the UI/UX is clumsy. It's not really integrated into anything and you have to spoon feed it a lot of detail for it to be useful. Microsoft is promising that via office integration of course but they haven't really delivered much yet. Same with Google.

The next milestone in terms of UX for AIs is probably some kind of glorified AI secretary that is fully up to speed on your email, calendar, documents, and other online tools. Such an AI secretary can then start adding value in terms of suggesting/completing things when prompted, orchestrating meeting timeslots, replying to people on your behalf, digging through the information to answer questions, summarizing things for you, working out notes into reports, drawing your attention to things that need it, etc. I.e. all the things a good human secretary would do for you that free you up to do more urgent things. Most of that work is not super hard it just requires enough context to understand things.

This does not even require any AGIs or fancy improvements. Even with chat gpt 3.5 and a better ux, you'd probably be able to do something decent. It does require product innovation. And neither MS nor Google is very good at disruptive new products at this point. It takes them a long time and they have a certain fail of failure that is preventing them from moving quickly.

Re: Gemini AI

#169

The performance results here are interesting. G-Ultra seems to meet or exceed GPT4V on all text benchmark tasks with the exception of Hellaswag where there's a significant lag, 87.8% vs 95.3%, respectively.

I wonder how that weird HellaSwag lag is possible. Is there something really special about that benchmark?

Tech report seems to hint at the fact that GPT-4 may have had some training/testing data contamination and so GPT-4 performance may be overstated.

Re: Gemini AI

#170

So it's basically just GPT-4, according to the benchmarks, with a slight edge for multimodal tasks (ie audio, video). Google does seem to be quite far behind, GPT-4 launched almost a year ago.

Less than a year difference is "quite far behind"?

Lotus 1-2-3 came out 4 years before Microsoft Excel. WordPerfect came out 4 years before Microsoft Word.

Hotmail launched 8 years before Gmail. Yahoo! Mail was 7 years before Gmail.

Heck, AltaVista launched 3 years before Google Search.

I don't think less than a year difference is meaningful at all in the big picture.

Post reply on HN