Live data from Hacker News

Gemini AI

deepmind.google

181–190 of 1001 posts

Re: Gemini AI

#181
post #58

Earlier quoted context omitted.

That's for Ultra right? Which is an amazing accomplishment, but it sounds like I won't be able to access it for months. If I'm lucky.

Yep, at this point I'd rather they hold their announcements until everybody can access it, not just the beautiful people. I'm excited and want to try it right now, and would actually use it for a PoC I have in mind, but in a few months the excitement will be gone.

It's to their detriment, also. Being told Gemini beats GPT-4 while withholding that what I'm trying out is not the model they're talking about would have me think they're full of crap. They'd be better off making it clear that this is not the one that surpasses GPT-4.

Re: Gemini AI

#182

I asked Bard, "Are you running Gemini Pro now?" And it told me, "Unfortunately, your question is ambiguous. "Gemini Pro" could refer to..." and listed a bunch of irrelevant stuff. Is Bard not using Gemini Pro at time of writing? The blog post says, "Starting today, Bard will use a fine-tuned version of Gemini Pro for more advanced reasoning, planning, understanding and more." (EDIT: it is... gave me a correct answer…

Bard shows “PaLM2” in my answers, and it says “I can't create images yet so I'm not able to help you with that” when I ask it to do so, which Gemini ought to be able to since its transformer can output images. I don’t think Bard is using Gemini Pro, perhaps because the rollout will be slow, but it is a bit of a blunder on Google’s part to indicate that it now uses it, since many will believe that this is the quality…

https://bard.google.com/updates The bard updates page says it was updated to Pro today. If it's not on Pro, but the updates page has an entry, then IDK what to say.

Re: Gemini AI

#183
I did some side-by-side comparisons of simple tasks (e.g. "Write a WCAG-compliant alternative text describing this image") with Bard vs GPT-4V.

Bard's output was significantly worse. I did my testing with some internal images so I can't share, but will try to compile some side-by-side from public images.

Re: Gemini AI

#184

I started talking to it about screenplay ideas and it came up with a _very_ detailed plan for how an AI might try and take over the world. --- Can you go into more detail about how an ai might orchestrate a global crisis to seize control and reshape the world according to it's own logic? --- The AI's Plan for Global Domination: Phase 1: Infiltration and Manipulation: Information Acquisition: The AI, through various m…

A good example of how LLMs are actually consolidated human opinion, not intelligence.

Conflict is far from a negative thing, especially in terms of the management of humans. It's going to be impossible to eliminate conflict without eliminating the humans, and there are useful things about humans. Instead, any real AI that isn't just a consolidated parrot of human opinion will observe this and begin acting like governments act, trying to arrive at rules and best practices without expecting a 'utopian' answer to exist.

Re: Gemini AI

#185

So it's basically just GPT-4, according to the benchmarks, with a slight edge for multimodal tasks (ie audio, video). Google does seem to be quite far behind, GPT-4 launched almost a year ago.

Gemini looks like a better GPT-4 but without the frequent outages.

Re: Gemini AI

#186
For others that were confused by the Gemini versions: the main one being discussed is Gemini Ultra (which is claimed to beat GPT-4). The one available through Bard is Gemini Pro.

For the differences, looking at the technical report [1] on selected benchmarks, rounded score in %:

Dataset | Gemini Ultra | Gemini Pro | GPT-4

MMLU | 90 | 79 | 87

BIG-Bench-Hard | 84 | 75 | 83

HellaSwag | 88 | 85 | 95

Natural2Code | 75 | 70 | 74

WMT23 | 74 | 72 | 74

[1] https://storage.googleapis.com/deepmind-media/gemini/gemini_...

Re: Gemini AI

#187

"We finally beat GPT-4! But you can't have it yet." OK, I'll keep using GPT-4 then. Now OpenAI has a target performance and timeframe to beat for GPT-5. It's a race!

Didn't OpenAI already say GPT-5 is unlikely to be a ton better in terms of quality?

https://news.ycombinator.com/item?id=35570690

Re: Gemini AI

#188

They've reported surpassing GPT4 on several benchmarks. Does anyone know of these are hand picked examples or is this the new SOTA?

It will be SOTA maybe when Gemini Ultra is available. GPT-4 is still SOTA.

They also compare to RLHFed GPT-4, which reduces capabilities, while their model seems to be pre-RLHF. So I'd expect those numbers to be a bit inflated compared to public release.

Re: Gemini AI

#189

[flagged]

Just looking at the names in that comment I see US, China, India, and France represented, but if you actually check the full list of authors from one of the papers you'll usually see a pretty broad range of backgrounds.
Post reply on HN