Live data from Hacker News

Gemini 2.0: our new AI model for the agentic era

blog.google

321–330 of 512 posts

Re: Gemini 2.0: our new AI model for the agentic era

#321

Earlier quoted context omitted.

The fact that they're using Gemini with even their most important products shows that they trust it.

Again, that's covered by "all our products". Why do we need to be reminded that Google has a lot of users? Someone oblivious to that isn't going to care about this press release.

Scale and cost are defining considerations of LLMs. By saying they're rolling out to billions of users, they're pointing out they're doing something pretty unique and have confidence in a major competitive advantage. Point billions of devices at other high-performing competitors' offerings, and all of them would fall over.

Re: Gemini 2.0: our new AI model for the agentic era

#322

Gemini-2.0-Flash does extremely well on the Hallucination Evaluation Leaderboard, at 1.3% hallucination rate https://github.com/vectara/hallucination-leaderboard

Fascinating, thanks for calling that out: I found 1.0 promising in practice, but with hallucination problems. Then I saw it had gotten 57% of questions wrong on open book true/false and I wrote it off completely - no reason to switch to it for speed and cost if it's just a random generator. That's a great outcome.

Re: Gemini 2.0: our new AI model for the agentic era

#323

Earlier quoted context omitted.

I've started keeping an eye out for original brainteasers, just for that reason. GCHQ's Christmas puzzle just came out [1], and o1-pro got 6 out of 7 of them right. It took about 20 minutes in total. I wasn't going to bother trying those because I was pretty sure it wouldn't get any of them, but decided to give it an easy one (#4) and was impressed at the CoT. Meanwhile, Google's newest 2.0 Flash model went 0 for 7.…

Why are you comparing flash vs o1-pro, wouldn't a more fair comparison be flash vs mini?

It's the only Google model that my account has access to that accepts .PNG files. I assumed it was the latest/greatest experimental 2.0 release.

If they want a rematch, they'll need to bring their 'A' game next time, because o1-pro is crazy good.

Re: Gemini 2.0: our new AI model for the agentic era

#324
post #270

Earlier quoted context omitted.

I've started keeping an eye out for original brainteasers, just for that reason. GCHQ's Christmas puzzle just came out [1], and o1-pro got 6 out of 7 of them right. It took about 20 minutes in total. I wasn't going to bother trying those because I was pretty sure it wouldn't get any of them, but decided to give it an easy one (#4) and was impressed at the CoT. Meanwhile, Google's newest 2.0 Flash model went 0 for 7.…

Did it get the 8 right? The linked article provides the wrong answer btw.

I didn't see a straightforward way to submit the final problem, because I used different contexts for each of the 7 subproblems.

Given the right prompt, though, I'm sure it could handle the 'find the corresponding letter from the landmarks to form an anagram' part. That's easier than most of the other problems.

You're saying the ultimate answer isn't 'PROTECTING THE UNITED KINGDOM'?

Re: Gemini 2.0: our new AI model for the agentic era

#326

I know this isn't really a useful comment, but, I'm still sour about the name they chose. They MUST have known about the Gemini protocol. I'm tempted to think it was intentional, even. It's like Microsoft creating an AI tool and calling it Peertube. "Hurr durr they couldn't possibly be confused; one is a decentralised video platform and the other is an AI tool hurr durr. And ours is already more popular if you 'bing'…

I'd never heard of this protocol, and I try to keep up so I can't blame Google.

Re: Gemini 2.0: our new AI model for the agentic era

#327

Earlier quoted context omitted.

> Remains to be seen how well they will be able to productize and market The challenge is trust. Google is one of the leaders in AI and are home to incredibly talented developers. But they also have an incredibly bad track record of supporting their products. It's hard to justify committing developers and money to a product when there's a good chance you'll just have to pivot again once they get bored. Say what you w…

Putting your trust in Google is a fools errand. I don't know anyone that doesn't have a story.

Google has 4 Billion users. It's delusional to think that you don't know anyone or you live in an incredibly small bubble

Re: Gemini 2.0: our new AI model for the agentic era

#328
Agents are the worst idea. I think AI will start to progress again when we get better models that drop this whole chat idea, and just focus on completions. Building the tools on top should be the work that is given over to the masses. It's sad that instead the most powerful models have been hamstrung.

Re: Gemini 2.0: our new AI model for the agentic era

#329

Earlier quoted context omitted.

Again, that's covered by "all our products". Why do we need to be reminded that Google has a lot of users? Someone oblivious to that isn't going to care about this press release.

Scale and cost are defining considerations of LLMs. By saying they're rolling out to billions of users, they're pointing out they're doing something pretty unique and have confidence in a major competitive advantage. Point billions of devices at other high-performing competitors' offerings, and all of them would fall over.

That's not what the sentence says. "Now millions of developers are building with Gemini. And it’s helping us reimagine all of our products — including all 7 of them with 2 billion users — and to create new ones." That does not imply that Gemini will start receiving requests from billions of users. At best it says that they'll start using it in some unspecified way.

Re: Gemini 2.0: our new AI model for the agentic era

#330

Big companies can be slow to pivot, and Google has been famously bad at getting people aligned and driving in one direction. But, once they do get moving in the right direction the can achieve things that smaller companies can't. Google has an insane amount of talent in this space, and seems to be getting the right results from that now. Remains to be seen how well they will be able to productize and market, but hard…

> but hard to deny that their LLM models aren't really, really good though

Although I do still pay for ChatGPT, I find it dog slow. ChatGPT is simply way too slow to generate answers. It feels like --even though of course it's not doing the same thing-- I'm back to the 80s with my 8-bit computer printing thing line by line.

Gemini OTOH doesn't feel like that: answers are super fast.

To me low latency is going to be the killer feature. People won't keep paying for models that are dog slow to answer.

I'll probably be cancelling my ChatGPT subscription soon.

Post reply on HN