Live data from Hacker News

Gemini 2.0: our new AI model for the agentic era

blog.google

401–410 of 512 posts

Re: Gemini 2.0: our new AI model for the agentic era

#403
Gemini 1120, 1206, and Gemini 2.0 flash have better coding results than ChatGPT o1 and Claude Sonnet 3.5.

They did it: from now on Google will keep a leadership position.

They have too much data (Search, Maps, Youtube, Chrome, Android, Gmail, etc.), and they have their own servers (it's free!) and now the Willow QPU.

To me, it is evident how the future will look. I'll buy some more Alphabet stocks

Re: Gemini 2.0: our new AI model for the agentic era

#404

Earlier quoted context omitted.

Majority of people want better performance, running locally is just a nice to have feature.

Latency is a huge factor in performance, and local models often have a huge edge. Especially on mobile devices that could be offline entirely.

Definitely not when it comes to LLM's, the larger more useful local models are not that fast and latency is not an issue, just look at this Google models voice function or even openai's advanced voice.

Re: Gemini 2.0: our new AI model for the agentic era

#405

Earlier quoted context omitted.

> they also have an incredibly bad track record of supporting their products Incredibly bad track record of supporting products that don't grow . I'm not saying this to defend Google, I'm still (perhaps unreasonably) angry because of Reader, it's just that there is a pattern and AI isn't likely to fit that for a long while.

> products that don't grow. I think we all acknowledge this. The question is seldom "why" they kill it (I'd argue ultimately it doesn't matter), it's about how fast and what they offer as a migration path for those who boarded the train. That also means the minute Gemini stops looking like a growing product it's gone from this world, where Microsoft backed alternatives have a fighting chance to get some leeway to rec…

Yeah, MS Azure DevOps is still alive, though stagnant. I thought everyone would be moved to GitHub in few years after MS acquired GitHub. Yet here we are, 6 years later.

Re: Gemini 2.0: our new AI model for the agentic era

#407

Earlier quoted context omitted.

Leaderboards are not that useful for measuring real-life effectiveness of the models atleast in my day-today usage. I am currently struggling to diagnose an ipv6 mis-configuration in my enormous aws cloudformation yaml code. I gave the same input to Claude Opus, Gemini and ChatGPT ( o1 and 4o). 4o was the worst. verbose and waste of my time. Claude completely went off-tangent and began recommending fixes for ipv4 whi…

Sonnet 3.5 as of today is superior to Opus, curious if sonnet could have solved your problem

Yes. I find it a bit funny how much people care about leaderboards. I see models going up and down, winning this or that benchmark and yet, for me, Sonnet 3.5 still beats the crap out of all of them.

Re: Gemini 2.0: our new AI model for the agentic era

#408

OT: I’m not entirely sure why, but "agentic" sets my teeth on edge. I don't mind the concept, but the word itself has that hollow, buzzwordy flavor I associate with overblown LinkedIn jargon, particularly as it is not actually in the dictionary...unlike perfectly serviceable entries such as "versatile", "multifaceted" or "autonomous"

[dead]

Re: Gemini 2.0: our new AI model for the agentic era

#409

The Gemini 2 models support native audio and image generation but the latter won't be generally available till January. Really excited for that as well as 4o's image generation (whenever that comes out). Steerability has lagged behind aesthetics in image generation for a while now and it's be great to see a big advance in that. Also a whole lot of computer vision tasks (via LLMs) could be unlocked with this. Think In…

I asked Gemini 2.0 Flash (with my voice) whether it natively understands audio or is converting my voice to text. It replied:

"That's an insightful question. My understanding of your speech involves a pipeline first. Your voice is converted to text and then I process the text to understand what you're saying. So I don't understand your voice directly but rather through a text representation of it."

Unsure if this is a hallucination, but is disappointing if true.

Edit: Looking at the video you linked, they say "native audio output", so I assume this means the input isn't native? :(

Post reply on HN