Live data from Hacker News

Gemini-2.5-pro-preview-06-05

deepmind.google

101–110 of 237 posts

Re: Gemini-2.5-pro-preview-06-05

#101
post #6

As if 3 different preview versions of the same model is not confusing enough, the last two dates are 05-06 and 06-05. They could have held off for a day:)

> the last two dates are 05-06 and 06-05 they are clearly trolling OpenAI's 4o and o4 models.

ChatGPT itself suggests better names than that!

Re: Gemini-2.5-pro-preview-06-05

#102
post #70

Is there a no brainer alternative to Claude Code where I can try other models?

People quite like aider! I’m not as much of a fan of the CLI workflow but it’s quite comparable, I think.

I've heard about it but is the outcome as good as claude code?

Re: Gemini-2.5-pro-preview-06-05

#103
post #12

Impressive seeing Google notch up another ~25 ELO on lmarena, on top of the previous #1, which was also Gemini! That being said, I'm starting to doubt the leaderboards as an accurate representation of model ability. While I do think Gemini is a good model, having used both Gemini and Claude Opus 4 extensively in the last couple of weeks I think Opus is in another league entirely. I've been dealing with a number of gn…

[deleted]

Re: Gemini-2.5-pro-preview-06-05

#104
post #6

As if 3 different preview versions of the same model is not confusing enough, the last two dates are 05-06 and 06-05. They could have held off for a day:)

> the last two dates are 05-06 and 06-05 they are clearly trolling OpenAI's 4o and o4 models.

Don't repeat the same mistake if you want to troll somebody.

It makes you look even more stupid.

Re: Gemini-2.5-pro-preview-06-05

#105
post #6

As if 3 different preview versions of the same model is not confusing enough, the last two dates are 05-06 and 06-05. They could have held off for a day:)

At what point will they move from Gemini 2.5 pro to Gemini 2.6 pro? I'd guess Gemini 3 will be a larger model.

Re: Gemini-2.5-pro-preview-06-05

#106

I have two issues with Gemini that I don't experience with Claude: 1. It RENAMES VARIABLE NAMES even in places where I don't tell it to change (I pass them just as context). and 2. Sometimes it's missing closing square brackets. Sure I'm a lazy bum, I call the variable "json" instead of "jsonStringForX", but it's contextual (within a closure or function), and I appreciate the feedback, but it makes reviewing the chan…

Gemini loves to add idiotic non-functional inline comments. "# Added this function" "# Changed this to fix the issue" No, I know, I was there! This is what commit messages for, not comments that are only relevant in one PR.

And it sure loves removing your carefully inserted comments for human readers.

Re: Gemini-2.5-pro-preview-06-05

#107
post #88

I have two issues with Gemini that I don't experience with Claude: 1. It RENAMES VARIABLE NAMES even in places where I don't tell it to change (I pass them just as context). and 2. Sometimes it's missing closing square brackets. Sure I'm a lazy bum, I call the variable "json" instead of "jsonStringForX", but it's contextual (within a closure or function), and I appreciate the feedback, but it makes reviewing the chan…

i've noticed with ChatGPT is will 100% ignore certain instructions and I wonder if it's just an LLM thing. For example, I can scream and yell in caps at ChatGPT to not use em or en dashes and if anything it makes it use them even more . I've literally never once made it successfully not use them, even when it ignored it the first time, and my follow up is "output the same thing again but NO EM or EN DASHES!" i've not…

There are some things so ubiquitous in the training data that it is really difficult to tell models to not so them. Simply because it is so ingrained in their core training. Em dashes are apparently one of those things.

It's something I read a lottle while ago in a larger article but can't remember which article it was.

Re: Gemini-2.5-pro-preview-06-05

#109
post #12

Impressive seeing Google notch up another ~25 ELO on lmarena, on top of the previous #1, which was also Gemini! That being said, I'm starting to doubt the leaderboards as an accurate representation of model ability. While I do think Gemini is a good model, having used both Gemini and Claude Opus 4 extensively in the last couple of weeks I think Opus is in another league entirely. I've been dealing with a number of gn…

I think the only way to be particularly impressed with new leading models lately is to hold the opinion all of the benchmarks are inaccurate and/or irrelevant and it's vibes/anecdotes where the model is really light years ahead. Otherwise you look at the numbers on e.g. lmarena and see it's claiming a ~16% preference win rate for gpt-3.5-turbo from November of 2023 over this new world-leading model from Google.

Re: Gemini-2.5-pro-preview-06-05

#110
post #68

Earlier quoted context omitted.

OpenAI has already forecast $12B in revenue by the end of this year. I agree that Google is well-positioned, but the mindshare/product advantage OpenAI has gives them a stupendous amount of leeway

The hurdle for OpenAI is going to be on the profit side. Google has their own hardware acceleration and their own data centers. OpenAI has to pay a monopolist for hardware acceleration and beholden to another tech giant for data centers. Never mind that Google can customize it's hardware specifically for it's models. The only way for OpenAI to really get ahead on solid ground is to discover some sort of absolute game…

OpenAI has now partnered with Jony Ive now and they are going to have thinnest data centers with thinnest servers mounted on thinnest racks. And since everything is so thin, servers can just whisper to each other instead of communicating via fat cables.

I think that will be the game changer OpenAI will show us soon.

Post reply on HN