Live data from Hacker News

Gemini-2.5-pro-preview-06-05

deepmind.google

231–237 of 237 posts

Re: Gemini-2.5-pro-preview-06-05

#231

Earlier quoted context omitted.

We are trading tokens and mental health for time? I used Cursor well over a year ago. It gave me a headache. It was very immature. Used cursor more recently: the headache intensity increased. It's not cursor it is the senseless loops hoping for the LLM to spit out something somewhat correct. Revisiting the prompt. Trying to become an elite in language protocols because we need that machine to understand us. Leaving a…

> We are trading tokens and mental health for time? I have bipolar disorder. This makes programming incredibly difficult for me at times. Almost all the recent improvements to code generation tooling have been a tremendous boon for me. Coding is now no longer this test of how frustrated I can get over the most trivial of tasks. I just ask for what I want precisely and treat responses like a GitHub PR where mistakes m…

If your handicap make coding difficult perhaps another profession would suit you better.

Now if Ai assistance allow you to perform well then that is a different story and I take my advice back of course.

There is a lot to say, positive things about how LLMs enables people to perform at tasks that would be impossible for them. Whether due to handicaps or simply lacking the abilities, or opportunity to train.

My comment was on the impact on "healthy" individuals who remain the majority of the population. And I only spoke for myself, I have no clue maybe it is just me or due to how I use the thing. Thanks for sharing your experience though, I had not considered what might be a concern for the majority with this might very well be an enabler.

Re: Gemini-2.5-pro-preview-06-05

#232

Earlier quoted context omitted.

Yeah, exactly. For everyone who might not know, the chat apps add lots of complex system prompting to handle and shape personality, tone, general usability, etc. IDE's also do this (with Claude Code being one of the ones that are closest to "bare" model that you can get) but at they are at least guiding it's behavior to be really good at coding tasks. Another reason is using the Agent feature that IDE's have had for…

IDE's are intimidating to non-tech people. I'm surprised there isn't a VibeIDE yet that is purpose build to make it possible for your grandmother to execute code output by an LLM.

https://firebase.studio comes to mind

Re: Gemini-2.5-pro-preview-06-05

#233
post #126

Earlier quoted context omitted.

I think the only way to be particularly impressed with new leading models lately is to hold the opinion all of the benchmarks are inaccurate and/or irrelevant and it's vibes/anecdotes where the model is really light years ahead. Otherwise you look at the numbers on e.g. lmarena and see it's claiming a ~16% preference win rate for gpt-3.5-turbo from November of 2023 over this new world-leading model from Google.

Not sure I follow - Gemini has ELO 1470, GPT3.5-turbo is 1206, which is an 86% win rate. https://chatgpt.com/share/6841f69d-b2ec-800c-9f8c-3e802ebbc0...

gpt-3.5-turbo-1106 from November 2023 was 1170, 1206 is for the March variant.

Change that and you get ~84%, flip the order (i.e. the win rate of GPT-3.5 is ~16%). I.e. the point is a two year old model still wins far too often to be excited about each new top model for the last two years, not that the two year old model is better.

Re: Gemini-2.5-pro-preview-06-05

#234
post #70

Earlier quoted context omitted.

People quite like aider! I’m not as much of a fan of the CLI workflow but it’s quite comparable, I think.

I've heard about it but is the outcome as good as claude code?

In short - it's like comparing Ubuntu with MacOS/Windows

Open source power vs corp - if you are eager to do some stuff yourself, aider is probably a better pick (or even a strategic investment); otherwise you'll simply find much more tools that don't require setting them up in Cline/Claude Code

But man, if you're on ycombinator.com, how not to stick with open source?

Re: Gemini-2.5-pro-preview-06-05

#235

Earlier quoted context omitted.

"People have no idea what claude or gemini are" One well-placed ad campaign could easily change all that. Doesn't hurt that Google can bundle Gemini into Android.

If it were that simple to sway markets through marketing, we would see Pepsi/Coca-Cola or McDonalds/BurgerKing swing like crazy all the time from "one well-placed ad campaign" to the next. We do not.

Thanks to well-placed ad campaigns, people have a very good idea what Pepsi, Coca-Cola, McDonald's and Burger King are. They also know what Siri is. And it would be similarly easy to establish Gemini as a household name.

Re: Gemini-2.5-pro-preview-06-05

#237

Earlier quoted context omitted.

Engineers are surprisingly bad at naming things!

I rather like date codes as versions.

But it's not clear how to interpret the date code: 05-06 could be 5th June or 6th May; same sorry for 06-05. Very confusing due to American-style date formatting. Versions number are at least sequential, with a bigger number being a later version.
Post reply on HN