Live data from Hacker News

Gemini 2.5 Deep Think

blog.google

151–160 of 259 posts

Re: Gemini 2.5 Deep Think

#151

Earlier quoted context omitted.

Mainframes are the only viable way to build computers. Micro processors will never figure out how to get small and fast enough for personal computers to reach escape velocity.

Why do you think the analogy hold?

Hardware typically gets faster and cheaper over time. Unless we hit hard a wall because of physics then I don't see any reason that won't continue to be true.

Re: Gemini 2.5 Deep Think

#152

Earlier quoted context omitted.

it turns out that AI at this level is very expensive to run (capex, energy). my bet is that AI itself won't figure out how to overcome these constraints and reach escape velocity.

our minds are incredibly energy efficient, that leads me to believe it is possible to figure out, but it might be a human rather than an AI that gives us something more akin to a biological solution.

This could fix my main gripe with The Matrix. ”Humans are used as batteries” always felt off, but it totally would make sense if the human brains have uniquely energy efficient pattern matching abilities that an emerging AI organism would harvest. That would also strengthen the spiritual humanist subtext.

Re: Gemini 2.5 Deep Think

#153
post #57

Been using Gemini for a few months, somehow it's gotten much, much worse in that time. Hallucinations are very common, and it will argue with you when you point it out. So, don't have much confidence.

Is the problem mainly with tool use ? and are you using it through AI studio or through the API ?.

I've found that it hallucinates tool use for tools that aren't available and then gets very confident about the results.

Re: Gemini 2.5 Deep Think

#154

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

Several years ago I thought a good litmus test for mastery of coding is not finding a solution using internet search nor getting well written questions about esoteric coding problems answered on StackOverflow. For a while, I would post a question and answer my own question after I solved the problem for posterity (or AI bots). I always loved getting the "I've been working on this for 3 days and you saved my life" com…

They're remarkably useless on stuff they've seen but not had up-weighted in the training set. Even the best ones (Opus 4 running hot, Qwen and K2 will surprise you fairly often) are a net liability in some obscure thing.

Probably the starkest example of this is build system stuff: it's really obvious which ones have seen a bunch of `nixpkgs`, and even the best ones seem to really struggle with Bazel and sometimes CMake!

The absolute prestige high-end ones running flat out burning 100+ dollars a day and it's a lift on pre-SEO Google/SO I think... but it's not like a blowout vs. a working search index. Back when all the source, all the docs, and all the troubleshooting for any topic on the whole Internet were all above the fold on Google? It was kinda like this: type a question in the magic box and working-ish code pops out. Same at a glory-days FAANG with the internal mega-grep.

I think there's a whole cohort or two who think that "type in the magic box and code comes out" is new. It's not new, we just didn't have it for 5-10 years.

Re: Gemini 2.5 Deep Think

#155

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

The part I cannot understand is why for many AI offerings, I cannot make out what each pricing tier does with a quick glance.

What happened to the simplicity of Steve Jobs' 2x2 (consumer vs.pro, laptop vs. desktop)?

Re: Gemini 2.5 Deep Think

#156

Earlier quoted context omitted.

No, you're talking about costs to user, which are oversimplifications of the costs that providers bear. One output token with a million input tokens is incredibly cheap for providers

> One output token with a million input tokens is incredibly cheap for providers Source? Afaik this is incorrect.

Chevk out any LLM API providers pricing. Output tokens are always significantly more expensive than input (which can also be cached).

Re: Gemini 2.5 Deep Think

#157
post #124

Earlier quoted context omitted.

It's not particularly interesting if Deep Mind comes to the same (correct) conclusion on a single problem as o3 but costs more. You could ask gpt 2.5 and gpt4 what 1+1= and would get the same response with gpt 4 costing more, but this doesn't tell us much about model capability or value. It would be more interesting to know if it can handle problems that o3 can't do, or if it is 'correct' more often than o3 pro on th…

> It would be more interesting to know if it can handle problems that o3 can't do Suppose it can't. How will you know? All the datapoints will be "not particularly interesting".

> Suppose it can't. How will you know?

By finding and testing problems that o3 can't do on Deep Think, and also testing the reverse? Or by large benchmarks comparing a whole suite of questions with known answers.

Problems that both get correct will be easy to find and don't say much about comparative performance. That's why some of the benchmarks listed in the article (e.g. Humanity's Last Exam / AIME 2025) are potentially more insightful than one person's report on testing one question (which they don't provide) where both models replied with the same answer.

Re: Gemini 2.5 Deep Think

#158
post #120

Earlier quoted context omitted.

Easily the best one yet!

Saw one today from gpt5 (via some api trick someone found) that was better than this, let me see if I can find it. Pelican: https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd.... Longer thread re gpt5: https://old.reddit.com/r/OpenAI/comments/1mettre/gpt5_is_alr...

Uh that doesn't look better. it has more texture but the composition is bad/incomplete

Re: Gemini 2.5 Deep Think

#160
post #108

Earlier quoted context omitted.

OK that is recognizably a pelican, pretty great!

This feels like the best pelicanbike yet. The singularity might be closer than we imagine. Time for a leaderboard?

Ask and you'll receive: https://pelicans.borg.games/
Post reply on HN