Live data from Hacker News

Gemini 3.1 Pro

blog.google

931–940 of 951 posts

Re: Gemini 3.1 Pro

#931

Earlier quoted context omitted.

https://one.google.com/about/google-ai-plans/?utm_source=g1&...

At least Anthropic tells you how many more tokens you’re paying for! 5x 10x 20x whatever. Google seems to just say more, higher, highest.

The pricing page for Claude literally says "More usage" for the $17/month pro plan. Doesn't really quantify anything. The usage is whatever they feel like it should be.

And then the very expensive plan says "Choose 5x or 20x more usage than Pro". It's all arbitrary.

Re: Gemini 3.1 Pro

#932
post #768

Earlier quoted context omitted.

> In any case they aren't trying to hide it particularly hard. What does that mean? Are you able to read the raw cot? how?

My guess they mean Google create those summaries via tool use and not trying to filter actual chain of thoughts on API level or return errors if model start leaking it. If you work with big contexts in AI Studio (like 600,000-900,000 tokens) it sometimes just breaks downs on its own and starts returning raw cot without any prompt hacking whatsoever. I believe if you intentionally try to expose it that would be pretty…

3.1 bugged and gave CoT for me yesterday

Re: Gemini 3.1 Pro

#933
post #694
post #522

Earlier quoted context omitted.

GPT 5.2 loses at everything but they included that

Who are they supposed to compare it to? I'm not sure what makes you think that Grok is even remotely comparable to the frontier models right now.

Grok has been and still is the best at incorporating search.

4.20 with its 4 agents puts it back at the top for reasoning as well. As soon as it's added to the API, the benchmarks should show that.

Re: Gemini 3.1 Pro

#934

Earlier quoted context omitted.

If that's what you're after, tou MITM it and setup a proxy so Claude Code or whatever sends to your program, and then that program forwards it to Anthropics's server (or whomever). That way, you get everything .

I'm aware that this is possible, and thank you for the suggestion, but surely you can see that it's a relatively large lift; may not work in controlled enterprise environments; and compared to just right click -> view source it's basically inaccessible to anyone who might have wanted to dabble.

If you can't be bothered to build it youself, use someone else's. https://github.com/jmuncor/tokentap made the rounds here ~three weeks ago.

https://news.ycombinator.com/item?id=46799898

Re: Gemini 3.1 Pro

#935
After 2 days of giving it a go, I find that Gemini CLI is still considerably worse than both Codex and Claude Code.

The model itself also has strange behaviors that seem like it gets randomly replaced with Gemini-3-Flash or something else. I'll explain.

Once agentic coding was a bust, I gave it a run as a daily driver for AI assistant. It performed fairly well but then began behaving strangely. It would lose context mid conversation. For instance, I said "In san francisco I'm looking for XYZ". Two turns later I'm asking about food and it gives me suggestions all over the world.

Another time, I asked it about the likelihood of the pending east coast winter storm of affecting my flight. I gave it all the details (flight, stops, time, cities).

Both GPT-5.2 and Claude crunched and came back with high quality estimations and rationale. Gemini 3.1 Pro... 5 times, returned a weather forecast widget for either the layover or final destination. This was on "Pro" reasoning, the highest exposed on the Gemini App/WebApp. I've always suspected Google swaps out models randomly so this.. wasn't surprising.

I then asked Gemini 3.1 Pro via the API and it returned a response similar to Claude and GPT-5.2 -- carefully considering all factors.

This tells me that a Google AI Ultra subscription gives me a sub-par coding agent which often swaps in Flash models, a sub-par web/app AI experience that also isn't using the advertised SOTA models, and a bunch of preview apps for video gen, audio gen (crashed every time I attempted), and world gen (Genie was interesting but a toy).

This will be a quick cancel as soon as the intro rate is done.

It's like Google doesn't ACTUALLY want to be the leader in AI or serve people their best models. They want to generate hype around benchmarks and then nerf the model and go silent.

Gemini 3 Pro Preview went from exceptional in the first month to mediocre and then out of my rotation within a month.

Re: Gemini 3.1 Pro

#936

Earlier quoted context omitted.

Lol Ive admitted im a google employee, not hiding my bias. Most things aren't worth commenting on except the gemini posts here, which I find insane. And pretty much every example you gave Id expect quite a lot more for 2x the amount? Idk man

So you are an anthropic / openai employee, actually.

Lol i wish.

Re: Gemini 3.1 Pro

#937
post #741

Earlier quoted context omitted.

Google is a cloud provider so API usage is funneled thru GCP. It's the same for Microsoft and Amazon.

By that logic, G Suite should be funneled through GCP. Also, are you sure you meant to mention Microsoft? Microsoft has this Copilot thing that they will gladly sell you, with generally inoffensive commercial terms, through more channels than you can shake a stick at. Got a $4 GitHub for Teams subscription? Add $20 or so and you will be swimming in Copilot outputs, and all you have to do is check the checkbox.

Got a free Gmail account? Add $20 or so and you'll be swimming in Gemini outputs. Yet both companies also have a cumbersome onboarding process if all you want to do is get an API token. So yeah, quite similar!

Re: Gemini 3.1 Pro

#938

Earlier quoted context omitted.

Even though I don't like the privacy implications, make sure you use the option to save and use past chats for context. After a few months of back and forth (hundreds of 'chat' sessions), the responses are much higher quality. It sometimes does 'callbacks' to things discussed in past chats, which are typically awkward non-sequiturs, but it does improve it overall. When I play with it in 'temporary chat' mode that ign…

You must be joking. I’ve turned that off after first month of use. It’s unbearable. “Oh since you are in {place i mentioned a week ago while planning trip but ultimately didnt go} the home assistant integration question changes completely”. Or ending every answer with “since you are salesforce consultant, would you like to learn more about iron smelting?”

I told Gemini I'm a software engineer and it explains absolutely everything in programming metaphors now. I think it's way undertrained with personalization.

Re: Gemini 3.1 Pro

#939

Earlier quoted context omitted.

I understand. But isn't it a sign of "smarts" that one can generalize from analoguous tasks?

LLMs are great at knowledge transfer, the real question is how well can they demonstrate intelligence with "unknown unknown" types of questions. This model has the benefit of being released after that issue became public knowledge, so it's hard to know how it would've performed pre-hoc.

There's a long delay ("knowledge cutoff") in model training, so it probably hasn't seen the question before.

Re: Gemini 3.1 Pro

#940
post #611

Earlier quoted context omitted.

Tell me more about Codex. I'm trying to understand it better. I have a pretty crude mental model for this stuff but Opus feels more like a guy to me, while Codex feels like a machine. I think that's partly the personality and tone, but I think it goes deeper than that. (Or maybe the language and tone shapes the behavior, because of how LLMs work? It sounds ridiculous but I told Claude to believe in itself and suddenl…

Your intuition is exactly correct - it's not just 'tone' it's 'deeper than that'. Codex is a 'poor communicator' - which matters surprisingly a lot in these things. It's overly verbose, it often misses the point - but - it is slightly stronger in some areas. Also - Codex now has 'Spark' which is on Cerebras, it's wildly fast - and this absolutely changes 'workflow' fundamentally. With 'wait-thinking' - you an have 3-…

> Also - Codex now has 'Spark' which is on Cerebras, it's wildly fast - and this absolutely changes 'workflow' fundamentally.

In my AI coding experience, reviewing and making sure AI didn't screw up something (eg: by writing tutorial grade code) takes most of the time. It's still useful but I don't see how speeding up the non-bottleneck part can change the workflow fundamentally.

Post reply on HN