Live data from Hacker News

Gemini 3.1 Pro

blog.google

841–850 of 951 posts

Re: Gemini 3.1 Pro

#841

People underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

Gemini is the most paradoxical model because it benchmarks great even in private benchmarks done by regular people, Deep Mind is unquestionably full of capable engineers with incredible skill, and personally Gemini has been great for my day job and my coding for fun (not for profit) endeavors. Switching between it and 4.6 in antigravity and I don't see much of a difference, they both do what I ask. But man, people ar…

I feel like a lot of this is just Googles tooling - if you're using Antigravity/Gemini CLI and then use Claude Code it feels like a huge difference. I can say from experience though (using Cline + OpenCode) that they are really close.

The harness is just much better on the Anthropic side.

Re: Gemini 3.1 Pro

#842
post #151
post #82

Google seems to really pull ahead in this AI race. For me personally they offer the best deal and although the software is not quiet there compared to openai or anthropic (in regards to 1. web GUI, 2. agent-cli). I hope they can fix that in the future and I think once Gemini 4 or whatever launches we will see a huge leap again

I hope they fail. I honestly do not wish Google to have the best model out there and be forced to use their incomprehensible subscription / billing / project management whatever shit ever again. I don’t know what their stuff cost. I don’t know why would I use vertex or ai studio. What is included in my subscription what is billed per use. I pray that whatever they build fails and burns.

after using aistudio fine for months suddenly my billing was cancelled and a week later im still waiting for it to be re-enabled.

Im at a total loss to how google can function this way, my only explanation is they somehow have a Philosophers Stone they generate wealth with because they sure as hell make it impossible to give them money.

Re: Gemini 3.1 Pro

#843

Earlier quoted context omitted.

It matters to me. I pay for it and I like using it. I pick my models to keep my spend reigned in.

What do you use it for? What is your time worth that you'd settle for a lesser model to save a few bucks?

Homelab and hobby assistant. I have spent $300 for 12 months of tokens. If I'm burning up more than $25 a month then I'd have to pay more or curb use at the end of the year. $25 / month as a new expense is something I can accept for a toy that is letting me accelerate my fun stuff. I can't justify more than that. So I'm left constantly evaluating if my current task is worth more than future tasks and if it is expected to be harder than future tasks. Speculative execution is already one of the harder things I do at work.

Re: Gemini 3.1 Pro

#844

People underrate Google's cost effectiveness so much. Half price of Opus. HALF. Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight ____ Update: 3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed. https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

You can pay 1 cent for a mediocre answer or 2 cents for a great answer. So a lot of these things are relative. Now if that equation plays out 20K times a day, well that's one thing, but if it's 'once a day' then the cost basis becomes irrelevant. Like the cost of staplers for the Medical Device company. Obviously it will matter, but for development ... it's probably worth it to pay $300/mo for the best model, when th…

Quality is Anthropic's game.

Quantity is OpenAi's.

Google's is... specialized hardware? (For now.)

Also deeper crawls, and Google Books! (Though it's unclear if they're making good use of those.)

Re: Gemini 3.1 Pro

#845

Earlier quoted context omitted.

Plus they started making AI processors 11 years ago and invented the math behind “GPTs” 9 years ago. Gemini is way cheaper to run for them than it does for everyone else. I think Gemini is really built for their biggest market — Google Search. You ask questions and get answers. I’m sure they’ll figure out agentic flows. Google is always a mess when it comes to product. Don’t forget the Google chat sagas where it seem…

Who they? Do the engineers who actually did that work at Google still? I heard that the guy who made TPUs has his own startup now.

Only one guy built the TPUs?

Re: Gemini 3.1 Pro

#846
post #449

I asked Gemini 3.1 Pro to generate some of the modern artworks in my "Pelican Art Gallery". I particularly like the rendition of the Sunflowers: https://pelican.koenvangilst.nl/gallery/category/modern

Nice collection of visible bits that have no relation at all with art

Re: Gemini 3.1 Pro

#847

Earlier quoted context omitted.

LLMs are pretty mediocre for a lot of money queries like searching to buy shoes, looking at flights etc due to them not being up to date. So sure you can use them as a wrapper on top of Google but I assume a huge chunk of people will just go to Google to do that or use Google agents. Chrome will prove a very valuable asset for that - the whole experience can become agentic and Google is very well positioend to conver…

LLMs can execute searches? You can absolutely send ChatGPT to look for a cheap flight and it will do pretty well. And because I am paying for ChatGPT rather than the advertiser's, I am the customer and not the product.

You may pay to ChatGPT, but sooner or later you will become their product too. All the conversations you had or will have will be turned into signals to match you with products from advertisers, maybe not directly in the conversation with them, but anywhere else. It's not a mater of if, but looking at the pace things are going, and how financially pressured openai is, it's only a matter of time that their conversations with them will be turned into profit in some way or another, they basically have no choice financially.

Re: Gemini 3.1 Pro

#848

Every time I've used Gemini models for anything besides code or agentic work they lean so far into the RLHF induced bold lettering and bullet point list barf that everything they output reads as if the model was talking _at_ me and not _with_ me. In my Openclaw experiment(s) and in the Gemini web UI, I've specifically added instructions to avoid this type of behavior, but it only seemed to obey those rules when I rem…

You just articulated why I struggle to personally connect with Gemini. It feels so unrelatable and exhausting to read its output. I prefer to read Opus/Deepseek/GLM over Gemini, Qwen and the open source GPT models. Maybe it is RLHF that is creating my distaste from using it. (I pay for Gemini; I should be using it more... but the outputs just bug me and feel more work to get actionable insight.)

> feel more work to get actionable insight

WHAT?! I find that exactly the nice sharp formatting are what makes it EASIER to get actionable insight from it...

(Plus the weird-but-cute unrequested analogies are nice to occassionally elicit a smile and keep you motivated :P)

Re: Gemini 3.1 Pro

#849

Earlier quoted context omitted.

You are holding it wrong! No but for real, what is your usecase? Do you acutely think something like gpt3 was best?

I dont have a real special usecase, i just use it whenever i think it will give better results than googling or thinking or i dont feel like getting annoyed by cookie popups. And i dont think gpt3 was best, but it felt like it actually listened. Now i tell it: "You did this and this wrong, i specifically told u the exact opposite. Can you please do what i asked you?" And then it says something like: "Oh yes my bad, y…

...you sound like a typical opus-person :P Just use anthropic's flagships if you want good instruction following, focus in long convos, and proper understanding of guidance-when-wrong.

Re: Gemini 3.1 Pro

#850
post #529

Earlier quoted context omitted.

the agentic benchmarks for 3.1 indicate Gemini has caught up. the gains are big from 3.0 to 3.1. For example the APEX-Agents benchmark for long time horizon investment banking, consulting and legal work: 1. Gemini 3.1 Pro - 33.2% 2. Opus 4.6 - 29.8% 3. GPT 5.2 Codex - 27.6% 4. Gemini Flash 3.0 - 24.0% 5. GPT 5.2 - 23.0% 6. Gemini 3.0 Pro - 18.0%

Benchmarks are basically straight up meaningless at this point in my experience. If they mattered and were the whole story, those Chinese open models would be stomping the competition right now. Instead they're merely decent when you use them in anger for real work. I'll withhold judgement until I've tried to use it.

Does anyone know what this "APEX-Agents benchmark for long time horizon investment banking, consulting and legal work" actually evaluates?

That sounds so broad that creating a meaningful benchmark is probably as difficult as creating an AI that actually "solves" those domains.

Post reply on HN