Live data from Hacker News

Gemini 3.1 Pro

blog.google

181–190 of 951 posts

Re: Gemini 3.1 Pro

#181

In an attempt to get outside of benchmark gaming I had it make Platypus on a Tricycle. It's not as good as pelican on bicycle. https://www.svgviewer.dev/s/BiRht5hX

For a moment I assumed the output would look like Perry the Platipus from the Disney (I think?) show. It's suprising to me (as a layman) that a show with lots of media that would've made it to the training corpus didn't show up.

Re: Gemini 3.1 Pro

#182

Earlier quoted context omitted.

100% agreed. I wish someone would make a test for how reliably the LLMs follow tool use instructions etc. The pelicans are nice but not useful for me to judge how well a model will slot into a production stack.

At first when I got started with using LLMs I read/analyzed benchmarks, looked at what example prompts people used and so on, but many times, a new model does best at the benchmark, and you think it'll be better, but then in real work, it completely drops the ball. Since then I've stopped even reading benchmarks, I don't care an iota about them, they always seem more misdirected than helpful. Today I have my own priv…

share

Re: Gemini 3.1 Pro

#183

Earlier quoted context omitted.

It's an excellent demonstration of the main issue I have with the Gemini family of models, they always go "above and beyond" to do a lot of stuff, even if I explicitly prompt against it. In this case, most of the SVG ends up consisting not just of a bike and a pelican, but clouds, a sun, a hat on the pelican and so much more. Exactly the same thing happens when you code, it's almost impossible to get Gemini to not do…

> it's almost impossible to get Gemini to not do "helpful" drive-by-refactors Just asking "Explain what this service does?" turns into [No response for three minutes...] +729 -522

"I don't know what did it, but here's what it does now"

Re: Gemini 3.1 Pro

#184

I really want to use google’s models but they have the classic Google product problem that we all like to complain about. I am legit scared to login and use Gemini CLI because the last time I thought I was using my “free” account allowance via Google workspace. Ended up spending $10 before realizing it was API billing and the UI was so hard to figure out I gave up. I’m sure I can spend 20-40 more mins to sort this ou…

So much this. It's absolutely amazing how hostile Google is to releasing billing options that are reasonable, controllable, or even fucking understandable. I want to do relatively simple things like: 1. Buy shit from you 2. For a controllable amount (ex - let me pick a limit on costs) 3. Without spending literally HOURS trying to understand 17 different fucking products, all overlapping, with myriad project configs,…

You think AWS is better?

Re: Gemini 3.1 Pro

#185

To use in OpenCode, you can update the models it has: opencode models --refresh Then /models and choose Gemini 3.1 Pro You can use the model through OpenCode Zen right away and avoid that Google UI craziness. --- It is quite pricey! Good speed and nailed all my tasks so far. For example: @app-api/app/controllers/api/availability_controller.rb @.claude/skills/healthie/SKILL.md Find Alex's id, and add him to the block…

I don't see it even after refresh. Are you using the opencode-gemini-auth plugin as well?

No I am not just vanilla OpenCode. I do have OpenCode Zen credits, and I did opencode login whatever their command is to auth against opencode itself. Maybe that's the reason I see these premium models.

Re: Gemini 3.1 Pro

#186
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Less pretty and more practical, it's really good at outputting circuit designs as SVG schematics. https://www.svgviewer.dev/s/dEdbH8Sw

that's pretty amazing for an LLM but as an EE, if my intern did this i would sigh inwardly and pull up some existing schematics for some brief guidance on symbol layout.

Re: Gemini 3.1 Pro

#187
post #82

Google seems to really pull ahead in this AI race. For me personally they offer the best deal and although the software is not quiet there compared to openai or anthropic (in regards to 1. web GUI, 2. agent-cli). I hope they can fix that in the future and I think once Gemini 4 or whatever launches we will see a huge leap again

I don't understand this sentiment. It may hold true for other LLM use cases (image generation, creative writing, summarizing large texts), but when it comes to coding specifically, Google is *always* behind OpenAI and Anthropic, despite having virtually infinite processing power, money, and being the ones who started this race in the first place.

Until now, I've only ever used Gemini for coding tests. As long as I have access to GPT models or Sonnet/Opus, I never want to use Gemini. Hell, I even prefer Kimi 2.5 over it. I tried it again last week (Gemini Pro 3.0) and, right at the start of the conversation, it made the same mistake it's been making for years: it said "let me just run this command," and then did nothing.

My sentiment is actually the opposite of yours: how is Google *not* winning this race?

Re: Gemini 3.1 Pro

#188

Earlier quoted context omitted.

Models are soon going to start benchmaxxing generating SVGs of pelicans on bikes

Soon? I'd be willing to bet it's been included in the training set at least 6 months by now. Not so obvious so it generates always perfect pelicans on bikes, but sufficiently for the "minibench" to be less useful today than in the past.

[deleted]
Post reply on HN