Earlier quoted context omitted.
Any tips for working with Gemini through its chat interface? I’ve worked with ChatGPT and Claude and I’ve generally found them pleasant to work with, but everytime I use Gemini the output is straight dookie
Even though I don't like the privacy implications, make sure you use the option to save and use past chats for context. After a few months of back and forth (hundreds of 'chat' sessions), the responses are much higher quality. It sometimes does 'callbacks' to things discussed in past chats, which are typically awkward non-sequiturs, but it does improve it overall. When I play with it in 'temporary chat' mode that ign…
Gemini 3.1 Pro
811–820 of 951 posts
Re: Gemini 3.1 Pro
#812Earlier quoted context omitted.
share
The harness? Trivial to build yourself, ask your LLM for help, it's ~1000 LOC you could hack together in 10-15 minutes. As for the test cases themselves, that would obviously defeat the purpose, so no :)
Re: Gemini 3.1 Pro
#813Earlier quoted context omitted.
> Knowledge cutoff is unchanged at Jan 2025. Isn't that a bit old?
Old relative to its competitors, but the Search tool can compensate for it.
Gemini 3.0 was convinced that my dependency versions pinned in package.json were hallucinated by an AI, because they "shouldn't yet exist". I just hope this kind of behavior is gone.
Re: Gemini 3.1 Pro
#814Earlier quoted context omitted.
Francois Chollet accuses the big labs of targeting the benchmark, yes. It is benchmaxxed.
Didn't the same Francois Chollet claim that this was the Real Test of Intelligence? If they target it, perhaps they target... real intelligence?
He also said that the "real test of intelligence" is being unable to come up with new tests that a human can easily do that the AI can't, not in being able to pass any specific benchmark.
Re: Gemini 3.1 Pro
#815Earlier quoted context omitted.
I'm not denying any progress, I'm saying that reasoning failures that are simple which have gone viral are exactly the kind of thing that they will toss in the training data. Why wouldn't they? There's real reputational risks in not fixing it and no costs in fixing it.
Given that Gemini 3 Pro already did solid on that test, what exactly did they improve? Why would they bother? I double checked and tested on AI Studio, since you can still access the previous model there: >You should drive. >If you walk there, your car will stay behind, and you won't be able to wash it. Thinking models consistently get it correct and did when the test was brand new (like a week or two ago). It is the…
Re: Gemini 3.1 Pro
#816My enthusiasm is a bit muted this cycle because I've been burned by Gemini CLI. These models are very capable but Gemini CLI just doesn't seem to be able to work for one it never follows instructions strictly like its competitors do, and it hallucinates even which is a rarity. More importantly feels like Google is stretched thin across different Gemini products and pricing reflects this, I still have no idea how to p…
> I still have no idea how to pay for Gemini CLI, in codex/claude its very simple $20/month for entry and $200/month for ton of weekly usage. This! I would like to sign up for a paid plan for Gemini CLI. But I have not been able to figure out how. I already have Codex and Claude plans. Those were super easy to sign up for.
Re: Gemini 3.1 Pro
#817I hope this works better than 3.0 Pro I'm a former Googler and know some people near the team, so I mildly root for them to at least do well, but Gemini is consistently the most frustrating model I've used for development. It's stunningly good at reasoning, design, and generating the raw code, but it just falls over a lot when actually trying to get things done, especially compared to Claude Opus. Within VS Code Copi…
Re: Gemini 3.1 Pro
#818These models are so powerful. It's totally possible to build entire software products in the fraction of the time it took before. But, reading the comments here, the behaviors from one version to another point version (not major version mind you) seem very divergent. It feels like we are now able to manage incredibly smart engineers for a month at the price of a good sushi dinner. But it also feels like you have to b…
Re: Gemini 3.1 Pro
#819I’m keen to know how and where are you using Gemini. Anthropic is clearly targeted to developers and OpenAI is general go to AI model. Who are the target demographic for Gemini models? ik that they are good and Flash is super impressive. but i’m curious
Re: Gemini 3.1 Pro
#820Does well on SVGs outside of "pelican riding on a bicycle" test. Like this prompt: "create a svg of a unicorn playing xbox" https://www.svgviewer.dev/s/NeKACuHj Still some tweaks to the final result, but I am guessing with the ARC-AGI benchmark jumping so much, the model's visual abilities are allowing it to do this well.