Live data from Hacker News

Gemini 3 Deep Think

blog.google

691–700 of 722 posts

Re: Gemini 3 Deep Think

#691

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

Yet it still can't solve a Pokle hand for me

Re: Gemini 3 Deep Think

#692

Earlier quoted context omitted.

https://www.dwarkesh.com/p/francois-chollet (June 2024, about ARC-AGI-1. Note the AGI right in the name) > I’m pretty skeptical that we’re going to see an LLM do 80% in a year. That said, if we do see it, you would also have to look at how this was achieved. If you just train the model on millions or billions of puzzles similar to ARC, you’re relying on the ability to have some overlap between the tasks that you trai…

He has been wrong about timelines and about what specific approaches would ultimately solve ARC-AGI 1 and 2. But he is hardly alone in that. I also won't argue if you call him smug. But he was right about a lot of things, including most importantly that scaling pretraining alone wouldn't break ARC-AGI. ARC-AGI is unique in that characteristic among reasoning benchmarks designed before GPT-3. He deserves a lot of cred…

Totally agree. And I hope he continues to be a sort of confident red-teamer like he has been, it's immensely valuable. At some level if he ever drinks the AGI kool-aid we will just be looking for another him to keep making up harder tests.

Re: Gemini 3 Deep Think

#693
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Their models might be impressive, but their products absolutely suck donkey balls. I’ve given Gemini web/cli two months and ran away back to ChatGPT. Seriously, it would just COMPLETELY forget context mid dialog. When asked about improving air quality it just gave me a list of (mediocre) air purifiers without asking for any context whatsoever, and I can list thousands of conversations like that. Shopping or comparing…

It’s all context/ use case; I’ve had weird things too but if you only use markdown inputs and specific prompts Gemini 3 Pro is insane, not to mention the context window

Also because of the long context window (1 mil tokens on thinking and pro! Claude and OpenAI only have 128k) deep research is the best

That being said, for coding I definitely still use Codex with GPT 5.3 XHigh lol

Re: Gemini 3 Deep Think

#694
post #597
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

I'd personally bet on Google and Meta in the long run since they have access to the most interesting datasets from their other operations.

Agree. Anyone with access to large proprietary data has an edge in their space (not necessarily for foundation models): Salesforce, adobe, AutoCAD, caterpillar

Re: Gemini 3 Deep Think

#695
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Their models might be impressive, but their products absolutely suck donkey balls. I’ve given Gemini web/cli two months and ran away back to ChatGPT. Seriously, it would just COMPLETELY forget context mid dialog. When asked about improving air quality it just gave me a list of (mediocre) air purifiers without asking for any context whatsoever, and I can list thousands of conversations like that. Shopping or comparing…

I've used their Pro models very successfully in demanding API workloads (classification, extraction, synthesis). On benchmarks it crushed the GPT-5 family. Gemini is my default right now for all API work.

It took me however a week to ditch Gemini 3 as a user. The hallucinations were off the charts compared to GPT-5. I've never even bothered with their CLI offering.

Re: Gemini 3 Deep Think

#696

Earlier quoted context omitted.

It's basically bunch of people who see themselves as too smart to believe in God, instead they have just replaced it with AI and Singularity and attribute similar stuff to it eg. eternal life which is just heaven in religion. Amodei was hawking doubling of human lifespan to a bunch of boomers not too long ago. Ponce de León also went to search for the fountain of youth. It's a very common theme across human history.…

You realize that science and technology does in fact produce medical breakthroughs that cure disease, right? On the other hand, prayer doesn’t heal anybody and there’s no proof of supernatural beings.

The boomers he was talking to will be long underground before we will have any major cures for the diseases they will die from lmao. Maybe in 200 years?

Btw, so will you and I most likely.

Re: Gemini 3 Deep Think

#697

Earlier quoted context omitted.

Are there benchmarks if we allow the LLM to practice and study the game?

You can make one, the balatro bench is open source. But I'm quite sure it'd be crazily expensive for a hobby project. At the end of the day, LLM can't actually 'practice and learn.'

I've gotten pretty good results by prompting "What did you struggle on? Please update the instructions in " and "Here's your conversation , please see what you struggled with and update ".

It's hit or miss, but I've been able to have it self improve on prompts. It can spot mistakes and retain things that didn't work. Similar to how I learned games like Balatro. Playing Balatro blind, you wouldn't know which jokers are coming and have synergy together, or that X strategy is hard to pull off, or that you can retain a card to block it from appearing in shops.

If the LLM can self discover that, and build prompt files that gradually allow it to win at the highest stake, that's an interesting result. And I'd love to know which models do best at that.

Re: Gemini 3 Deep Think

#698
post #484

Earlier quoted context omitted.

Antigravity is an embarrassment. The models feel terrible, somehow, like they're being fed terrible system prompts. Plus the damn thing kept crashing and asking me to "restart it". What?! At least Kiro does what it says on the tin.

My experience with Antigravity is the opposite. It's the first time in over 10 years that an IDE has managed to take me out a bit out of the jetbrain suite. I did not think that was something possible as I am a hardcore jetbrain user/lover.

Have you tried Cursor or VS Code with Github Copilot in agent mode (recently, not 3 or 6 months ago)?

I've recently tried a buuuuunch of stuff (including Antigravity and Kiro) and I really, really, could not stomach Antigravity.

Re: Gemini 3 Deep Think

#699

Earlier quoted context omitted.

> at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs Every time anyone has tried that it excludes one or more classes of human life, and sometimes led to atrocities. Let's just skip it this time.

Having trouble parsing this one. Is it meant to be a WWII reference? If anything I would say consciousness research has expanded our understanding of living beings understood to be conscious. And I don't think it's fair or appropriate to treat study of the subject matter of consciousness like it's equivalent to 20th century authoritarian regimes signing off on executions. There's a lot of steps in the middle before y…

> Is it meant to be a WWII reference?

The sum total of human history thus far has been the repetition of that theme. "It's OK to keep slaves, they aren't smart enough to care for themselves and aren't REALLY people anyhow." Or "The Jews are no better than animals." Or "If they aren't strong enough to resist us they need our protection and should earn it!"

Humans have shown a complete and utter lack of empathy for other humans, and used it to justify slavery, genocide, oppression, and rape since the dawn of recorded history and likely well before then. Every single time the justification was some arbitrary bar used to determine what a "real" human was, and consequently exclude someone who claimed to be conscious.

This time isn't special or unique. When someone or something credibly tells you it is conscious, you don't get to tell it that it's not. It is a subjective experience of the world, and when we deny it we become the worst of what humanity has to offer.

Yes, I understand that it will be inconvenient and we may accidentally be kind to some things that didn't "deserve" kindness. I don't care. The alternative is being monstrous to some things that didn't "deserve" monstrosity.

Re: Gemini 3 Deep Think

#700
post #188

Earlier quoted context omitted.

> . I don't think there are many people who posted their Balatro playthroughs in text form online There are * tons * of balatro content on YouTube though, and it makes absolutely zero doubt that Google is using YouTube content to train their model.

Yeah, or just the steam text guides would be a huge advantage. I really doubt it's playing completely blind

Thanks to another comment here I went looking for the strategy guides that are injected. To save everyone else the trouble, here [0]. Look at (e.g.) default/STRATEGY.md.jinja. Also adding a permalink [1] for future readers' sake.

[0]: https://github.com/coder/balatrollm/tree/main/src/balatrollm...

[1]: https://github.com/coder/balatrollm/blob/a245a0c2b960b91262c...

Post reply on HN