Earlier quoted context omitted.
Wartime Google gave us Google+. Wartime Google is still bumbling, and despite OpenAI's numerous missteps, I don't think it has to worry about Google hurting its business yet.
Google+ was fun. Failed in the market though. Apple made a social network called Ping. Disaster. MobileMe was silly. Microsoft made Zune and the Kin 1 and Kin 2 devices and Windows phone and all sorts of other disasters. These things happen.
Gemini 3 Deep Think
661–670 of 722 posts
Re: Gemini 3 Deep Think
#662Earlier quoted context omitted.
Their models are absolutely not impressive. Not a single person is using it for coding (outside of Google itself). Maybe some people on a very generous free plan. Their model is a fine mid 2025 model, backed by enormous compute resources and an army of GDM engineers to help the “researchers” keep the model on task as it traverses the “tree of thoughts”. But that isn’t “the model” that’s an old model backed by massive…
Uhh, just false.
Come on.
Worthless.
Do you have any market counter points.
Market counter points that aren't really just a repackaging of:
1. "Google has the world's best distribution" and/or
2. "Google has a firehose of money that allows them to sell their 'AI product' at an enormous discount?
Good luck!Re: Gemini 3 Deep Think
#663Earlier quoted context omitted.
And after i do that, how do i combine the output of 1000 subagents into one output? (Im not being snarky here, i think it's a nontrivial problem)
You just pipe it to another agent to do the reduce step (i.e. fan-in) of the mapreduce (fan-out) It's agents all the way down.
Re: Gemini 3 Deep Think
#664Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.
Re: Gemini 3 Deep Think
#665Earlier quoted context omitted.
Not without people later saying "you shared that on Hacker News last year clearly the AI labs are training for it now!"
Couldn't you just make up new combinations, or new caveats indefinitely to mitigate that? It would be nice to see maybe 3-4 good examples for validation. I'd do it myself, but I don't have $200 to play around with this model.
Re: Gemini 3 Deep Think
#666I learned a lot about Gemini last night. Namely that I have lead it like a reluctant bull to understand what I want it to do (beyond normal conversations, etc).
Don't get me wrong, ChatGPT didn't do any better.
It's an important spreadsheet so I'm triple checking on several LLM's and, of course, comparing results with my own in depth understanding.
For running projects, and making suggestions, and answering questions and being "an advisor", LLM's are fantastic ... feed them a basic spreadsheet and it doesn't know what to do. You have to format the spreadsheet just right so that it "gets it".
I dread to think of junior professionals just throwing their spreadsheets into LLM's and runninng with the answers.
Or maybe I'm just shit at prompting LLM's in relation to spreadsheets. Anyone had better results in this scenario?
Re: Gemini 3 Deep Think
#667Earlier quoted context omitted.
Couldn't you just make up new combinations, or new caveats indefinitely to mitigate that? It would be nice to see maybe 3-4 good examples for validation. I'd do it myself, but I don't have $200 to play around with this model.
Here's what it gave me for a kakapo on a skateboard https://gist.github.com/simonw/5e2041c32333effd090e3df42b64d...
Re: Gemini 3 Deep Think
#668Earlier quoted context omitted.
Let's come back in 12 months and discuss your singularity then. Meanwhile I spent like $30 on a few models as a test yesterday, none of them could tell me why my goroutine system was failing, even though it was painfully obvious (I purposefully added one too many wg.Done), gemini, codex, minimax 2.5, they all shat the bed on a very obvious problem but I am to believe they're 98% conscious and better at logic and math…
It's basically bunch of people who see themselves as too smart to believe in God, instead they have just replaced it with AI and Singularity and attribute similar stuff to it eg. eternal life which is just heaven in religion. Amodei was hawking doubling of human lifespan to a bunch of boomers not too long ago. Ponce de León also went to search for the fountain of youth. It's a very common theme across human history.…
On the other hand, prayer doesn’t heal anybody and there’s no proof of supernatural beings.
Re: Gemini 3 Deep Think
#669Earlier quoted context omitted.
It beats ante eight 9 times out of 15 attempts. I do consider 60% winning chance very good for a first time player. The average is only 19.3 rounds because there is a bugged run where Gemini beats round 6 but the game bugs out when it attempts to sell Invisible Joker (a valid move)[0]. That being said, Gemini made a big mistake in round 6 that would have costed it the run at higher difficulty. [0]: given the existenc…
Are there benchmarks if we allow the LLM to practice and study the game?
Re: Gemini 3 Deep Think
#670Earlier quoted context omitted.
Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…
Hi, BalatroBench creator here. Yeah, Google models perform well (I guess the long context + world knowledge capabilities). Opus 4.6 looks good on preliminary results (on par with Gemini 3 Pro). I'll add more models and report soon. Tbh, I didn't expect LLMs to start winning runs. I guess I have to move to harder stakes (e.g. red stake).
It's what I did for my game benchmark https://d.erenrich.net/paperclip-bench/index.html