Earlier quoted context omitted.
I don't even know what the fuck "agentic" is or why the hell I would want it all over my software. So tired of everything in the computing world today.
Prompting, planning, iteration, coding, and tool use over an entire code base until a problem is solved.
Gemini 3
931–940 of 1001 posts
Re: Gemini 3
#932Re: Gemini 3
#933Re: Gemini 3
#934Earlier quoted context omitted.
I also used Gemini 3 Pro Preview. It finished it 271s = 4m31s. Sadly, the answer was wrong. It also returned 8 "sources", like stackexchange.com, youtube.com, mpmath.org, ncert.nic.in, and kangaroo.org.pk, even though I specifically told it not to use websearch. Still a useful tool though. It definitely gets the majority of the insights. Prompt: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...
Terrence Tao claims [0] contributions by the public are counter -productive since the energy required to check a contribution outweighs its benefit: > (for) most research projects, it would not help to have input from the general public. In fact, it would just be time-consuming, because error checking Since frontier LLMs make clumsy mistakes, they may fall into this category of 'error-prone' mathematician whose net c…
Re: Gemini 3
#935It still failed my image identification test ([a photoshopped picture of a dog with 5 legs]...please count the legs) that so far every other model has failed agonizingly, even failing when I tell them they are failing, and they tend to fight back at me. Gemini 3 however, while still failing, at least recognized the 5th leg, but thought the dog was...well endowed. The 5th leg however is clearly a leg, despite being wh…
> Gemini 3 however, while still failing, at least recognized the 5th leg, but thought the dog was...well endowed. I see that AI is reaching the level of a middle school boy...
Re: Gemini 3
#936Re: Gemini 3
#937Earlier quoted context omitted.
Unlike general public the models can be trained. I mean if you train a member of general public, you've got a specialist, who is no longer a member of general public.
Unlike the general public though, these models have advanced dementia when it comes to learning from corrections, even within a single session. They keep regressing and I haven't found a way to stop that yet. What boggles the mind: we have gone for so long to try to strive for correctness and suddenly being right 70% of the time and wrong the remaining 30% is fine. The parallel with self driving is pretty strong here…
Re: Gemini 3
#938Earlier quoted context omitted.
> We are cheering for a product sold back to us at a 60% markup (input costs up to $2.00/M) that was built on our own private correspondence. That feels like something between a hallucination and an intentional fallacy that popped up because you specifically said "intense discussion". The increase is 60% on input tokens from the old model, but it's not a markup, and especially not "sold back to us at X markup". I've…
60% probably felt like a lot to Gemini. However, I liked the doomerism and how google was using our data to train its models. Nonetheless, Gemini 3 failed this test. It failed to start a discussion. Its points were shallow, and too aiesque.
Looking at it again it's actually a completely nonsensical sentence that just happens to resemble a sensible statement in a way that would fool most people.
RL is definitely showing some busting seams at this point.
Re: Gemini 3
#939Earlier quoted context omitted.
Terrence Tao claims [0] contributions by the public are counter -productive since the energy required to check a contribution outweighs its benefit: > (for) most research projects, it would not help to have input from the general public. In fact, it would just be time-consuming, because error checking Since frontier LLMs make clumsy mistakes, they may fall into this category of 'error-prone' mathematician whose net c…
But he actually uses frontier LLMs in his own work. Probably that's stronger evidence.
Re: Gemini 3
#940Earlier quoted context omitted.
Unlike general public the models can be trained. I mean if you train a member of general public, you've got a specialist, who is no longer a member of general public.
Unlike the general public though, these models have advanced dementia when it comes to learning from corrections, even within a single session. They keep regressing and I haven't found a way to stop that yet. What boggles the mind: we have gone for so long to try to strive for correctness and suddenly being right 70% of the time and wrong the remaining 30% is fine. The parallel with self driving is pretty strong here…