Live data from Hacker News

Gemini 3 Deep Think

blog.google

681–690 of 722 posts

Re: Gemini 3 Deep Think

#681
post #626

Earlier quoted context omitted.

> at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs Every time anyone has tried that it excludes one or more classes of human life, and sometimes led to atrocities. Let's just skip it this time.

I excluded all right handed, blue eyed people yesterday before breakfast. No atrocities happened because of it.

Exactly, there's a few extra steps between here and there, and it's possible to pick out what those steps are without having to conclude that giving up on all brain research is the only option.

Re: Gemini 3 Deep Think

#682

Earlier quoted context omitted.

Here is what the original paper for ARC-AGI-1 said in 2019: > Our definition, formal framework, and evaluation guidelines, which do not capture all facets of intelligence, were developed to be actionable, explanatory, and quantifiable, rather than being descriptive, exhaustive, or consensual. They are not meant to invalidate other perspectives on intelligence, rather, they are meant to serve as a useful objective fun…

https://www.dwarkesh.com/p/francois-chollet (June 2024, about ARC-AGI-1. Note the AGI right in the name) > I’m pretty skeptical that we’re going to see an LLM do 80% in a year. That said, if we do see it, you would also have to look at how this was achieved. If you just train the model on millions or billions of puzzles similar to ARC, you’re relying on the ability to have some overlap between the tasks that you trai…

He has been wrong about timelines and about what specific approaches would ultimately solve ARC-AGI 1 and 2. But he is hardly alone in that. I also won't argue if you call him smug. But he was right about a lot of things, including most importantly that scaling pretraining alone wouldn't break ARC-AGI. ARC-AGI is unique in that characteristic among reasoning benchmarks designed before GPT-3. He deserves a lot of credit for identifying the limitations of scaling pretraining before it even happened, in a precise enough way to construct a quantitative benchmark, even if not all of his other predictions were correct.

Re: Gemini 3 Deep Think

#683

Earlier quoted context omitted.

When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.

I think no matter what happens with AI in the future, there will always be a subset of people with elaborate conspiracies about how it's all fake/a hoax.

I'm not saying it's a hoax. If it gets better because of that data, tant mieux, but we have to be clear eyed about what these models are actually doing. Especially when companies don't explain what they've done.

Re: Gemini 3 Deep Think

#685
Is xAI out of the race? I’m not on a subscription, but their Ara voice model is my favorite. Gemini on iOS is pretty terrible in voice mode. I suspect because they have aggressive throttling instructions to keep output tokens low.

Re: Gemini 3 Deep Think

#686
post #133

it is interesting that the video demo is generating .stl model. I run a lot of tests of LLMs generating OpenSCAD code (as I have recently launched https://modelrift.com text-to-CAD AI editor) and Gemini 3 family LLMs are actually giving the best price-to-performance ratio now. But they are very, VERY far from being able to spit out a complex OpenSCAD model in one shot. So, I had to implement a full fledged "screensho…

If you want that to get better, you need to produce a 3d model benchmark and popularize it. You can start with a pelican riding a bicycle with working bicycle.

I am building pretty much the same product as OP, and have a pretty good harness to test LLMs. In fact I have run a tons of tests already. It’s currently aimed for my own internal tests, but making something that is easier to digest should be a breeze. If you are curious: https://grandpacad.com/evals

Re: Gemini 3 Deep Think

#687

Earlier quoted context omitted.

An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.

My dog wags his tail hard when I ask "hoosagoodboi?". Pretty definitive I'd say.

I'm fairly sure he'd have the same response if you asked them "who's a good lion" in the same tone of voice.

*I tried hard to find an animal they wouldn't know. My initial thought of cat was more likely to fail.

Re: Gemini 3 Deep Think

#688

this is like the doomsday clock 84% is meaningless if these things can't reason getting closer and closer to 100%, but still can't function

> if these things can't reason

I see people talk about "reasoning". How do you define reasoning such that it is clear humans can do it and AI (currently) cannot?

Re: Gemini 3 Deep Think

#689

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

I read somewhere that Google will ultimately always produce the best LLMs, since "good AI" relies on massive amounts of data and Google owns the most data. Is that a based assumption?

No.
Post reply on HN