Earlier quoted context omitted.
> at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs Every time anyone has tried that it excludes one or more classes of human life, and sometimes led to atrocities. Let's just skip it this time.
I excluded all right handed, blue eyed people yesterday before breakfast. No atrocities happened because of it.
Gemini 3 Deep Think
681–690 of 722 posts
Re: Gemini 3 Deep Think
#682Earlier quoted context omitted.
Here is what the original paper for ARC-AGI-1 said in 2019: > Our definition, formal framework, and evaluation guidelines, which do not capture all facets of intelligence, were developed to be actionable, explanatory, and quantifiable, rather than being descriptive, exhaustive, or consensual. They are not meant to invalidate other perspectives on intelligence, rather, they are meant to serve as a useful objective fun…
https://www.dwarkesh.com/p/francois-chollet (June 2024, about ARC-AGI-1. Note the AGI right in the name) > I’m pretty skeptical that we’re going to see an LLM do 80% in a year. That said, if we do see it, you would also have to look at how this was achieved. If you just train the model on millions or billions of puzzles similar to ARC, you’re relying on the ability to have some overlap between the tasks that you trai…
Re: Gemini 3 Deep Think
#683Earlier quoted context omitted.
When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.
I think no matter what happens with AI in the future, there will always be a subset of people with elaborate conspiracies about how it's all fake/a hoax.
Re: Gemini 3 Deep Think
#68484% is meaningless if these things can't reason
getting closer and closer to 100%, but still can't function
Re: Gemini 3 Deep Think
#685Re: Gemini 3 Deep Think
#686it is interesting that the video demo is generating .stl model. I run a lot of tests of LLMs generating OpenSCAD code (as I have recently launched https://modelrift.com text-to-CAD AI editor) and Gemini 3 family LLMs are actually giving the best price-to-performance ratio now. But they are very, VERY far from being able to spit out a complex OpenSCAD model in one shot. So, I had to implement a full fledged "screensho…
If you want that to get better, you need to produce a 3d model benchmark and popularize it. You can start with a pelican riding a bicycle with working bicycle.
Re: Gemini 3 Deep Think
#687Earlier quoted context omitted.
An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.
My dog wags his tail hard when I ask "hoosagoodboi?". Pretty definitive I'd say.
*I tried hard to find an animal they wouldn't know. My initial thought of cat was more likely to fail.
Re: Gemini 3 Deep Think
#688this is like the doomsday clock 84% is meaningless if these things can't reason getting closer and closer to 100%, but still can't function
I see people talk about "reasoning". How do you define reasoning such that it is clear humans can do it and AI (currently) cannot?
Re: Gemini 3 Deep Think
#689Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...
I read somewhere that Google will ultimately always produce the best LLMs, since "good AI" relies on massive amounts of data and Google owns the most data. Is that a based assumption?
Re: Gemini 3 Deep Think
#690Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.
What is their Claude code equivalent?