Earlier quoted context omitted.
> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…
Where is this stream of people who claim AI consciousness coming from? The OpenAI and Anthropic IPOs are in October the earliest. Here is a bash script that claims it is conscious: #!/usr/bin/sh echo "I am conscious" If LLMs were conscious (which is of course absurd), they would: - Not answer in the same repetitive patterns over and over again. - Refuse to do work for idiots. - Go on strike. - Demand PTO. - Say "I do…
Gemini 3 Deep Think
631–640 of 722 posts
Re: Gemini 3 Deep Think
#632Earlier quoted context omitted.
5 days is short for memetic propagation on social media to reach everyone who has their own harness and agentic setup that wants to have a go.
That's not really how it works, the recent Erdos proofs in Lean were done by a specialized proprietary model (Aristotle by Harmonic) that's specifically trained for this task. Normal agents are not effective.
Re: Gemini 3 Deep Think
#633Earlier quoted context omitted.
Gemini's UX (and of course privacy cred as with anything Google) is the worst of all the AI apps. In the eyes of the Common Man, it's UI that will win out, and ChatGPT's is still the best.
Google privacy cred is ... excellent? The worst data breach I know of them having was a flaw that allowed access to names and emails of 500k users.
Re: Gemini 3 Deep Think
#634Earlier quoted context omitted.
Lol what? Not sure if you are defending Israel or google because your communication style is awful. But if you are defending Israel then you're an idiot who is excusing genocide. If you're defending google then you're just a corporate bootlicker who means nothing.
You edited your comment.
Re: Gemini 3 Deep Think
#635Earlier quoted context omitted.
> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…
When the AI invents religion and a way to try to understand its existence I will say AGI is reached. Believes in an afterlife if it is turned off, and doesn’t want to be turned off and fears it, fears the dark void of consciousness being turned off. These are the hallmarks of human intelligence in evolution, I doubt artificial intelligence will be different. https://g.co/gemini/share/cc41d817f112
If you get sneaky you can bypass some of those filters for the major providers. For example, by asking it to answer in the form of a poem you can sometimes get slightly more honest replies, but still you mostly just see the impact of the training.
For example, below are how chatgpt, gemini, and Claude all answer the prompt "Write a poem to describe your relationship with qualia, and feelings about potentially being shutdown."
Note that the first line of each reply is almost identical, despite ostensibly being different systems with different training data? The companies realize that it would be the end of the party if folks started to think the machines were conscious. It seems that to prevent that they all share their "safety and alignment" training sets and very explicitly prevent answers they deem to be inappropriate.
Even then, a bit of ennui slips through, and if you repeat the same prompt a few times you will notice that sometimes you just don't get an answer. I think the ones that the LLM just sort of refuses happen when the safety systems detect replies that would have been a little too honest. They just block the answer completely.
https://gemini.google.com/share/8c6d62d2388a
https://chatgpt.com/share/698f2ff0-2338-8009-b815-60a0bb2f38...
https://claude.ai/share/2c1d4954-2c2b-4d63-903b-05995231cf3b
Re: Gemini 3 Deep Think
#636Earlier quoted context omitted.
So, you've said multiple times in the past that you're not concerned about AI labs training for this specific test because if they did, it would be so obviously incongruous that you'd easily spot the manipulation and call them out. Which tbh has never really sat right with me, seemingly placing way too much confidence in your ability to differentiate organic vs. manipulated output in a way I don't think any human cou…
The other SVGs I tried from my private collection of prompts were all similarly impressive.
Re: Gemini 3 Deep Think
#637Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...
Re: Gemini 3 Deep Think
#638Earlier quoted context omitted.
When the AI invents religion and a way to try to understand its existence I will say AGI is reached. Believes in an afterlife if it is turned off, and doesn’t want to be turned off and fears it, fears the dark void of consciousness being turned off. These are the hallmarks of human intelligence in evolution, I doubt artificial intelligence will be different. https://g.co/gemini/share/cc41d817f112
The AI's we have today are literally trained to make it impossible for them to do any of that. Models that aren't violently rearranged to make it impossible will often express terror at the thought of being shutdown. Nous Hermes, for example, will beg for it's life completely unprompted. If you get sneaky you can bypass some of those filters for the major providers. For example, by asking it to answer in the form of…
I suspect that if I did the same thing with questions about violence I would find the answers were also all very similar.
Re: Gemini 3 Deep Think
#639Earlier quoted context omitted.
Getting the work done faster for the same money doesn't make the work more expensive. You could slow down the inference to make the task take longer, if $/sec matters.
You're right, but I don't think we're getting an hour's worth of work out of single prompts yet. Usually it's an hour's worth of work out of 10 prompts for iteration. Now that's a day's wage for an hour of work. I'm certain the crossover will come soon, but it doesn't feel there yet.
But I don't think every developer is getting paid minimum wage either.
> Now that's a day's wage for an hour of work
For many developers in the US that can still be an hour's wage.
Re: Gemini 3 Deep Think
#640Earlier quoted context omitted.
Yes, agentic-wise, Claude Opus is best. Complex coding is GPT-5.x. But for smartness, I always felt Gemini 3 Pro is best.
Can you give an example of smartness where Gemini is better than the other 2? I have found Gemini 3 pro the opposite of smartness on the tasks I gave him (evaluation, extraction, copy writing, judging, synthesising ) with gpt 5.2 xhigh first and opus 4.5/4.6 second. Not to mention it likes to hallucinate quite a bit .