[flagged]
Claude Sonnet 4.6
741–750 of 1001 posts
Re: Claude Sonnet 4.6
#742Re: Claude Sonnet 4.6
#743Earlier quoted context omitted.
You can always make it go back and forth with "Are you sure?". The fact that these are still issues ~6 years into this tech is bewildering.
...is it though? Fundamentally, these are statistical models with harnesses that try to conform them to deterministic expectations via narrow goal massaging. They're not improving on the underlying technology. Just iterating on the massaging and perhaps improved data accuracy, if at all. It's still a mishmash of code and cribbed scifi stories. So, of course it's going to hit loops because it's not fundamentally consc…
> So, of course it's going to hit loops because it's not fundamentally conscience.
Wait, I was told that these are superintelligent agents with sophisticated reasoning skills, and that AGI is either here or right around the corner. Are you saying that's wrong?
Surely they can answer a simple question correctly. Just look at their ARC-AGI scores, and all the other benchmarks!
Re: Claude Sonnet 4.6
#744Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…
“Walk. It’s 50 meters—a 30-second stroll. Driving that distance to a car wash would be slightly absurd, and you’d presumably need to drive back anyway. “
Opus 4.6 nailed it: “Drive. You’re going to a car wash. ”
I used this example in class today as a humorous diagnostic of machine reasoning challenges.
Re: Claude Sonnet 4.6
#745Re: Claude Sonnet 4.6
#746I ran the same test I ran on Opus 4.6: feeding it my whole personal collection of ~900 poems which spans ~16 years It is a far cry from Opus 4.6. Opus 4.6 was (is!) a giant leap, the largest since Gemini 2.5 pro. Didn't hallucinate anything and produced honestly mind-blowing analyses of the collection as a whole. It was a clear leap forward. Sonnet 4.6 feels like an evolution of whatever the previous models were doin…
Re: Claude Sonnet 4.6
#747Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…
My answer was (for which it did zero thinking and answered near-instantaneously): "Drive. You're going there to use water and machinery that require the car to be present. The question answers itself." I tried it 3 more times with extended thinking explicitly off: "Drive. You're going to a car wash." "Drive. You're washing the car, not yourself." "Drive. You're washing the car — it needs to be there." Guess they're s…
Re: Claude Sonnet 4.6
#748They use the word "Sonnet" 60+ times on that page but never give the casual reader any context of what a "Sonnet model" actually is. Neither does their landing page. You have to scroll all the way to the footer to find a link under the "Models" section. You click it and you finally get the description "Hybrid reasoning model with superior intelligence for agents, featuring a 1M context window" You then compare that t…
"Sonnet" only makes sense relative to other things but not by itself. If you don't know those other things, it is difficult to understand.
But, if you were asking (and I'm not sure that you are): "Sonnet 4.6 is a cheaper, but worse, version of Opus 4.6 which itself is like GPT-5.3 Codex with Thinking High. Making Sonnet 4.6 like a ChatGPT 5.3 Thinking Standard model."
Re: Claude Sonnet 4.6
#749[flagged]
Re: Claude Sonnet 4.6
#750They use the word "Sonnet" 60+ times on that page but never give the casual reader any context of what a "Sonnet model" actually is. Neither does their landing page. You have to scroll all the way to the footer to find a link under the "Models" section. You click it and you finally get the description "Hybrid reasoning model with superior intelligence for agents, featuring a 1M context window" You then compare that t…