Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

191–200 of 1001 posts

Re: Claude Sonnet 4.6

#191
post #101

I'm a bit surprised it gets this question wrong (ChatGPT gets it right, even on instant). All the pre-reasoning models failed this question, but it's seemed solved since o1, and Sonnet 4.5 got it right. https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af This was sonnet 4.6 with extended thinking.

Off-by-one errors are one of the hardest problems in computer science.

That is not an off-by-one error in a computer science sense, nor is it "one of the hardest problems in computer science".

Re: Claude Sonnet 4.6

#192

The weirdest thing about this AI revolution is how smooth and continuous it is. If you look closely at differences between 4.6 and 4.5, it’s hard to see the subtle details. A year ago today, Sonnet 3.5 (new), was the newest model. A week later, Sonnet 3.7 would be released. Even 3.7 feels like ancient history! But in the gradient of 3.5 to 3.5 (new) to 3.7 to 4 to 4.1 to 4.5, I can’t think of one moment where I saw e…

If you've been using each new step is very noticeable and so have the mindshare. Around Sonnet 3.7 Claude Code-style coding became usable, and very quickly gained a lot of marketshare. Opus 4 could tackle significant more complexity. Opus 4.6 has been another noticable step up for me, suddenly I can let CC run significantly more independently, allowing multiple parallel agents where previously too much babysitting was required for that.

Re: Claude Sonnet 4.6

#193
post #28
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

simonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...

Can’t beat Gemini’s which was basically perfect.

Re: Claude Sonnet 4.6

#195

Just used Sonnet 4.6 to vibe code this top-down shooter browser game, and deployed it online quickly using Manus. Would love to hear feedback and suggestions from you all on how to improve it. Also, please post your high scores! https://apexgame-2g44xn9v.manus.space

That was fun, reminded me of some flash games I used to play. Got a bit boring after like level 6. It'd be nice to have different power-ups and upgrades. Maybe you had that at later levels, though!

Re: Claude Sonnet 4.6

#196

Earlier quoted context omitted.

An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…

> Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. That's why I have a functioning brain, to discern between ethical and unethical, among other things.

Yes, and most of us won’t break into other people’s houses, yet we really need locks.

Re: Claude Sonnet 4.6

#197
post #101

Earlier quoted context omitted.

Off-by-one errors are one of the hardest problems in computer science.

That is not an off-by-one error in a computer science sense, nor is it "one of the hardest problems in computer science".

This was in reference to a well-known joke, see here: https://martinfowler.com/bliki/TwoHardThings.html

Re: Claude Sonnet 4.6

#198

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…

>Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question.

Thanks for the successful pitch. I am seriously considering them now.

Re: Claude Sonnet 4.6

#199

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

idk, codex 5.3 frankly kicks opus 4.6 ass IMO... opus i can use for about 30 min - codex i can run almost without any break

Re: Claude Sonnet 4.6

#200

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

The funny thing is that Anthropic is the only lab without an open source model

They are, at the same time I considered their model more specialized than everyone trying to make a general purpose model.

I would only use it for certain things, and I guess others are finding that useful too.

Post reply on HN