Live data from Hacker News

Claude Sonnet 5

anthropic.com

141–150 of 822 posts

Re: Claude Sonnet 5

#142

Earlier quoted context omitted.

Why do you think they are bragging? Anthropic has long been the company to give us by far the most in-depth information about their models, both positive and negative. I read this as them just stating a fact about this model that users would want to know.

I'm absolutely certain that their marketing team has input on (if not owning) these announcements.

I find it interesting how two different directly opposed messages seem to have both been interpreted as being nothing but marketing speak.

Re: Claude Sonnet 5

#143
post #123
post #68

Earlier quoted context omitted.

I've been largely disappointed how much the Claude models ignore custom instructions, and sometimes even prompts on the chat interface. It sometimes feels like talking to a wall, or as if there was a third person in the chatroom whose messages I can't see. I can't help but feel this is intentional towards the 'Agentic' workflow.

> or as if there was a third person in the chatroom whose messages I can't see. If you set off a classifier, that's how it looks to Claude.

I wasn't working with anything sensitive, but it really does feel like it sometimes condenses even something low like three bullet points to two.

IMO, they were quite good with checklists even a year ago, and tried to tick off each one.

Re: Claude Sonnet 5

#144
post #109

Earlier quoted context omitted.

Totally agreed. I sometimes wonder if they are making the model "lazy" with each iteration, it keeps getting better at avoiding work.

This is why Fable was so good. It followed instructions and it was in no way lazy.

I've been seeing LLMs act lazy from the very beginning. They got a little better but smaller models really only want to have a single task given to them. Mythos at least does work. RIP

Re: Claude Sonnet 5

#145

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

They are obviously trying to avoid getting Sonnet 5 blocked.

Re: Claude Sonnet 5

#146

Earlier quoted context omitted.

"Lower ability to perform cybersecurity-related tasks" makes me super concerned it will leave my codebase like Swiss cheese for any American granny with access to Fable 5, when we non-American Brits, or rest-of-worlders, don't have access to it to clean our codebases.

That’s not even close to true. Unless you’re vibe coding trash that a better model might catch.

I don't think so. During the time I was using Fable 5, I was getting it to clean security bugs that Opus 4.8 had introduced ... bugs which weren't localised to a single PHP file but were caused by cascading data flow through multiple PHP files. I'm not an expert on security but I know I wouldn't have found these myself. I knew from day one of Fable's release that it would do thorough security audits and fix loads of flaws, even offering up PoCs to help show that it fixed them, as long as I didn't explicitly ask it to do a security audit. I just said, "My codebase is a mess," and it went on for an hour doing a thorough security audit and helping plug numerous holes. This was before the "fix my code" story came out.

Re: Claude Sonnet 5

#147

Earlier quoted context omitted.

"Lower ability to perform cybersecurity-related tasks" makes me super concerned it will leave my codebase like Swiss cheese for any American granny with access to Fable 5, when we non-American Brits, or rest-of-worlders, don't have access to it to clean our codebases.

I think they don’t understand that cybersecurity skills are what prevent bad code from making it into production. It’s like telling a chef to cook without a knife because knives can kill people. Dario and his lackeys at Anthropic aren’t visionaries.

I think you misunderstood what their vision is, or rather what their possible futures are. They are many steps ahead of almost everyone, both in wargaming possibilities and the actual realized path. What doesn’t make sense to you may be the only safe option for them.

Re: Claude Sonnet 5

#148

Is it just me or is there a huge difference between how much one can accomplish in a 5-hour window with GPT 5.5 on xhigh versus any Claude model?

I exclusively use 5.5-xhigh-fast within Codex and find it superior to Opus 4.8.

Re: Claude Sonnet 5

#150
I'd love if they would include speed (though I know there are difficulties involved). At this point the quality of Opus 4.8 is no longer my limiting factor, it's the speed, so a faster model would be great.
Post reply on HN