Live data from Hacker News

Claude Sonnet 5

anthropic.com

321–330 of 822 posts

Re: Claude Sonnet 5

#321
post #157

Earlier quoted context omitted.

Where is gpt 5.6?

Victim of the same hype generated by Dario. Now everyone has to walk on eggshells, do limited releases to trusted partners, and nerf their cybersecurity capabilities lest they get deemed “too powerful to release”.

Yeah and our government is continuing to take pages from China's playbook for the last fucking decade... and not the plays that work.

Re: Claude Sonnet 5

#322
post #73

Earlier quoted context omitted.

They will release it eventually. Once they see the Chinese models are close to Mythos level they will release it before, so it will be "revolutionary".

It was already released. US government is the only reason it's not available to us mere mortals anymore

Obviously I meant released for public use.

Re: Claude Sonnet 5

#323

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

if you like that, use gpt models instead.

Re: Claude Sonnet 5

#325
post #253

Earlier quoted context omitted.

Some napkin math -- total global labor compensation is about 50% of the GDP, which puts it in the USD 50 - 60 Trillion range: https://ourworldindata.org/grapher/labor-share-of-gdp This source claims that knowledge workers alone (probably because they are paid much more) account for 35 - 50 Trillion of that: https://github.com/danielmiessler/Substrate/blob/main/Data/K... If LLMs can boost their productivity even by an…

I am deeply surprised by the silence of philosophers, sociologists, liberal arts majors, economists. Where are the think tanks who contemplate and debate the societal aspects? The tech is advancing full steam but the "other side" doesn't feel anywhere nearly ready.

Reid Blackmun has written several books and has a consultanting agency to guide ethical implementation of AI

Re: Claude Sonnet 5

#326

Earlier quoted context omitted.

I was surprised to learn that Sonnet generally has the same tokens per second as Opus

I would indeed be more inclined to use it if the tokens per second were better. Though I would be then using their more expensive Opus less though. Perhaps it is strategy.

They should add a Sonnet 5 fast mode at ~Opus pricing

Re: Claude Sonnet 5

#327

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

“Hey I saw some messed up function commented out that at face value is a bad idea… so here it is again with some nonsense assumptions ….”

I ask “where did you get that?” … too often if I’m not constantly guiding it, and even then it still goes off the rails.

Re: Claude Sonnet 5

#328

Earlier quoted context omitted.

Which is your own harness and your own evals for your tasks I guess

Maybe. But that sounds like a large amount of bespoke work for what seems like a common problem?

I was talking about enterprise agents and then realized the question is more about coding agents.

Re: Claude Sonnet 5

#329

Earlier quoted context omitted.

This is code for "this model can't be used to hack other systems as effectively as Opus or Mythos."

"dangerous cyber skills, such as developing software exploits" is very plainly referring to the same thing you are, but is more precise industry terminology rather than the loaded slang "hack".

I was referring to "Lower ability to perform cybersecurity-related tasks," which is newspeak for hacking.

Re: Claude Sonnet 5

#330
I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https://revise.io/errata-bench
Post reply on HN