Live data from Hacker News

Claude Sonnet 5

anthropic.com

731–740 of 822 posts

Re: Claude Sonnet 5

#731
post #729
post #727

Earlier quoted context omitted.

Try it 25 more times and let us know how it averages out. It's non-deterministic, remember?

I tried 3 more times. Two were nearly identical and 1 recognized Latitude 41 as a restaurant but had a similar useless reply

"Latitude 41" is a business name?

If you gave me this prompt, I'd say "Which Columbus? None of the ones I know about are at a latitude of 41 degrees north or south?"

Re: Claude Sonnet 5

#732
post #731
post #729

Earlier quoted context omitted.

I tried 3 more times. Two were nearly identical and 1 recognized Latitude 41 as a restaurant but had a similar useless reply

"Latitude 41" is a business name? If you gave me this prompt, I'd say "Which Columbus? None of the ones I know about are at a latitude of 41 degrees north or south?"

It's an interesting question to ask a machine. It's definitely ambiguous. As a person I might respond "What's Latitude 41"?

Re: Claude Sonnet 5

#733
post #731
post #729

Earlier quoted context omitted.

I tried 3 more times. Two were nearly identical and 1 recognized Latitude 41 as a restaurant but had a similar useless reply

"Latitude 41" is a business name? If you gave me this prompt, I'd say "Which Columbus? None of the ones I know about are at a latitude of 41 degrees north or south?"

That is fine and would have even been a better reply. I connected to a VPN in Belgium, opened a private browsing window, and Yandex, Google, and DuckDuckGo all find "latitude 41 in Columbus" as the first search result.

It figured out I was looking for a place but completely failed to identify I already provided proximity information and asked for proximity information in its reply.

Re: Claude Sonnet 5

#734

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

My experience with Opus in the last weeks is the opposite. I have the feeling Opus got smarter since they released and blocked Fable. Maybe they got more compute available since a) they finished Training Mythos/Fable and b) couldn't provide inference for it?

Re: Claude Sonnet 5

#735

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

I would use Sonnet instead of Opus because it's faster. Isn't it? It's a smaller model

Re: Claude Sonnet 5

#736
post #560

Earlier quoted context omitted.

I appreciate the suggestion! But it isn't clear to me, from reading their marketing site, what they bring to the table from this perspective. Can you give me a more targeted pitch?

This page buried in their docs is a bit better than the homepage imo: https://ampcode.com/manual#why-amp I haven't used them in a while so my info may be out of date, but they tended to track whatever models were the best and auto-use them for each task (eg, one for planning, subagent for a code search, other frontier for implementing). Their CLI seemed very well thought out to make you do things "the correct way" --…

Thanks!

Re: Claude Sonnet 5

#737

Earlier quoted context omitted.

I appreciate the suggestion! But it isn't clear to me, from reading their marketing site, what they bring to the table from this perspective. Can you give me a more targeted pitch?

Sorry for the late answer and the missing context. usef- is right, the manual is probably the better page to share. Amp tries to give you a plug and play experience, where you can always see the actual costs and models/effort are autoselected for you. Some of my colleagues are big fans and use it a lot. I also like it, but prefer OpenCode.

Great. Amp wasn't even on my radar! Appreciate the pointer!

Re: Claude Sonnet 5

#738
What's interesting is that Claude Sonnet 5 costs more per task ($2.29) than Opus 4.8 ($1.80), while the latter is obviously better!

It actually costs more per task than every other model. It's only cheaper than Claude Fable 5.

Source: https://artificialanalysis.ai/?cost=cost-per-task#price-and-..., as of writing this comment (the results are frequently changing)

Re: Claude Sonnet 5

#739
post #132

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am n…

We have found similar when plugging GLM 5.2 into actual benchmarks in our product. The open-source models are really dialled into the public benchmarks, until you try them in context you won't have a solid idea of how they perform (Sonnet is a higher quality model than 5.2, both in prose, reasoning, and alignment).

Re: Claude Sonnet 5

#740
post #692

Earlier quoted context omitted.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

> I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't think so. Expect that in a market with high vendor lock-in but that's not the case here. The market is extremely competitive and switching cost are near zero. Anthropic can't afford to pull shit like this and sacrifice quality.

My employer just finalized a contract with Anthropic, for enterprise Claude Code use. Which means that unless there is a _major_ downgrade in service quality, we are now locked in for the next few years (but at least for a year, although vendor contracts are renegotiated less frequently).

Just checked the dashboard, and we seem to have the exact same $200 credit as others, enterprise or not. Token inflation affects us just like everyone else.

It feels a bit like buying the same box of chocolates every day, but the size / weight of the box is shrinking... the price remains unchanged!

Post reply on HN