Live data from Hacker News

Claude Sonnet 5

anthropic.com

541–550 of 822 posts

Re: Claude Sonnet 5

#541

Earlier quoted context omitted.

I'm biased because I run an inference company, https://synthetic.new . That being said I think we're pretty good at serving at GLM-5.2 — and other models, like Kimi K2.7! — and our privacy policy is quite good: zero data retention for prompts and completions on API requests. Our average streaming TPS for GLM-5.2 (aka, tokens after factoring out time-to-first-token, which varies based on geography) is 97tps over the l…

got a 500 error page on the site's chat, but I'll try the API

Interesting: I don't see anything in our error logs but we could be missing something (and personally the chat works for me + my unsubscribed test account). If you email us at hi@synthetic.new though we should be able to fix anything you're running into!

Re: Claude Sonnet 5

#542

Earlier quoted context omitted.

Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me

This guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.

I get similar results in my own tests. And Gemini 3.1 Pro is consistently on top of my ratings. Not everyone is coding monkey, I prefer staying a programmer.

Re: Claude Sonnet 5

#543
post #395

Earlier quoted context omitted.

As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.

What’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins

Run it on Amazon Bedrock or GCP vertex. No problems at all.

Re: Claude Sonnet 5

#545

Earlier quoted context omitted.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

> There’s no way to justify their valuations if they get downgraded to a pair programming tool. Honestly I still don't see how they justify their valuations, period. If anything they're serious liabilities. Open-weight models are improving and reaching "good enough" levels for more and more tasks. They're also known quantities; you know what you're getting with them and don't have to worry about the model silently (o…

[deleted]

Re: Claude Sonnet 5

#547

Earlier quoted context omitted.

Okay I don’t care about “eventually”, I want Fable now.

Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?

What if you're using it for Mathematics (e.g., making progress on unsolved problems) instead of writing software? Would you consider that a valid use-case?

Re: Claude Sonnet 5

#548

Earlier quoted context omitted.

This guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.

I get similar results in my own tests. And Gemini 3.1 Pro is consistently on top of my ratings. Not everyone is coding monkey, I prefer staying a programmer.

They're referencing Gemini 3.5 Flash being the top model, you must not be great with detail.

And no (strong) programmer would jump to assuming other people are coding monkeys just because they disagree on what a strong LLM is: that's the kind of thinking reserved for the glorified coding monkeys who wasted their life getting better at writing CRUD apps and are now upset that someone's tooling is dropping the already very low bar there.

Re: Claude Sonnet 5

#549

Earlier quoted context omitted.

They're actively trying to use lobbying power to make open weight models illegal. So I'm just not going to use their services at all anymore. I don't think they're a net gain if you're a skilled senior, and the hidden cost in terms of technical debt and skill atrophy is just being swept under the rug. I'll be okay without their bullshit generator.

Sure about Dario (and all billionaire) weirdness, but no gains if you are a skilled senior is well, very far out in our experience (our company is 30 years old with mostly the original employees and founders): what we deliver now at the speed and quality we deliver it would have been impossible 10 years ago with our team size of skilled seniors. We replaced all the commercial products our clients and ourselves used w…

> We replaced all the commercial products our clients and ourselves used with our own

You’ll never guess what product your clients are looking to replace with their own next.

Post reply on HN