Earlier quoted context omitted.
I'm biased because I run an inference company, https://synthetic.new . That being said I think we're pretty good at serving at GLM-5.2 — and other models, like Kimi K2.7! — and our privacy policy is quite good: zero data retention for prompts and completions on API requests. Our average streaming TPS for GLM-5.2 (aka, tokens after factoring out time-to-first-token, which varies based on geography) is 97tps over the l…
got a 500 error page on the site's chat, but I'll try the API
Claude Sonnet 5
541–550 of 822 posts
Re: Claude Sonnet 5
#542Earlier quoted context omitted.
Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
This guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.
Re: Claude Sonnet 5
#543Earlier quoted context omitted.
As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
What’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins
Re: Claude Sonnet 5
#544Re: Claude Sonnet 5
#545Earlier quoted context omitted.
There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…
> There’s no way to justify their valuations if they get downgraded to a pair programming tool. Honestly I still don't see how they justify their valuations, period. If anything they're serious liabilities. Open-weight models are improving and reaching "good enough" levels for more and more tasks. They're also known quantities; you know what you're getting with them and don't have to worry about the model silently (o…
Re: Claude Sonnet 5
#546Re: Claude Sonnet 5
#547Earlier quoted context omitted.
Okay I don’t care about “eventually”, I want Fable now.
Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?
Re: Claude Sonnet 5
#548Earlier quoted context omitted.
This guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.
I get similar results in my own tests. And Gemini 3.1 Pro is consistently on top of my ratings. Not everyone is coding monkey, I prefer staying a programmer.
And no (strong) programmer would jump to assuming other people are coding monkeys just because they disagree on what a strong LLM is: that's the kind of thinking reserved for the glorified coding monkeys who wasted their life getting better at writing CRUD apps and are now upset that someone's tooling is dropping the already very low bar there.
Re: Claude Sonnet 5
#549Earlier quoted context omitted.
They're actively trying to use lobbying power to make open weight models illegal. So I'm just not going to use their services at all anymore. I don't think they're a net gain if you're a skilled senior, and the hidden cost in terms of technical debt and skill atrophy is just being swept under the rug. I'll be okay without their bullshit generator.
Sure about Dario (and all billionaire) weirdness, but no gains if you are a skilled senior is well, very far out in our experience (our company is 30 years old with mostly the original employees and founders): what we deliver now at the speed and quality we deliver it would have been impossible 10 years ago with our team size of skilled seniors. We replaced all the commercial products our clients and ourselves used w…
You’ll never guess what product your clients are looking to replace with their own next.