Earlier quoted context omitted.
GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…
If a good SWE is $150/hour, does the model cost actually matter? Surely you'd be willing to spend $10/hour to make that SWE 20% more productive? The model cost is still much less than the salary.
GLM 5.2 beats Claude in our benchmarks
361–370 of 559 posts
Re: GLM 5.2 beats Claude in our benchmarks
#362Earlier quoted context omitted.
If a good SWE is $150/hour, does the model cost actually matter? Surely you'd be willing to spend $10/hour to make that SWE 20% more productive? The model cost is still much less than the salary.
I don’t think any engineers who cost $150/hr are having their productivity moved by 20% depending on a $10/hr gap between models on or near the frontier. Most of the gains right now come from tooling and process and any big post 2025 language model. The specific model isn’t that important right now.
Re: GLM 5.2 beats Claude in our benchmarks
#363Earlier quoted context omitted.
Why are you spending on API for GPT coding instead of stacking 20x subs and using codex-lb?
Company pays API prices so we can use daily the best model for our job without being locked in. Also the team subscriptions started to be more like X per seat + usage...
I understand the reasons to use team/enterprise accounts, but apart from the policy/management/billing side of it, I still don't understand the value in spending thousands for API instead of hundreds - even when there's argument that one provider is better than another depending on the use case, I don't think that credibly extends much beyond OpenAI + Anthropic frontiers, which both have $200 subs you can stack.
Re: GLM 5.2 beats Claude in our benchmarks
#364Re: GLM 5.2 beats Claude in our benchmarks
#365Re: GLM 5.2 beats Claude in our benchmarks
#366Re: GLM 5.2 beats Claude in our benchmarks
#367Earlier quoted context omitted.
GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…
> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way mo…
One group is trying to get the LLM to basically one shot everything and not properly reviewing the output.
Others are using the LLM to assist their human intelligence in a tight loop.
If you’re doing the former you really do need the best model available because that’s still right on the edge of what LLMs can do at best, and at worst you’re just shipping pure unmaintainable slop.
If you’re doing the latter then you can get away with a slightly less powerful model without it making a material difference because your human intelligence is filling in gaps
Re: GLM 5.2 beats Claude in our benchmarks
#368It’s hard to argue against the open weight models if your only concern is coding. Which, for many of us hackers here in this forum, it is. But I would like to point out that the overwhelming majority of people using LLMs aren’t programmers, don’t care about coding, and couldn’t even be bothered to “vibe code”. So we should consider the bias of the output of these open weight models, and what that looks like, outside…
There is no money made from these people though .. people who are using ChatGPT to plan for their next week-end or their next vacation aren't paying a $100 or $200 monthly subscription. As for non coder office workers (accountants, PMs, etc.), they use Microsoft or Google products which all integrate AI to some extent within their products - with RAG for Sharepoint to some basic AIs to generate text or automate work…
I don’t agree with “Software development is where money is made for these labs”. Coders will inevitably eat up the most tokens & buy the bigger $200 subscriptions because we want to keep working.
But us coders are still the small minority of users. They aren’t counting on us to get to trillion dollar evaluations.
They are counting on the regular folks to buy the $20/ month subscription. It’s really easy to run out your free tier usage these days, asking questions that have nothing to do with coding.
So my point is what does that output look like for someone asking a question about politics or world news?
Re: GLM 5.2 beats Claude in our benchmarks
#369So one can see businesses owning their own such cluster, next to their database infra, in the near future.