Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

531–540 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#531

> beats Claude in our Cyber Benchmarks Beats which model in Claude? Whenever a "benchmark" doesn't put precise model numbers in their headlines I am immediately skeptical. Either they don't know the difference (bad) or they are benchmarking against weaker models (misleading, also bad). It's like when studies say "AI is bad at X" and they used GPT-3.5 in current year.

Anthropic's own models perform differently under the same version depending on how much they've decided to quietly downgrade them.

Re: GLM 5.2 beats Claude in our benchmarks

#532

Earlier quoted context omitted.

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

Something I don't see in your charts is acknowledgement of the difference, sometimes paradoxical, in strength between the same model at different reasoning levels. Do you have charts that include low/med/high/xhigh/max for the various models?

This is something we omit for a few reasons but it's probably the biggest blind spot in our evaluations; we opt-in to auto-reasoning/adaptive reasoning or max thinking token budgets where supported (supported by most models now), but when an explicit reasoning level is required, we fall back to High reasoning. In practice, we've found most models scale High-> pretty consistently, but if one vendor started throwing 10x the resources into max reasoning and they didn't support auto-reasoning, they would be unfairly penalized in our evaluations.

Re: GLM 5.2 beats Claude in our benchmarks

#533

Earlier quoted context omitted.

You'd need to multiply that $10k by 8 minimum.

Why? 4 x DGX sparks should be enough. That's way less than $80k.

It would run so slow it would be functionally unusable or at such a low quant that it wouldn't be useful for serious work.

Re: GLM 5.2 beats Claude in our benchmarks

#535

Earlier quoted context omitted.

That’s a uniquely US issue - in NZ you can get a 100A single phase at 230V nominal without any issue. 23kw, straight to your door. A single circuit using 10mm TPS would technically be enough to run what you’re describing. Might be pricey though, I’d probably take the excuse to get 3 phase installed so I could get access to the stock of used 3 phase machinery.

> That’s a uniquely US issue - in NZ you can get a 100A single phase at 230V nominal without any issue. 23kw, straight to your door In the US it's common to get 200A 120/240V split-phase service. We're talking about the wiring inside the house, though. How do you think everyone here is charging their electric cars at home and running our AC and electric cooktops at the same time if we didn't also have that? :) You ne…

So where does the necessity to install multiple 15A circuits come from?

Re: GLM 5.2 beats Claude in our benchmarks

#536
post #526
post #413

Earlier quoted context omitted.

Big fan of Amp but pretty sure it only uses Flash for search: https://ampcode.com/models As for Fable: I used it as much as I could while we had it. It was a step change over Opus with my work.

Or maybe it was supposed to be the OMP (OhMy Pi) harness. Pi can do just about anything for you. Use most models in most ways possible. You just tell Pi what you want, and it builds an extension for itself.

I guess? I think most people just call it Pi though.

Re: GLM 5.2 beats Claude in our benchmarks

#537

Earlier quoted context omitted.

Hi Scott! Was just considering signing up, NW looks great (fp8 GLM 5.2 is good !) Standard cached token pricing for GLM 5.2 is pretty high, I'm wondering whether the KV cache for that model actually is that expensive to serve on average, or if Neuralwatt's energy pricing for long-running GLM 5.2 agents is especially competitive? The live energy stats don't break down by token type, would love to see that. And 2/3 of…

Thanks for the feedback! Our primary focus is charging by energy, for token pricing we really just try to be close to the market. That being said I'll take a look at our token pricing to see if we need an update there https://portal.neuralwatt.com/energy-pricing Generally our users get much lower cost on energy than token pricing though on a typical request with a high prefix cache hit the input, cached costs is very…

[dead]

Re: GLM 5.2 beats Claude in our benchmarks

#538

Earlier quoted context omitted.

Only you get things done lost faster and don't need to pay entire years salary?

How so? If you’re doing a one-off project, sure. But if you’re coding like this full time, you’re paying the year’s salary anyway. And faster? Not so sure about that. Sure they can write code faster, but writing code is a small part of building something.

That's a good point.. Writing code vs building

Re: GLM 5.2 beats Claude in our benchmarks

#539

Earlier quoted context omitted.

They postponed that change, here is the email they sent out: > In May, we sent you an email announcing that starting today, the Claude Agent SDK, claude -p, and third-party apps built on the Agent SDK would stop drawing from subscription rate limits and move to a dedicated monthly credit. We're writing to let you know that we’re not making this change today. We’re working to update the plan to better support how user…

Something I haven't been able to figure out.... How are you supposed to actually get an API key to use quota from your subscription? The terms of service still forbid using OAuth authentication and the API keys from the console indicate that you need to pre-load your account with funds when you try to use them.

This is for using claude -p, which is claude code, and you authenticate by running /login.

You can't get an API key for subscription based usage of claude, because you are supposed to use claude code only.

Re: GLM 5.2 beats Claude in our benchmarks

#540

Earlier quoted context omitted.

And just like Linux lost to Windows in consumer market due to devs/creator's stubbornness, same will happen with closed vs open LLM. In the end the one that is used the most will be the one that you train your kids on and therefore the one that wins the market. Eventually the closed one with too much guardrail will be left behind because people will stop using it. You need to read the market. Linus didn't read it in…

Is this 2006? Linux is present on literal billions of android phones, servers, supercomputers and other embedded devices. It's the most ubiquitous OS on the planet and it's not even close, even Microsoft contributes to it. The only niche where it doesn't utterly dwarf the competition is personal computers and it looks like we're all getting priced out of that anyway

I was 100% sure that somebody will throw this, but I didn't actually expect to do it from a throwaway account. Maybe because you know calling Android Linux is like calling a human just an ape (but mirrored because you know, Android in this case is the ape). Oh, and Microsoft Loves Linux, right? Because that's why they invented WSL, to make people go and use Linux, right? riight!!
Post reply on HN