Earlier quoted context omitted.
Do you actually believe this?
The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so? [1] https://www.washingtonpost.com/graphics/2020/world/national-...
GLM-5.3: Frontier coding with emergent cyber capabilities
281–290 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#282Earlier quoted context omitted.
The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.
I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this. Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volum…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#283This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
Fable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#284Earlier quoted context omitted.
There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.
Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.
Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which makes "banning" maybe not completely useless for certain things.
"Banning" things that can be easily copied over cyberweb transportation pipes feels like an fool's errand though, regardless of what it is. It's just too easy to get around, compared to actual physical items I suppose.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#285Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…
You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities.
Personally, I've had Fable 5, GPT 5.6 Sol, and GLM 5.2 all looking for correctness issues in an old abandoned WIP codebase of mine and all of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#286Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#287One htought I had; if The chinese allow unfettered access to cyber capabilties while th US does it's best to neuter it's model releases, from China's point of view they have the US all tied up in knots dealing with problems they don't give people the tools to solve. China giggles as it watches the US under threat from people using it's models. The US is restricting citizens from owning this particular kind of weapon,…
I suspect Anthropic wanted the US gov to ban Mythos for marketing. If it turns out to be bad for them, the US gov will likely suddenly unban models.
The threat is real.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#288Earlier quoted context omitted.
The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so? [1] https://www.washingtonpost.com/graphics/2020/world/national-...
I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.) I'm curious: to any professional vulnerability researchers reading this, what do you think?
Now i know that reality is worse than the worst conspiracy theorist.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#289Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the…
You're gonna have a had time getting straight answer to that out of the internet. There are now 4 different flavours of the Max plan floating around (Legacy V1, Legacy V2, New plans, and the current credit ones). And on top of that they have peak times. So ~8 scenarios, 24 in total across all feedback for their coding plans.
So when someone tells you they're having a good time on a GLM coding plan it's damn near unusable as a datapoint unless both parties are very clear about what precisely is being discussed
[It's been good for me though...V1 Max off peak...which is basically the best of the 24]
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#290This will be roughly on pair with Kimi K3, but using a third of its parameters. Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third. Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
GLM 5.2 has token efficiency problems. It's not a stupid model, but it takes a lot of "thinking" to produce not-stupid results. ("But wait...").
Which makes its pricing deceptive.
I tried to get by through the month of June on just GLM 5.2 and it was ... fine-ish for about two weeks. But the provider situation wasn't ideal.