Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

51–60 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#51

Earlier quoted context omitted.

We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried?

Have you tried to use Fable for anything even remotely security related, when the refusals kick in as soon as you even fart in the vague direction of anything security or biology-adjacent?

For this comment to have value, you should indicate whether or not you applied for cybersecurity approval, and were approved or not.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#52

OpenAI and Anthropic need to just go ahead and give people access to the cyber models. Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.

At least OpenAI seems to want to do that, but the US is now forcing them to go through approvals. Anthropic seems much more hesitant.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#53
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize.

These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.

I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#55

Earlier quoted context omitted.

> I don't understand all this spite about "rich friends" Okay, here's a challenge: I assume you're not a rich and powerful entity, so try to gain access to Mythos. I'll wait. > I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it. Well, first I'd suggest they stop with the constant fear mongering. Here's my…

This is a lot of words to say "you're right, Anthropic does not have any legal way to release frontier cyber capabilities to the public"

Right, so according to you it's because of the US government that they don't release it to the public? Have you missed their constant and incessant fear mongering?

The causality chain here was not "US government says its dangerous -> Anthropic can't release it", it was "Anthropic is fear mongering -> US government listens to their fear mongering".

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#57
post #47

Earlier quoted context omitted.

Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.

Are those watermarks why claude suddenly started being even more unbearable to work with lately? Man. That would make a lot of sense indeed.

I'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation.

It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.

It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#58
post #38

Earlier quoted context omitted.

Realistically, you're looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it's just better to run DSv4 flash at full precision. 4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama

How fast are 2x or 4x DGX? I only have one and am wondering what the benefits are of getting another. I feel I will be disappointed…

If you can afford it, another DGX spark is worth it imo. Especially since, owning just one, you have a $1000 ConnectX7 card that's unused. You can find speeds here: https://spark-arena.com/leaderboard

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#59
post #49

Earlier quoted context omitted.

I know all this? Only a few corporations have Mythos because the US government is whitelisting them one at a time . Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried.

Before the US government had anything to do with this, Anthropic were fear mongering Mythos (BTW, Amodei also fear-mongered GPT-2, so this is a normal pattern in their operation) calling it "too dangerous to release", and back then only Anthropic was in charge of the whitelist. Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted.

Sorry but if you stepped back for a moment you'd realize this is all contrived nonsense to let to have your cake and eat it too.

No, Anthropic did not mind-game the US government into being worried about cybersecurity. The NSA has been paranoid about cyber controls for longer than you've been alive. If Anthropic had come out of the gate saying "no don't worry man, our model is TOTALLY COOL", while simultaneously attacking HAWK and finding core Linux vulnerabilities, I assure you the US government would have caught up about ten minutes later and we'd be in exactly the same spot minus your ability to tell Anthropic they were wearing the wrong dress and asking for it.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#60
post #7

> Scaling post-training is all we did for GLM-5.3. Love this opening line. And wow, great results. > As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.

What actually is "scaling post-training"?
Post reply on HN