Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

241–250 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#241
post #127
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

Fable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet.

> At this point, Anthropic only needs to release models to the public when the competition forces them to.

Assuming the government allows them to lol

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#242
post #138

Earlier quoted context omitted.

We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried?

Why should I apply for *cybersecurity* approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, "programming") somehow is cybersecurity now?

Your tools refusing to do your bidding is an absurd idea in the first place

Imagine asking for permission to use your hammer

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#243
This will be roughly on pair with Kimi K3, but using a third of its parameters.

Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.

Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#244

Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.

The RAM shortage situation won’t be sorted out within the next year.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#245

Earlier quoted context omitted.

That makes them at best temporary middlemen. It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't). Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.

The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.

If you have exponentially increasing use of your harness, then it's true that every day you capture exponentially more data, but it's also true that every day exponentially more data will slip through the cracks of your would-be monopoly and that data arrives at your competitors via various channels (competitor harnesses, subsidized reselling, etc)

The very exponential that you are relying on to give you runaway improvement is also giving exponentially increasing data to your competitors. All else being equal your competitors stay a step behind but you never develop a monopoly either. That's the best case for Anthropic/OpenAI. In reality, training data is just one variable, exponentials don't last forever, and your competitors will get better at capturing a bigger slice of training data.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#246
post #181

Earlier quoted context omitted.

Interesting... So Chinese models are not so bad ?

There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.

How does banning the models in the US prevent this?

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#248
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

> and Anthropic's Project Glasswing is supposed to find them quite a while ago? We cannot trust a single company to report security issues, it’s good to see competition in that domain

Open source competition no less.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#249
post #135

Such an interesting times we are in, We just had amazing releases this past two months kimi k3, glm5.3 qwen3.8 and now glm5.3 These open models are getting really good

You wrote GLM5.3 two times :)

Because it's doubly good.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#250

Earlier quoted context omitted.

Not exciting for anyone directly or indirectly invested in a frontier lab or its partners. And that is a lot of people, including you.

The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.

I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this.

Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volume they need to sell.

Finally, there is a data wall. Sure, they can keep scaling RL on math problems and code. But with everything else, where will the supervision come from when they need several orders of magnitude more?

Post reply on HN