Live data from Hacker News

GLM-5.3-Flash

z.ai

591–600 of 605 posts

Re: GLM-5.3-Flash

#591

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training…

> Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc

No? A large model obviously can be dumb, I don't think you can infer much other than by testing it.

These small models are almost certainly worse at some things than the big models. They prize is making them dumber at things no-one cares about while retaining the capabilities people do care about. A model probably does not need to be able to give me a political treatise on the late 19th century "silver question" to be able to write me code.

Re: GLM-5.3-Flash

#592
Benchmarks and cost don't really help me understand how good a model actually is. Anyone have hands on experience working with the current newest models? What work did you do, and how did the model performance ?

Re: GLM-5.3-Flash

#593
post #588

Earlier quoted context omitted.

They innovated a lot.

Hardware constrains forced this?

Possible. But by looking at other industries Chinese don't seem to need to be forced to innovate. They just can and do. Unlike the West they seem to be on the way up and it seems like sky is the limit. In the West, the interests of the shareholders and other types of rent seekers seems to be the hard limit. Chinese have no qualms about making the cow obsolete before they milk it dry.

Re: GLM-5.3-Flash

#594
post #529
post #527

Earlier quoted context omitted.

Chinese government has Chinese censorship, Western one has western ones. You are not concluding what you think you do here.

Do you have an example of western AI censorship? It would help to give context.

The cybersecurity refusals, the biology refusals: https://blog.stephenturner.us/p/benchmarking-ai-biosecurity-... and the sexual activity refusals are all a form of censorship. Their purpose is the same as what the Chinese government would claim as being the reason to censor information on things like Tianenmen square.

Re: GLM-5.3-Flash

#595
post #248
post #185

Earlier quoted context omitted.

I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…

The one time I tried asking Sol to use subagents for a small project, it took a surprisingly long time, used up the entire usage limit in one go, and basically failed the project. I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickl…

I'm in the same boat. I haven't found sub agents flows useful.

Re: GLM-5.3-Flash

#596
post #102

Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is. What irks me about this is that the harnesses seem to be just an afterthought here. Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, e…

I've recently started PI, and after installing a couple extensions + writing a couple more it became enough for me. All I need is just a plan mode, and I installed that.

Re: GLM-5.3-Flash

#598
post #590

Earlier quoted context omitted.

Consider the reputation implications of them banning someone who can get their complaint about it to the front page of HN and into the YouTube drama loop. They’d get swarmed with activist cancellations. At most I suspect the A.I. providers will just come up with yellow banners like Anthropic did where naughty smut writers get put in the time out corner.

In the non China countries. But the Chinese populace would defend chinase companies unless they do something against China themselves

> In the non China countries.

That's a pretty big addressable market.

Re: GLM-5.3-Flash

#599
post #592

Benchmarks and cost don't really help me understand how good a model actually is. Anyone have hands on experience working with the current newest models? What work did you do, and how did the model performance ?

I have been using it since it was available in opencode Go plan.

It replaced all other models for me.

It's much less chatty than DeepSeek V4. And it feels noticeably smarter. I also use Claude at work and GLM 5.3 feels like Opus 4.8 (which I consider better than Opus 5).

It tends to be more proactive with suggestions after a task is done also.

DS4 Flash is awesome but GLM 5.3 is better despite being a bit more expensive.

Re: GLM-5.3-Flash

#600
post #301

Earlier quoted context omitted.

They only went up from 3000€ to 4000€ which isn't a lot. For comparison the cheapest Strix Halo 128GB went from 1600€ to 2600€ in the same timeframe.

So, literally the same 1k increase?

Yes! But, you know, percentages...
Post reply on HN