Live data from Hacker News

GLM-5.3-Flash

z.ai

531–540 of 605 posts

Re: GLM-5.3-Flash

#531
post #100

Earlier quoted context omitted.

Isn't this practically every TOS though? Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it. HN's for example > We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time. > You acknowledge that Y Combinator may establish general practices and limits concerning use o…

> Isn't this practically every TOS though? Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs. > HN's for example You're not paying to use HN. Getting banned here has essentially zero consequences. If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant…

That's a TOS thing though. I don't think that works just as trivially in e.g. the EU, you can't just single-handedly stop a service someone is paying for.

Re: GLM-5.3-Flash

#532

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.

> silent sound fingerprinting mechanisms

I hope you also block the other half of the internet because sure as hell google and facebook are way worse than that.

Re: GLM-5.3-Flash

#533

Earlier quoted context omitted.

> In fact, I don't think I've ever even had a prompt refused. I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.

That's alarming. I want to use these models to red team my own computers. How are people getting around this?

In my experience some of these models may have learnt some censoring during distillation of Western models, but these are mostly just there as a probable response. So if they first happen to respond "I ain't doing this because legality", then you will have a hard time "convincing" it, but either rolling the dice again (so that it may not come up with the I can't do that text) or rewriting the conversation history a bit will get it going.

I sometimes just switch to a model I know is less smart to block stuff so that it has a text agreeing to do that, and then switch to a stronger model to actually go at the task.

Your mileage may vary though.

Re: GLM-5.3-Flash

#534
post #301

Earlier quoted context omitted.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

They only went up from 3000€ to 4000€ which isn't a lot. For comparison the cheapest Strix Halo 128GB went from 1600€ to 2600€ in the same timeframe.

So, literally the same 1k increase?

Re: GLM-5.3-Flash

#535

Earlier quoted context omitted.

The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed. The ‘frontier’ models rely on scale to achieve their results but that’s not t…

Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc. It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.

I think the assumption here that might not hold is simply that increases in efficiency and smaller size will be achieved by linearly just training smaller models better.

You are absolutely right that there is a physical limit about these things, but very often I find that the solution is a clever way to work around the problem. Maybe the problem with knowledge of the models will be improved by them looking the information up in a better way - so smaller models will not have to have the knowledge trained in but will default to checking. Maybe Models will, I dunno, focus on training in assembler and start to only ever check the compiled output so they only ever need to learn assembler and will then compile the solution to reason about the assembler code.

Obviously that last part is a ridiculous example because I'm not gonna be able to come up with a solution myself - I'm not nearly smart enough for that. But I h ope you get what I mean. Not going the direct route but instead finding solutions people didn't think of before.

Re: GLM-5.3-Flash

#536
post #419

Earlier quoted context omitted.

Yes Chinese models censor some historical events. This is nothing knew and well known thing. To me, that does not do any difference since my usage is outside of that domain. Any competition against the western models are welcome and benefits us in terms of pricing and availability. If they have to comply with CCP to be able to do it, then so be it. I have zero sympathy for Anthropic and OAI being so secretive and act…

What it shows is that the CCP has enough oversight and control (either explicitly or by the companies making these decisions by default) that they will alter the models to benefit China. Who is to say they aren't doing it in other ways as well? That they aren't, or won't be, subtly hamstrung in engineering work? OAI and Anthropic have their own issues, you're right to be suspicious of them, but it's not like their mo…

Did we all forget that the US administration has ultimate control over US AI models? That was just a few months ago

Re: GLM-5.3-Flash

#537
post #448
post #417

Earlier quoted context omitted.

> I'm sure they have nothing to rival this on a price/performance basis and have already given up on that How can you be sure about this? They have unbelievable capital. OpenAI is starting to preview its own chips, which could dramatically change the price/performance. We don't know what else Anthropic has cooked up right now that could rival this if they wanted to. Yes, others will _also_ continue to innovate, but m…

> They have unbelievable capital. all those Chinese labs are backed by the Chinese government which can just print money. time to wake up.

> the Chinese government which can just print money

I'm gonna let you in on a secret: it works the same in most countries in the world.

Re: GLM-5.3-Flash

#538

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

0 days later: Qwen 3.8 Flash Next: Let's cut GLM 5.3 Flash parmeters in half and active parameters to a third!

Chinese models had 94% reduction in parameters (from 2.8T/104B to 180B/6B) in 6 weeks, while staying close to the same quality.

Re: GLM-5.3-Flash

#539

Earlier quoted context omitted.

Deepseek and GLM answered correctly on Openrouter when using non-Chinese endpoints. I hope it stays that way!

correctly is very loaded here. correct according to whom? the truth is different though.

There‘s only one truth, the rest is interpretation. I was looking for the non-Chinese interpretation.

Re: GLM-5.3-Flash

#540

Earlier quoted context omitted.

Exactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.

Do they publish their distillation strategies on the private frontier models? Just curious.

I'm not sure if they publicly admitted to doing that. Would be interesting though.
Post reply on HN