Live data from Hacker News

GLM-5.3-Flash

z.ai

461–470 of 605 posts

Re: GLM-5.3-Flash

#461
post #420

Earlier quoted context omitted.

Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology. Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium t…

Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this: https://news.ycombinator.com/item?id=49413456 We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there.

It's crazy that sometimes I ask Opus 5 to explain what it just wrote to me, and it declares "that was word salad" (its words, not mine, without any hint from me other than "explain it").

Re: GLM-5.3-Flash

#462
post #405
post #349

Earlier quoted context omitted.

They claim it applies to EU citizens when both they and the service are outside the EU as well.

No they don’t. The GDPR is location scoped so the data of an EU citizen given to an hotel while on vacation in the US isn’t protected by GDPR but the data of an US citizen giving their data to a hotel in the EU while on vacation is.

That is my understanding as well. Not a lawyer but recently looked into it since we have US, EU, and non-EU European employees and we were looking at token usage tracking tools (which would appear to potentially fall under employee surveillance or at the very least require explicit consent)

Not sure the exact legalese but legal put together a consent form for anyone wanting to do the PoC

Re: GLM-5.3-Flash

#463

Earlier quoted context omitted.

The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed. The ‘frontier’ models rely on scale to achieve their results but that’s not t…

Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc. It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.

It doesn't really matter though. Hardware performance is still growing. The new Mac Studio could just about run this model locally (rather slowly) - something that sits on your desk, that you as a consumer can buy.

Imagine prosumer desktop hardware 10 years from now. The 2036 DGX Spark. For a few thousand dollars you will be able to buy something with hundreds of GB (maybe TB if manufacturers step up) of unified RAM, memory bandwidth in the 10-20TB/s range. Overall AI "compute" will increase 10-20x, while at the same time AI model capability per byte will increase 5-10x.

The hardware would fit today's models, something like Kimi K3, quite comfortably and give performance of maybe 100 tokens/second. So what needs data center hardware today will run on your desk.

But if we also assume the models become more efficient, a 2036 Fable-class model (in terms of intelligence/capabilities, not size) will easily run on this thing at hundreds of tokens per second.

Unfortunately it'll still slow to a crawl with 5 Chrome tabs open, and every Electron app will need at least 200GB of RAM.

Re: GLM-5.3-Flash

#464

Earlier quoted context omitted.

The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed. The ‘frontier’ models rely on scale to achieve their results but that’s not t…

Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc. It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.

Sounds like you are describing a quantized model which is a naive form of compression, not a model that is trained more efficiently.

Additionally the information theory angle is for information storage, but a model can access resources and tools to gain information and what we are really seeking to train is reasoning not information retrieval. We reduce the needs to the right capabilities and we don’t get upset if it does not know the lyrics to every song ever written.

Re: GLM-5.3-Flash

#467

Earlier quoted context omitted.

yeah i mostly agree, especially compared to subsidized subscription cost. But for a heavy user who has enough work to be done so that the box runs almost 24/7 at say 50tok/sec, the math gets interesting against API prices. And it can be interesting compared to subscription in the sense that you don't have the quota anymore. That means there's probably a lot of things you're not doing because of the quotas that you co…

At 50tps for single stream you are going to get 50 * 60 * 60 * 24 * 30 = 130M out tokens of GLM 5.3 Flash... That's less than what 40$ at current API rates... So if you are willing to pay 200$ per month you will get much better limits paying API rates. You can't run large Kimi K3 models on 10K worth of hardware either way, you need to spend like 50K USD minimum. Just pay for the API rates or get a low cost provider t…

Your point isn't lost on me, but a few other considerations:

1) Rates are theoretically discounted for GLM 5.3 Flash right now, by 50%.

2) Hardware costs have continued ascending with no sign of letting off, so it's unlikely that a DGX Spark depreciates to zero in one year.

3) Compare performance in terms of difficult tasks/$ over the last 6 months, 3 months, etc. Open weights are a ratchet. In terms of intelligence per $, a Spark is never going to be a worse deal tomorrow than it is today, at least until the entire platform is replaced or obsoleted.

71 days ago the best model you could run on two Sparks was an aggressive Q3 quant of Qwen 3.5 397B (AA 34). 70 days ago it was a mixed-quant of GLM 5.2 (AA 53). 30 days ago it was full fat DeepSeek 4 Flash (AA 53). Today it's GLM 5.3 Flash (AA57) and/or Qwen 3.8 Next (Unknown). Sometime this week it will likely become mixed-quant GLM 5.3 (AA 60).

So in < 80 days we have almost doubled the benchmark score. And that curve is still accelerating. If you view it as "cost per token of model vs API" then yes it's a bad deal. If you view it as "cost of task per $" then it has almost doubled in value in less than 3 months. All of this, imo, API and hardware, is still massively underpriced.

Re: GLM-5.3-Flash

#468
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it. https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...

[dead]

Re: GLM-5.3-Flash

#469

Earlier quoted context omitted.

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…

Yes Chinese models censor some historical events. This is nothing knew and well known thing. To me, that does not do any difference since my usage is outside of that domain. Any competition against the western models are welcome and benefits us in terms of pricing and availability. If they have to comply with CCP to be able to do it, then so be it. I have zero sympathy for Anthropic and OAI being so secretive and act…

In my experience, it’s their chat harnesses/website that filters historical events. The model itself doesn’t.

For example you can point opencode at DeepSeek v4 and ask, it will accurately tell you about Tiananmen square.

Re: GLM-5.3-Flash

#470

Earlier quoted context omitted.

> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of…

I have exactly the same opinion Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled…

This is really insightful, thank you.

Can you share a bit more about how you shifted to be more aligned with P/L? And how to accurately estimate incremental upside?

I'm an early PhD student with interest in ibdustrial research/R&D, and currently struggling to understand how to think about how to navigate through my career.

Post reply on HN