Live data from Hacker News

GLM-5.3-Flash

z.ai

471–480 of 605 posts

Re: GLM-5.3-Flash

#471

Earlier quoted context omitted.

https://github.com/bethington/ghidra-mcp . Works flawlessly.

Ha - I saw this, but took one look at the slopfest README and it sorta scared me away. Will give it a go thanks.

i tought we can just use AI to read the readme and figure it out? install things.. i always do this

Re: GLM-5.3-Flash

#472

Earlier quoted context omitted.

At 50tps for single stream you are going to get 50 * 60 * 60 * 24 * 30 = 130M out tokens of GLM 5.3 Flash... That's less than what 40$ at current API rates... So if you are willing to pay 200$ per month you will get much better limits paying API rates. You can't run large Kimi K3 models on 10K worth of hardware either way, you need to spend like 50K USD minimum. Just pay for the API rates or get a low cost provider t…

Your point isn't lost on me, but a few other considerations: 1) Rates are theoretically discounted for GLM 5.3 Flash right now, by 50%. 2) Hardware costs have continued ascending with no sign of letting off, so it's unlikely that a DGX Spark depreciates to zero in one year. 3) Compare performance in terms of difficult tasks/$ over the last 6 months, 3 months, etc. Open weights are a ratchet. In terms of intelligence…

> 2) Hardware costs have continued ascending with no sign of letting off, so it's unlikely that a DGX Spark depreciates to zero in one year.

If someone told me that costs for X will keep increasing because they have been increasing rapidly in the last 1.5 years, but they have a history of continuously decreasing for decades before that.

I am not sure if I will take anything they say serious, I am not sure if it's HN or AI but people are delusional if they think compute costs will keep increasing from now on...

Either AI will be really good, hence compute and everything will materially depreciate or it won't be much better than it is today and token volumes will plateau compared to compute.

For instance the amount of token compute that's to come online in 6-12 months is several times what we have today...

Second 3) Compare performance in terms of difficult tasks/$ over the last 6 months, 3 months, etc. Open weights are a ratchet. In terms of intelligence per $, a Spark is never going to be a worse deal tomorrow than it is today, at least until the entire platform is replaced or obsoleted.

This is a bad take because again this assumes DGX Spark will not depreciate in price, we will have something better for far cheaper surely in the next couple years. M5 Max & Ultra are already arguably it, but will have to see.

> 71 days ago the best model you could run on two Sparks was an aggressive Q3 quant of Qwen 3.5 397B (AA 34). 70 days ago it was a mixed-quant of GLM 5.2 (AA 53). 30 days ago it was full fat DeepSeek 4 Flash (AA 53). Today it's GLM 5.3 Flash (AA57) and/or Qwen 3.8 Next (Unknown). Sometime this week it will likely become mixed-quant GLM 5.3 (AA 60).

This has nothing to do with DGX Spark's value, if models get cheaper the API costs also go down, this is not a defensible argument to cost to value.

Are people on HN really not thinking straight?

Tldr; no matter how you do the math compute is only getting more valuable because of a temporary crunch, don't expect this to continue permanently, sure you maybe able to time it and make money but so could you in stocks this is not for investments. Further second hand hardware sells for cheaper than sticker price, outside of a bubble..

And models getting cheaper == APIs getting cheaper == your hardware becoming worse value as your electricity & maintanence costs still remain.

I am not saying local models don't have their place but if someone is trying to use this logic to justify their purchase then I wish them all the best, as someone who is actively working on AI compute/inference/hardware stuff I personally don't have this level of courage.

But this is not a sound investment strategy that if something is going up and seems like it might keep going up, especially when investing in heavily depreciating assets like compute.

Re: GLM-5.3-Flash

#473

Earlier quoted context omitted.

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…

Deepseek and GLM answered correctly on Openrouter when using non-Chinese endpoints. I hope it stays that way!

correctly is very loaded here. correct according to whom? the truth is different though.

Re: GLM-5.3-Flash

#474
post #223

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips. Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified. whether it's t…

Weren't they giving free access? Not exacty a meaningful heuristic if so

Re: GLM-5.3-Flash

#475
post #96

Earlier quoted context omitted.

Compare to the cost of professional-grade tools in other trades and craft hobbies. Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies. And for some people, $4000 for a device you have complete control…

I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plan…

With multiple 200 a month subs you are getting a multiple of those subsidized tokens. At least if you tabulate at retail api prices.

This rent in the era of expensive hardware thing is not exclusive to inference.

I’ve needed x86 architecture for windows builds recently and have just hemmed and hawed over buying a decent windows 11 box.

I can’t make the math work against Azure instances.

I can spin up a nice one for build deallocate,spin up something cheaper for QA and then turn that off.

I can build all the devops around that, with a number of passes, with a skills based interface so working with the cloud is not too bad.

The only thing that still has me thinking about it is the prospect of price is going up even more, which is acid as far as I know.

And I’m hopefully going to need this x86 stuff enough that I don’t wanna wish I had gotten one for that high prices now.

Re: GLM-5.3-Flash

#476
post #419

Earlier quoted context omitted.

What it shows is that the CCP has enough oversight and control (either explicitly or by the companies making these decisions by default) that they will alter the models to benefit China. Who is to say they aren't doing it in other ways as well? That they aren't, or won't be, subtly hamstrung in engineering work? OAI and Anthropic have their own issues, you're right to be suspicious of them, but it's not like their mo…

If you don't want to use the Chinese model, then don't. Why attack it instead? Don't you want others to use it either?

The OP is pointing out issues that he thinks other people ought to consider before using Chinese models.

Re: GLM-5.3-Flash

#477

Earlier quoted context omitted.

The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

They are still industry leaders. They'll have to try to maintain that.

They have like a 3 month lead, and it takes unbelievable expenditure to maintain it.

Re: GLM-5.3-Flash

#478
post #426
post #228

Earlier quoted context omitted.

I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…

I can't get 3.8 to exit thinking loops. It will just think and think and think on the most trivial topics. I wanted it to port a speed test powershell script to c#. Claude opus 5 completes it under 60 seconds. I let 3.8 churn about 6 different times for 30+ minutes and it never wrote a single line of code to disk. It wrote lots of lines in thinking. unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL DSH (pi) Any tips?

That sounds like something is off - I'm using UD-Q4_K_XL on pi with xhigh thinking, and unless I'm vastly underestimating the complexity of the script that's the kind of task I would expect to take a couple of minutes (getting ~30t/s decode). What server are you running, and are you using the recommended parameters from qwen/unsloth?

Re: GLM-5.3-Flash

#479
post #393

Earlier quoted context omitted.

Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it. https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...

Qwen3.8 27B (which I adore) is nowhere near Opus 4.8 at puzzle games testing fluid intelligence, https://quesma.com/blog/baba-is-aug-2026/

yeah it's more like opus 4.6 iirc

Re: GLM-5.3-Flash

#480
post #330

Earlier quoted context omitted.

Yes, the terms are dubious. But they are also reasonably lenient with enforcement. They also don't require persona id verification, witch is wat turned me away from openai.

Yeah, you're probably right... > They also don't require persona id verification, witch is wat turned me away from openai. Could be worse. I was dumb enough to verify, only to get rejected for unknown reasons with no retries and no appeals. Had my privacy violated and have nothing to show for it.

Yeah, this is worse.

"Come verify your identity."

"Thanks, we've got your ID. Still not approving you, and there's no appeal."

Worst of both worlds.

Post reply on HN