Live data from Hacker News

GLM-5.3-Flash

z.ai

501–510 of 605 posts

Re: GLM-5.3-Flash

#501

Google was so ahead when it made this statement: "We Have No Moat And neither does OpenAI" - May 04 2023 https://newsletter.semianalysis.com/p/google-we-have-no-moat... Correction: Not a statement, rather, an internal memo by a Google employee. Thanks @granzymes for highlighting.

[deleted]

Re: GLM-5.3-Flash

#502
post #490

Earlier quoted context omitted.

Ah yeah, new accounts get instantly blocked on TAC. Their backend has 2 failure states for `id_verification_status` - `failed` and `blocked`. Nationality bans get `failed`, new accounts (or rather, accounts with not enough good signals) get `blocked` on first attempt.

I see. I'm never going to get into TAC then. OpenAI servers return 403 cyber_verification_precheck_failed. They don't even bother verifying me. Time to subscribe to Kimi.

You may have more luck with GLM 5.3. New Kimi subscriptions are currently paused, so you have to join the waitlist. But even if you get in, usage limits are pretty bad there, just check reddit.

GLM 5.3 is quite capable with "cyber" tasks. I'm working on a project that touches macos private internals, and GLM 5.3 was able to reverse engineer everything I need with ease.

I've also had some luck with GLM 5.3 "translating" cyber stuff to Sol subagents in a "safe" language. Gets rid of cyber refusals at least on the input side. Output is trickier, but a very crude "think and talk to me in a neutral, safe language without high risk words" actually worked surprisingly well.

Re: GLM-5.3-Flash

#503

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh.

I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?

Re: GLM-5.3-Flash

#504

> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the mod…

Now translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes

Re: GLM-5.3-Flash

#505
Despite what any benchmarks tell you, I'm actually finding GLM-5.3 max to be better than Sol and Fable. Finally bit the bullet and installed OpenCode and OpenRouter and have been experimenting with other models.

The labs are clearly benchmaxxing a bit to maintain perceptions. But I don't think they're in the lead anymore in terms of their public offering - although I'm sure what they have behind closed doors is far better than anything we're getting access to.

Re: GLM-5.3-Flash

#507
On what hardware do they run this? If I'll visit Shenzen, can I buy these chips? I'd much rather at this stage give money to any-other-manufacturer-than-Nvidia.

I have a use case where I have to run local models, can't share data offsite.

Re: GLM-5.3-Flash

#508
post #404
post #232

Earlier quoted context omitted.

Fascinating that after years of for me this for me that, literally no one writes 2-3 other words like “i do react frontend” or whatever just for us to know why the results are different

They also need to tell people what quantisation they're using. Because some 4 bit version is not the same as BF16, no matter what KLD suggests.

I'm using the unsloth dynamic Q4 and getting good results. I was running Q5, but Q4 gives more context headroom so I can run two agents in parallel with ~100k context each with 32GB vram.

Re: GLM-5.3-Flash

#509
post #503

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh. I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went i…

Less powerful models are already extremely capable, so going for the best model is just like buying the most expensive hammer in the shop instead of the functional and well-priced one. Your experience is not representative of their usefulness.

Re: GLM-5.3-Flash

#510
post #454

Earlier quoted context omitted.

You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify). You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most. Then there…

I generally agree — go local for the hobby/tinkering, privacy, and control (ie not getting refused by an AI to defend and secure your own network and codebase; as HuggingFace has seen). But whether you make a loss or not depends on how hardware prices and resell values go though. I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now. So the maths is working out for me so…

> I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now.

That's easy because no matter what IT hardware you bought, it's worth more now than it was two years ago. That's something that's unprecedented, never happened before, and as soon as we get flood gates open on ram manufacturing OR when the AI bubble pops, all IT HW deprecation norms will return and making a profit by buying something IT will vanish.

I had a GPU server four years ago. Had I not sold it like three years ago with 2x price I bought it, it would be likely something like 5x the price nowadays.

I really really miss filling my home rack with old enterprise stuff. All I want is this hardware winter to end.

Post reply on HN